agora inbox for [email protected]help / color / mirror / Atom feed
[PATCH] Lock upgrade without deadlocks. 172+ messages / 2 participants [nested] [flat]
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 172+ messages in thread
* [PATCH v3 2/2] Try to avoid a rewrite when adding a stored generated column @ 2026-04-24 08:44 Alberto Piai <[email protected]> 0 siblings, 0 replies; 172+ messages in thread From: Alberto Piai @ 2026-04-24 08:44 UTC (permalink / raw) This builds upon basic support for ... ALTER COLUMN ... ADD GENERATED ALWAYS AS (expr) STORED If we can find a constraint which proves that the given column is already always equal to the new generated column expression, skip the expensive rewrite of the table. The check constraint must use an equality operator which is mergejoinable, and the expression must match exactly the generated column's default expression. --- src/backend/catalog/pg_constraint.c | 75 +++++++ src/backend/commands/tablecmds.c | 46 +++-- src/include/catalog/pg_constraint.h | 2 + src/test/regress/expected/alter_table.out | 232 +++++++++++++++++++++- src/test/regress/sql/alter_table.sql | 147 ++++++++++++++ 5 files changed, 481 insertions(+), 21 deletions(-) diff --git a/src/backend/catalog/pg_constraint.c b/src/backend/catalog/pg_constraint.c index b12765ae691..cfb0a0068cf 100644 --- a/src/backend/catalog/pg_constraint.c +++ b/src/backend/catalog/pg_constraint.c @@ -17,6 +17,7 @@ #include "access/genam.h" #include "access/gist.h" #include "access/htup_details.h" +#include "access/relation.h" #include "access/sysattr.h" #include "access/table.h" #include "catalog/catalog.h" @@ -29,6 +30,8 @@ #include "catalog/pg_type.h" #include "commands/defrem.h" #include "common/int.h" +#include "nodes/nodeFuncs.h" +#include "parser/parse_relation.h" #include "utils/array.h" #include "utils/builtins.h" #include "utils/fmgroids.h" @@ -694,6 +697,78 @@ findDomainNotNullConstraint(Oid typid) return retval; } +/* + * Given a relation, an attnum and a (cooked) expression, this returns true if + * it finds a CHECK constraint which proves that the given column is equal to + * the expression. + * + * The constraint must use a mergejoinable operator for the type of the column, + * a concept used by the planner as well to infer equivalence classes on the + * terms in a query (see op_mergejoinable()). + * + * The expressions are compared structurally, so they must match exactly for + * this check to succeed. + */ +bool +findStructuralCheckConstraintOnAttr(Oid relid, AttrNumber attnum, + const Node *target_expr) +{ + Relation pg_constraint; + HeapTuple conTup; + SysScanDesc scan; + ScanKeyData key; + bool found = false; + + pg_constraint = table_open(ConstraintRelationId, AccessShareLock); + ScanKeyInit(&key, + Anum_pg_constraint_conrelid, + BTEqualStrategyNumber, F_OIDEQ, + ObjectIdGetDatum(relid)); + scan = systable_beginscan(pg_constraint, ConstraintRelidTypidNameIndexId, + true, NULL, 1, &key); + + while (HeapTupleIsValid(conTup = systable_getnext(scan))) + { + Form_pg_constraint con = GETSTRUCT(conTup); + char *conbin; + Datum val; + Node *conexpr; + + if (con->contype != CONSTRAINT_CHECK) + continue; + if (!con->convalidated) + continue; + + val = SysCacheGetAttrNotNull(CONSTROID, conTup, + Anum_pg_constraint_conbin); + conbin = TextDatumGetCString(val); + conexpr = stringToNode(conbin); + + if (IsA(conexpr, OpExpr)) + { + OpExpr *op = (OpExpr *) conexpr; + + if (list_length(op->args) == 2 && IsA(linitial(op->args), Var)) + { + Var *var = linitial(op->args); + + if (var->varattno == attnum && + op_mergejoinable(op->opno, exprType((Node *) var)) && + equal(lsecond(op->args), target_expr)) + { + found = true; + break; + } + } + } + } + + systable_endscan(scan); + table_close(pg_constraint, AccessShareLock); + + return found; +} + /* * Given a pg_constraint tuple for a not-null constraint, return the column * number it is for. diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c index aa54c629f8b..c886e6e5ca7 100644 --- a/src/backend/commands/tablecmds.c +++ b/src/backend/commands/tablecmds.c @@ -8938,6 +8938,7 @@ ATExecAddGeneratedAsExprStored(AlteredTableInfo *tab, NewColumnValue *newval; RawColumnDefault *rawEnt; Relation pg_attribute; + bool rewrite; Assert(def->raw_expr != NULL); Assert(def->cooked_expr == NULL); @@ -9001,29 +9002,36 @@ ATExecAddGeneratedAsExprStored(AlteredTableInfo *tab, /* Make above changes visible */ CommandCounterIncrement(); - /* - * Clear all the missing values if we're rewriting the table, since this - * renders them pointless. - */ - RelationClearMissing(rel); - - /* Make above changes visible */ - CommandCounterIncrement(); - - /* Drop any pg_statistic entry for the column */ - RemoveStatistics(RelationGetRelid(rel), attnum); - /* Build a concrete expression for the new default (generated) value */ defval = (Expr *) build_column_default(rel, attnum); defval = expression_planner(defval); - /* Schedule a rewrite */ - newval = palloc0_object(NewColumnValue); - newval->attnum = attnum; - newval->expr = defval; - newval->is_generated = true; - tab->newvals = lappend(tab->newvals, newval); - tab->rewrite |= AT_REWRITE_DEFAULT_VAL; + rewrite = !findStructuralCheckConstraintOnAttr(RelationGetRelid(rel), + attnum, + (Node *) defval); + + if (rewrite) + { + /* + * Clear all the missing values if we're rewriting the table, since + * this renders them pointless. + */ + RelationClearMissing(rel); + + /* Make above changes visible */ + CommandCounterIncrement(); + + /* Drop any pg_statistic entry for the column */ + RemoveStatistics(RelationGetRelid(rel), attnum); + + /* Schedule a rewrite */ + newval = palloc0_object(NewColumnValue); + newval->attnum = attnum; + newval->expr = defval; + newval->is_generated = true; + tab->newvals = lappend(tab->newvals, newval); + tab->rewrite |= AT_REWRITE_DEFAULT_VAL; + } InvokeObjectPostAlterHook(RelationRelationId, RelationGetRelid(rel), attnum); diff --git a/src/include/catalog/pg_constraint.h b/src/include/catalog/pg_constraint.h index 1b7fedf1750..7ac9e00c28b 100644 --- a/src/include/catalog/pg_constraint.h +++ b/src/include/catalog/pg_constraint.h @@ -266,6 +266,8 @@ extern char *ChooseConstraintName(const char *name1, const char *name2, extern HeapTuple findNotNullConstraintAttnum(Oid relid, AttrNumber attnum); extern HeapTuple findNotNullConstraint(Oid relid, const char *colname); extern HeapTuple findDomainNotNullConstraint(Oid typid); +extern bool findStructuralCheckConstraintOnAttr(Oid relid, AttrNumber attnum, + const Node *target_expr); extern AttrNumber extractNotNullColumn(HeapTuple constrTup); extern bool AdjustNotNullInheritance(Oid relid, AttrNumber attnum, const char *new_conname, bool is_local, bool is_no_inherit, bool is_notvalid); diff --git a/src/test/regress/expected/alter_table.out b/src/test/regress/expected/alter_table.out index 08981f4e380..5346791fd8d 100644 --- a/src/test/regress/expected/alter_table.out +++ b/src/test/regress/expected/alter_table.out @@ -4955,6 +4955,233 @@ select :idx_filenode_before != :idx_filenode_after as did_rewrite_idx; t (1 row) +drop table testgen.t3; +-- turning a regular column into a stored generated column +-- without rewriting the table (when a check constraint proves it isn't needed) +create table testgen.t4 (a int, b int not null); +insert into testgen.t4 (a, b) select x, x * 2 from generate_series(0, 5) x; +alter table testgen.t4 add constraint chk_gen_clause check (b = a * 2); +select pg_relation_filenode('testgen.t4') as t4_filenode_before \gset +alter table testgen.t4 alter column b add generated always as (a * 2) stored; +select pg_relation_filenode('testgen.t4') as t4_filenode_after \gset +select :t4_filenode_before = :t4_filenode_after as did_skip_rewrite; + did_skip_rewrite +------------------ + t +(1 row) + +\d+ testgen.t4 + Table "testgen.t4" + Column | Type | Collation | Nullable | Default | Storage | Stats target | Description +--------+---------+-----------+----------+------------------------------------+---------+--------------+------------- + a | integer | | | | plain | | + b | integer | | not null | generated always as (a * 2) stored | plain | | +Check constraints: + "chk_gen_clause" CHECK (b = (a * 2)) +Not-null constraints: + "t4_b_not_null" NOT NULL "b" + +drop table testgen.t4; +-- turning a regular column into a stored generated column +-- same as the previous case, but a rewrite happens since the constraint is not +-- valid +create table testgen.t4 (a int, b int not null); +insert into testgen.t4 (a, b) select x, x * 2 from generate_series(0, 5) x; +alter table testgen.t4 add constraint chk_gen_clause check (b = a * 2) not valid; +select pg_relation_filenode('testgen.t4') as t4_filenode_before \gset +alter table testgen.t4 alter column b add generated always as (a * 2) stored; +select pg_relation_filenode('testgen.t4') as t4_filenode_after \gset +select :t4_filenode_before != :t4_filenode_after as did_rewrite; + did_rewrite +------------- + t +(1 row) + +\d+ testgen.t4 + Table "testgen.t4" + Column | Type | Collation | Nullable | Default | Storage | Stats target | Description +--------+---------+-----------+----------+------------------------------------+---------+--------------+------------- + a | integer | | | | plain | | + b | integer | | not null | generated always as (a * 2) stored | plain | | +Check constraints: + "chk_gen_clause" CHECK (b = (a * 2)) NOT VALID +Not-null constraints: + "t4_b_not_null" NOT NULL "b" + +drop table testgen.t4; +-- turning a regular column into a stored generated column +-- same as the previous case, but a rewrite happens since the constraint +-- operator is not mergejoinable +create table testgen.t4 (a int, b int not null); +insert into testgen.t4 (a, b) select x, x * 2 from generate_series(0, 5) x; +alter table testgen.t4 add constraint chk_gen_clause check (b >= a * 2); +select pg_relation_filenode('testgen.t4') as t4_filenode_before \gset +alter table testgen.t4 alter column b add generated always as (a * 3) stored; +select pg_relation_filenode('testgen.t4') as t4_filenode_after \gset +select :t4_filenode_before != :t4_filenode_after as did_rewrite; + did_rewrite +------------- + t +(1 row) + +\d+ testgen.t4 + Table "testgen.t4" + Column | Type | Collation | Nullable | Default | Storage | Stats target | Description +--------+---------+-----------+----------+------------------------------------+---------+--------------+------------- + a | integer | | | | plain | | + b | integer | | not null | generated always as (a * 3) stored | plain | | +Check constraints: + "chk_gen_clause" CHECK (b >= (a * 2)) +Not-null constraints: + "t4_b_not_null" NOT NULL "b" + +drop table testgen.t4; +-- test the whole process for adding a stored generated column without +-- long-lived exclusive locks +create table testgen.t5 (a int); +select pg_relation_filenode('testgen.t5') as t5_filenode_before \gset +insert into testgen.t5 select x from generate_series(1, 5) x; +alter table testgen.t5 add column b int; +-- take care of new and updated columns +create function testgen.gen () returns trigger language plpgsql as $$ +begin + new.b = new.a * 2; return new; +end +$$; +create trigger testgen_gen + before insert or update on testgen.t5 + for each row execute function testgen.gen(); +-- add the constraint as not valid: enforced only for new and updated rows +begin; +alter table testgen.t5 + add constraint chk_gen_clause check (b = a * 2) not valid; +select locktype, mode from pg_locks + where relation = 'testgen.t5'::regclass and granted; + locktype | mode +----------+--------------------- + relation | AccessExclusiveLock +(1 row) + +commit; +insert into testgen.t5 (a) values (100), (200), (300); +-- backfill existing rows at the appropriate pace +update testgen.t5 set b = a * 2 where b is null; +-- validate: this scans the table, but without an exclusive lock +begin; +alter table testgen.t5 validate constraint chk_gen_clause; +select locktype, mode from pg_locks + where relation = 'testgen.t5'::regclass and granted; + locktype | mode +----------+-------------------------- + relation | ShareUpdateExclusiveLock +(1 row) + +commit; +-- now the schema update, which skips the rewrite because of the check +begin; +drop trigger testgen_gen on testgen.t5; +alter table testgen.t5 alter column b + add generated always as (a * 2) stored; +select locktype, mode from pg_locks +where relation = 'testgen.t5'::regclass and granted; + locktype | mode +----------+--------------------- + relation | AccessShareLock + relation | AccessExclusiveLock +(2 rows) + +commit; +select pg_relation_filenode('testgen.t5') as t5_filenode_after \gset +select :t5_filenode_before = :t5_filenode_after as did_skip_rewrite; + did_skip_rewrite +------------------ + t +(1 row) + +\d+ testgen.t5 + Table "testgen.t5" + Column | Type | Collation | Nullable | Default | Storage | Stats target | Description +--------+---------+-----------+----------+------------------------------------+---------+--------------+------------- + a | integer | | | | plain | | + b | integer | | | generated always as (a * 2) stored | plain | | +Check constraints: + "chk_gen_clause" CHECK (b = (a * 2)) + +-- test support for partitioned tables and inheritance +create table testgen.tpart (a int, b int) partition by hash (a); +create table testgen.tpart_p1 partition of testgen.tpart + for values with (modulus 2, remainder 0); +create table testgen.tpart_p2 partition of testgen.tpart + for values with (modulus 2, remainder 1); +insert into testgen.tpart (a, b) select x, x from generate_series(1, 5) x; +-- altering the parent table, recursing +begin; +alter table testgen.tpart alter column b + add generated always as (a * 2) stored; +-- expected: all the partitions have been rewritten +select a, b, a * 2 as expected, b = (a * 2) as correct + from testgen.tpart_p1 order by a; + a | b | expected | correct +---+---+----------+--------- + 1 | 2 | 2 | t + 2 | 4 | 4 | t +(2 rows) + +select a, b, a * 2 as expected, b = (a * 2) as correct + from testgen.tpart_p2 order by a; + a | b | expected | correct +---+----+----------+--------- + 3 | 6 | 6 | t + 4 | 8 | 8 | t + 5 | 10 | 10 | t +(3 rows) + +rollback; +-- altering a single partition is not allowed +begin; +-- expected: error +alter table testgen.tpart_p1 alter column b + add generated always as (a * 2) stored; +ERROR: cannot change inherited column to be a stored generated column +rollback; +-- altering only the parent table is not allowed +begin; +-- expected: error +alter table only testgen.tpart alter column b + add generated always as (a * 2) stored; +ERROR: ALTER TABLE / ADD GENERATED ALWAYS AS (expr) STORED must be applied to child tables too +rollback; +drop table testgen.tpart; +-- subpartitions +create table testgen.tpart (a int, b int, c int) + partition by hash (a); +create table testgen.tpart_p1 partition of testgen.tpart + for values with (modulus 2, remainder 0) + partition by hash (b); +create table testgen.tpart_p1_1 partition of testgen.tpart_p1 + for values with (modulus 2, remainder 0); +create table testgen.tpart_p1_2 partition of testgen.tpart_p1 + for values with (modulus 2, remainder 1); +create table testgen.tpart_p2 partition of testgen.tpart + for values with (modulus 2, remainder 1) + partition by hash (b); +create table testgen.tpart_p2_1 partition of testgen.tpart_p2 + for values with (modulus 2, remainder 0); +create table testgen.tpart_p2_2 partition of testgen.tpart_p2 + for values with (modulus 2, remainder 1); +insert into testgen.tpart (a, b) + select x, y + from generate_series(1, 5) x + cross join generate_series(1, 5) y; +-- currently, it is not possible to change the generated state of an +-- inheritance tree of depth >= 2 (same as in DROP EXPRESSION), so we expect an +-- error here. This might be fixed later. +begin; +alter table testgen.tpart alter column c + add generated always as (a + b) stored; +ERROR: ALTER TABLE / ADD GENERATED ALWAYS AS (expr) STORED must be applied to child tables too +rollback; +drop table testgen.tpart; -- tests for invalid invocations alter table doesnotexist alter column foo add generated always as (bar * 2) stored; @@ -4998,6 +5225,7 @@ alter table testgen.t3 alter column b ERROR: cannot use subquery in column generation expression drop table testgen.t3; drop schema testgen cascade; -NOTICE: drop cascades to 4 other objects -DETAIL: drop cascades to table testgen.t3 +NOTICE: drop cascades to 3 other objects +DETAIL: drop cascades to table testgen.t5 +drop cascades to function testgen.gen() drop cascades to table testgen.t1 diff --git a/src/test/regress/sql/alter_table.sql b/src/test/regress/sql/alter_table.sql index 76187083289..efb08b9ff70 100644 --- a/src/test/regress/sql/alter_table.sql +++ b/src/test/regress/sql/alter_table.sql @@ -3198,6 +3198,153 @@ select pg_relation_filenode('testgen.idx_b') as idx_filenode_after \gset select :idx_filenode_before != :idx_filenode_after as did_rewrite_idx; drop table testgen.t3; +-- turning a regular column into a stored generated column +-- without rewriting the table (when a check constraint proves it isn't needed) +create table testgen.t4 (a int, b int not null); +insert into testgen.t4 (a, b) select x, x * 2 from generate_series(0, 5) x; +alter table testgen.t4 add constraint chk_gen_clause check (b = a * 2); +select pg_relation_filenode('testgen.t4') as t4_filenode_before \gset +alter table testgen.t4 alter column b add generated always as (a * 2) stored; +select pg_relation_filenode('testgen.t4') as t4_filenode_after \gset +select :t4_filenode_before = :t4_filenode_after as did_skip_rewrite; +\d+ testgen.t4 +drop table testgen.t4; + +-- turning a regular column into a stored generated column +-- same as the previous case, but a rewrite happens since the constraint is not +-- valid +create table testgen.t4 (a int, b int not null); +insert into testgen.t4 (a, b) select x, x * 2 from generate_series(0, 5) x; +alter table testgen.t4 add constraint chk_gen_clause check (b = a * 2) not valid; +select pg_relation_filenode('testgen.t4') as t4_filenode_before \gset +alter table testgen.t4 alter column b add generated always as (a * 2) stored; +select pg_relation_filenode('testgen.t4') as t4_filenode_after \gset +select :t4_filenode_before != :t4_filenode_after as did_rewrite; +\d+ testgen.t4 +drop table testgen.t4; + +-- turning a regular column into a stored generated column +-- same as the previous case, but a rewrite happens since the constraint +-- operator is not mergejoinable +create table testgen.t4 (a int, b int not null); +insert into testgen.t4 (a, b) select x, x * 2 from generate_series(0, 5) x; +alter table testgen.t4 add constraint chk_gen_clause check (b >= a * 2); +select pg_relation_filenode('testgen.t4') as t4_filenode_before \gset +alter table testgen.t4 alter column b add generated always as (a * 3) stored; +select pg_relation_filenode('testgen.t4') as t4_filenode_after \gset +select :t4_filenode_before != :t4_filenode_after as did_rewrite; +\d+ testgen.t4 +drop table testgen.t4; + +-- test the whole process for adding a stored generated column without +-- long-lived exclusive locks +create table testgen.t5 (a int); +select pg_relation_filenode('testgen.t5') as t5_filenode_before \gset +insert into testgen.t5 select x from generate_series(1, 5) x; +alter table testgen.t5 add column b int; +-- take care of new and updated columns +create function testgen.gen () returns trigger language plpgsql as $$ +begin + new.b = new.a * 2; return new; +end +$$; +create trigger testgen_gen + before insert or update on testgen.t5 + for each row execute function testgen.gen(); +-- add the constraint as not valid: enforced only for new and updated rows +begin; +alter table testgen.t5 + add constraint chk_gen_clause check (b = a * 2) not valid; +select locktype, mode from pg_locks + where relation = 'testgen.t5'::regclass and granted; +commit; +insert into testgen.t5 (a) values (100), (200), (300); +-- backfill existing rows at the appropriate pace +update testgen.t5 set b = a * 2 where b is null; +-- validate: this scans the table, but without an exclusive lock +begin; +alter table testgen.t5 validate constraint chk_gen_clause; +select locktype, mode from pg_locks + where relation = 'testgen.t5'::regclass and granted; +commit; +-- now the schema update, which skips the rewrite because of the check +begin; +drop trigger testgen_gen on testgen.t5; +alter table testgen.t5 alter column b + add generated always as (a * 2) stored; +select locktype, mode from pg_locks +where relation = 'testgen.t5'::regclass and granted; +commit; +select pg_relation_filenode('testgen.t5') as t5_filenode_after \gset +select :t5_filenode_before = :t5_filenode_after as did_skip_rewrite; +\d+ testgen.t5 + +-- test support for partitioned tables and inheritance +create table testgen.tpart (a int, b int) partition by hash (a); +create table testgen.tpart_p1 partition of testgen.tpart + for values with (modulus 2, remainder 0); +create table testgen.tpart_p2 partition of testgen.tpart + for values with (modulus 2, remainder 1); +insert into testgen.tpart (a, b) select x, x from generate_series(1, 5) x; + +-- altering the parent table, recursing +begin; +alter table testgen.tpart alter column b + add generated always as (a * 2) stored; +-- expected: all the partitions have been rewritten +select a, b, a * 2 as expected, b = (a * 2) as correct + from testgen.tpart_p1 order by a; +select a, b, a * 2 as expected, b = (a * 2) as correct + from testgen.tpart_p2 order by a; +rollback; + +-- altering a single partition is not allowed +begin; +-- expected: error +alter table testgen.tpart_p1 alter column b + add generated always as (a * 2) stored; +rollback; + +-- altering only the parent table is not allowed +begin; +-- expected: error +alter table only testgen.tpart alter column b + add generated always as (a * 2) stored; +rollback; + +drop table testgen.tpart; + +-- subpartitions +create table testgen.tpart (a int, b int, c int) + partition by hash (a); +create table testgen.tpart_p1 partition of testgen.tpart + for values with (modulus 2, remainder 0) + partition by hash (b); +create table testgen.tpart_p1_1 partition of testgen.tpart_p1 + for values with (modulus 2, remainder 0); +create table testgen.tpart_p1_2 partition of testgen.tpart_p1 + for values with (modulus 2, remainder 1); +create table testgen.tpart_p2 partition of testgen.tpart + for values with (modulus 2, remainder 1) + partition by hash (b); +create table testgen.tpart_p2_1 partition of testgen.tpart_p2 + for values with (modulus 2, remainder 0); +create table testgen.tpart_p2_2 partition of testgen.tpart_p2 + for values with (modulus 2, remainder 1); +insert into testgen.tpart (a, b) + select x, y + from generate_series(1, 5) x + cross join generate_series(1, 5) y; +-- currently, it is not possible to change the generated state of an +-- inheritance tree of depth >= 2 (same as in DROP EXPRESSION), so we expect an +-- error here. This might be fixed later. +begin; +alter table testgen.tpart alter column c + add generated always as (a + b) stored; +rollback; + +drop table testgen.tpart; + -- tests for invalid invocations alter table doesnotexist alter column foo add generated always as (bar * 2) stored; -- 2.47.0 --4wx636ozsu2eakgd-- ^ permalink raw reply [nested|flat] 172+ messages in thread
end of thread, other threads:[~2026-04-24 08:44 UTC | newest] Thread overview: 172+ messages (download: mbox mbox.gz follow: Atom feed) -- links below jump to the message on this page -- 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <[email protected]> 2026-04-24 08:44 [PATCH v3 2/2] Try to avoid a rewrite when adding a stored generated column Alberto Piai <[email protected]>
This inbox is served by agora; see mirroring instructions for how to clone and mirror all data and code used for this inbox