Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1vi78L-00FK0H-2k for pgsql-hackers@arkaria.postgresql.org; Tue, 20 Jan 2026 08:30:42 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1vi78K-00GyXb-1G for pgsql-hackers@arkaria.postgresql.org; Tue, 20 Jan 2026 08:30:40 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1vi78K-00GyXT-07 for pgsql-hackers@lists.postgresql.org; Tue, 20 Jan 2026 08:30:40 +0000 Received: from mail-wm1-x334.google.com ([2a00:1450:4864:20::334]) by magus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.96) (envelope-from ) id 1vi78H-001U8K-2r for pgsql-hackers@lists.postgresql.org; Tue, 20 Jan 2026 08:30:39 +0000 Received: by mail-wm1-x334.google.com with SMTP id 5b1f17b1804b1-47ee3a63300so47741375e9.2 for ; Tue, 20 Jan 2026 00:30:37 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cybertec.at; s=google; t=1768897834; x=1769502634; darn=lists.postgresql.org; h=message-id:date:mime-version:comments:references:in-reply-to :subject:to:from:from:to:cc:subject:date:message-id:reply-to; bh=KtmxOeUu5wg5aASJV0fpWLGVEmNd9ynMnxO0Jv4OsVU=; b=eq60aPX2GQWKWpD4n+DIXcsA7S7++FeadHESlH/itTmuhj3IrdGISiFXU7qGZ8S1dT 5yCVnP6ekE8G56oQOY4pwtRZ51XjKztI0TURlSnTM3w9E0zLVbmfYXfbDRGNeoFAxEue TsbmRcGnVz3YDZVr0/nsGFT0KZAlqaOC5+SXx46zP2YCa/LQ4Md5p0a/hxIYwe8/CB7G r0qiLj4LAiq4pK/SrG+HzBDdx4152a/KFqZ885IpkYU/a/wkhGLLFHXGktnCd3Z3nNik oA8feJuR0WmsxiY2TcOeJTSikDr+Uwz7Po7cC0fEC3PId4xH1KWVQolWg4P5QGKL5aA0 svzQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1768897834; x=1769502634; h=message-id:date:mime-version:comments:references:in-reply-to :subject:to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=KtmxOeUu5wg5aASJV0fpWLGVEmNd9ynMnxO0Jv4OsVU=; b=R7p6A3WyzhNkqYK935tDdq84Kp1WVRWIOcNd4b4qcsm8V29pqyl1dJmtHii/L+yzTd pPZwT3zNvbpIu8d0UcFiLF3ZP2YD6jZrCIvVIHeMxy0ypWnR74uU3hU9NgwQy1JahZmy CrUV2VxB2Ca9pUV26UXCMXdy3JdabQNKgbBrWJxyfMTxbSwfgDNhhnaSVx7xOtZOzxiJ Jk5FFMdvPG6C7hsQobpPWXJkVpE0jbAxGDxMPuSoW+rzgyIziF1eixXOsjfeTidJUuMO i8Q5bvKFVr3u2V/dnkxLxzEv16BsWcR6FGpJnHre3VSvEc2dihNYnjmVW4trPZq0tN3T Udfw== X-Gm-Message-State: AOJu0YxsJKvq/L3seb8AQQuJOLs6AbonzDeAS+bYgrZCL+b5L5BL6H5N qMJ60OpHooouXu0Bxt6VsRtIcP9CpRT77O5v3sEPIXV4H+AldDFPymTJdBRIbMpxI2IasbuGSqq bw02z X-Gm-Gg: AY/fxX5TV5uSlzCzbSwiIjG/PofBtn7FxGR/Rb5VmM9F/x3LxcjkpX4ts3YUGYKekKj cVtvOJ1yAnEKarKQq4WzwwAx/RU7CEYamMpN/6tfq3P+VBIFWlitR0PX9FLhWEQDer1PxOn6GYc 3rAk54ha78NjdL/n4HT2AKbw++OjPXJjt7bYs61H9cv0Y6dRB4oHc8HZazJSIcogT/FqeSw7lZX pv71G/izBeUYdgQgNG0Wt126NEEGce46y0EaccZMVqe/2omdUQ02z9WbmQbGs2Y6Zuc1d/yIFS1 +djGIEW+y7AYcIz4FjwQ11ZS2jyWAB/tKLMq4OK+ZJbl6OHqD8VT7R71d6s2sT4uSdyN0iybt32 bF7G9OBnCocWzmqu02jWmfgVJdgEnTmS6DgO2pUT9nV/kJB+/NgjAEWSfvKUlfZ7CF8A/sxJ1SB 4QynfJx/+SXsi6z8fnIQpwYwOs X-Received: by 2002:a05:600c:a013:b0:477:79c7:8994 with SMTP id 5b1f17b1804b1-4803e7f0e39mr13035575e9.30.1768897834397; Tue, 20 Jan 2026 00:30:34 -0800 (PST) Received: from localhost (109-81-168-246.rct.o2.cz. [109.81.168.246]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4801e9434d9sm107388465e9.0.2026.01.20.00.30.33 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 20 Jan 2026 00:30:33 -0800 (PST) From: Antonin Houska To: pgsql-hackers@lists.postgresql.org Subject: Re: Race conditions in logical decoding In-reply-to: <85833.1768840165@localhost> References: <85833.1768840165@localhost> Comments: In-reply-to Antonin Houska message dated "Mon, 19 Jan 2026 17:29:25 +0100." X-Mailer: MH-E 8.6+git; nmh 1.8; GNU Emacs 28.3 MIME-Version: 1.0 Content-Type: multipart/mixed; boundary="=-=-=" Date: Tue, 20 Jan 2026 09:30:33 +0100 Message-ID: <62335.1768897833@localhost> List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk --=-=-= Content-Type: text/plain Antonin Houska wrote: > I'm not sure yet how to fix the problem. I tried to call XactLockTableWait() > from SnapBuildAddCommittedTxn() (like it happens in SnapBuildWaitSnapshot()), > but it made at least one regression test (subscription/t/010_truncate.pl) > stuck - probably a deadlock. I can spend more time on it, but maybe someone > can come up with a good idea sooner than me. Attached here is what I consider a possible fix - simply wait for the CLOG update before building a new snapshot. Unfortunately I have no idea right now how to test it using the isolation tester. With the fix, the additional waiting makes the current test block. (And if a step is added that unblock the session, it will not reliably catch failure to wait.) -- Antonin Houska Web: https://www.cybertec-postgresql.com --=-=-= Content-Type: text/x-diff Content-Disposition: attachment; filename=0001-Fix-race-conditions-during-the-setup-of-logical-deco.patch From 5a6002215fb8ebeaf1dde120e5f6706bca7b62ae Mon Sep 17 00:00:00 2001 From: Antonin Houska Date: Mon, 19 Jan 2026 16:07:45 +0100 Subject: [PATCH] Fix race conditions during the setup of logical decoding. Although it's rather unlikely, it can happen that the snapshot builder considers transaction committed (according to WAL) before the commit could be recorded in CLOG. In an extreme case, snapshot can even be created and used in between. Since both snapshot and CLOG are needed for visibility checks, this inconsistency can make them work incorrectly. The typical symptom is that a transaction that the snapshot considers not running anymore is (per CLOG) considered aborted instead of committed. Thus a new tuple version can be evaluated as invisible (if xmin is incorrectly considered aborted) or a deleted tuple version can be evaluated as visible (if xmax is incorrectly considered aborted). This patch fixes the problem by checking if all the XIDs that the new snapshot considers committed are really committed per CLOG. If at least one is not, the check is repeated after a short delay. However, a single check is sufficient in almost all cases, so the performance impact should be minimal. --- src/backend/replication/logical/snapbuild.c | 41 +++++++++++++++++++ .../utils/activity/wait_event_names.txt | 1 + 2 files changed, 42 insertions(+) diff --git a/src/backend/replication/logical/snapbuild.c b/src/backend/replication/logical/snapbuild.c index 9b09dc8eac1..c02d08f3417 100644 --- a/src/backend/replication/logical/snapbuild.c +++ b/src/backend/replication/logical/snapbuild.c @@ -400,6 +400,47 @@ SnapBuildBuildSnapshot(SnapBuild *builder) snapshot->xmin = builder->xmin; snapshot->xmax = builder->xmax; + /* + * Although it's very unlikely, it's possible that a commit WAL record was + * decoded but CLOG is not aware of the commit yet. Should the CLOG update + * be delayed even more, visibility checks that use this snapshot could + * work incorrectly. Therefore we check the CLOG status here. + */ + while (true) + { + bool found = false; + + for (int i = 0; i < builder->committed.xcnt; i++) + { + /* + * XXX Is it worth remembering the XIDs that appear to be + * committed per CLOG and skipping them in the next iteration of + * the outer loop? Not sure it's worth the effort - a single + * iteration is enough in most cases. + */ + if (unlikely(!TransactionIdDidCommit(builder->committed.xip[i]))) + { + found = true; + + /* + * Wait a bit before going to the next iteration of the outer + * loop. The race conditions we address here is pretty rare, + * so we shouldn't need to wait too long. + */ + (void) WaitLatch(MyLatch, + WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH, + 10L, + WAIT_EVENT_SNAPBUILD_CLOG); + ResetLatch(MyLatch); + + break; + } + } + + if (!found) + break; + } + /* store all transactions to be treated as committed by this snapshot */ snapshot->xip = (TransactionId *) ((char *) snapshot + sizeof(SnapshotData)); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 5537a2d2530..b6318b0cf37 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -181,6 +181,7 @@ PG_SLEEP "Waiting due to a call to pg_sleep or a sibling fu RECOVERY_APPLY_DELAY "Waiting to apply WAL during recovery because of a delay setting." RECOVERY_RETRIEVE_RETRY_INTERVAL "Waiting during recovery when WAL data is not available from any source (pg_wal, archive or stream)." REGISTER_SYNC_REQUEST "Waiting while sending synchronization requests to the checkpointer, because the request queue is full." +SNAPBUILD_CLOG "Waiting for CLOG update before building snapshot." SPIN_DELAY "Waiting while acquiring a contended spinlock." VACUUM_DELAY "Waiting in a cost-based vacuum delay point." VACUUM_TRUNCATE "Waiting to acquire an exclusive lock to truncate off any empty pages at the end of a table vacuumed." -- 2.47.3 --=-=-=--