Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wklUC-000rSz-1Q for pgsql-bugs@arkaria.postgresql.org; Fri, 17 Jul 2026 16:32:28 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1wklUB-001BTa-1A for pgsql-bugs@arkaria.postgresql.org; Fri, 17 Jul 2026 16:32:27 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wkggl-00HG5A-3C for pgsql-bugs@lists.postgresql.org; Fri, 17 Jul 2026 11:25:07 +0000 Received: from mahout.postgresql.org ([2001:4800:3e1:1::227]) by makus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1wkggi-00000000f9S-3pNI for pgsql-bugs@lists.postgresql.org; Fri, 17 Jul 2026 11:25:07 +0000 DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=postgresql.org; s=20171124; h=Message-ID:Date:Reply-To:Cc:From:To:Subject: Content-Transfer-Encoding:MIME-Version:Content-Type:Sender:Content-ID: Content-Description:In-Reply-To:References; bh=S4BSx/I1X5QOU0MdP2nPBazB5L++O/iJPbVh/5aekoE=; b=mOdl4+EJHC+Ike5x74pUuuFefr 0Gq0Z/TZJSyl4d5alUVjt/hISpd6/V1B4+rqmd3KjecAanopIGA3LqgC0Jlst5I9Ey5Y3vhzU10xj hsGCE4cA0L/CVQbDs8zpPyU98p7zYVckmaDyi20JKMXDUzNs1IPEBY5wlHqfsOfsEFhN7nvlPDn/3 jHxniL/g0jx/G6/m/+wkUu5abxCQs16J3umubK6XJa5JEovWbd/3bCwaKkAzXFs1x3XFfCcrZftuZ 36j/w9//JzUH6Z29w9Xf6GnLdOxGKXzIdV+TqRS9P6qd5tIsA3dkVMJ8jwt+sqXLxsPJgaBIVpi/2 AJSgmY4Q==; Received: from wrigleys.postgresql.org ([2a02:16a8:dc51::60]) by mahout.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wkggi-001zfj-27 for pgsql-bugs@lists.postgresql.org; Fri, 17 Jul 2026 11:25:05 +0000 Received: from localhost ([127.0.0.1] helo=wrigleys.postgresql.org) by wrigleys.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1wkggg-004bsR-0b for pgsql-bugs@lists.postgresql.org; Fri, 17 Jul 2026 11:25:03 +0000 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Subject: BUG #19557: Parallel leader stuckat ParallelFinish leaked InterruptHoldoffCount To: pgsql-bugs@lists.postgresql.org From: PG Bug reporting form Cc: jacob@brejnbjerg.com Reply-To: jacob@brejnbjerg.com, pgsql-bugs@lists.postgresql.org Date: Fri, 17 Jul 2026 11:24:48 +0000 Message-ID: <19557-d88cb23f38eb9b91@postgresql.org> X-Auto-Response-Suppress: All Auto-Submitted: auto-generated List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk The following bug has been logged on the website: Bug reference: 19557 Logged by: Jacob M=C3=B8rk Brejnbjerg Email address: jacob@brejnbjerg.com PostgreSQL version: 18.4 Operating system: Ubuntu 24.04 (kernel 6.8.0-124-generic) Description: =20 A backend executing a parallel query has been stuck for 40+ hours in wait_event IPC/ParallelFinish. All of its parallel workers exited cleanly (exit code 0, nothing logged). pg_cancel_backend() and pg_terminate_backend() have no effect. We kept the backend alive and debugged it with gdb + debug symbols; the evidence shows a leaked InterruptHoldoffCount, plus a design interaction that turns any such leak into a permanently unkillable backend. Possibly relevant configuration: io_method =3D io_uring, max_parallel_workers_per_gather =3D 48, shared_preload_libraries =3D 'timescaledb,pg_parquet' (timescaledb 2.27.2, pg_parquet 0.5.1). PostgreSQL 18.4 (Ubuntu 18.4-1.pgdg24.04+1). The query is a COPY (SELECT GROUP BY ...) TO STDOUT (FORMAT parquet) issued via pg_parquet's COPY hook; the inner select runs a parallel plan (Gather, 10+ workers planned) over TimescaleDB compressed-hypertable chunks, i.e. an AIO-read-heavy scan. Backtrace of the stuck leader (main thread; the only other thread is an io_uring kernel worker, "iou-wrk-"): #0 epoll_wait (epfd=3D754, events=3D..., maxevents=3D1, timeout=3D-1) #1 WaitEventSetWait () #2 WaitLatch () #3 WaitForParallelWorkersToFinish () #4 ExecParallelFinish () #5-#8 (Gather shutdown path, static functions) #9 standard_ExecutorRun () #10-#11 PortalRun () #12-#17 pg_parquet.so (COPY TO hook driving the portal) #18 prev_ProcessUtility (timescaledb process_utility.c:124) #19 timescaledb_ddl_command_start (...) #20-#22 PortalRun () #23-#24 PostgresMain () Key state read from the live process via gdb: InterruptPending =3D 1 ProcDiePending =3D 1 <- pg_terminate_backend arrived, never serviced QueryCancelPending =3D 1 <- pg_cancel_backend arrived, never serviced InterruptHoldoffCount =3D 1 <- held, forever CritSectionCount =3D 0 num_held_lwlocks =3D 0 <- NO LWLock is held ParallelMessagePending =3D 1 <- worker messages queued, unprocessable /proc signal state: SigPnd/ShdPnd are all zero - the SIGTERM/SIGINT were delivered and consumed by the handlers (which set the flags above); the backend wakes from WaitLatch, but CHECK_FOR_INTERRUPTS() is a no-op with InterruptHoldoffCount > 0, so it just re-sleeps. Why this is a permanent hang and not just an uncancellable wait: worker-exit accounting is itself interrupt-driven. The workers' detach notifications are only processed via CHECK_FOR_INTERRUPTS() -> ProcessParallelMessages(), which is what clears each worker's error_mqh. With the holdoff leaked, ProcessParallelMessages never runs (ParallelMessagePending stays 1 - verified in the live process), so WaitForParallelWorkersToFinish() believes the workers are still attached and waits forever. The same leaked holdoff blocks ProcDiePending. One leaked holdoff therefore produces both symptoms; only SIGKILL (crash-restart) or a debugger can clear the backend. On the origin of the leak: num_held_lwlocks =3D 0 rules out a stuck LWLockAcquire, and the server log shows NO error from this backend for its entire lifetime, which we believe rules out the "ereport(ERROR) raised under HOLD, caught by PG_TRY" pattern (the raise would have been logged by errfinish). That points to a straight-line code path that called HOLD_INTERRUPTS() and returned without RESUME_INTERRUPTS(). We audited the read path: bufmgr.c, read_stream.c, shm_mq.c, condition_variable.c contain no bare holds; timescaledb 2.27.2 and pg_parquet 0.5.1 (and pgrx 0.16.1) contain no HOLD_INTERRUPTS at all. The remaining bare holds on this workload's path are in src/backend/storage/aio/aio.c (four HOLD/RESUME regions, including pgaio_io_reclaim which runs completion callbacks under HOLD) and the parallel-message path itself. Given io_method =3D io_uring and an AIO-heavy scan we suspect an AIO path, but we could not identify the exact site from the live process. Frequency: once in ~6 days of heavy production use of this COPY path (thousands of executions/day). Not deterministically reproducible so far. We have a gcore of the stuck backend and can keep it (and/or the live process) available - happy to run additional gdb commands against it. Separately, it may be worth considering a defensive measure: because WaitForParallelWorkersToFinish depends on interrupt processing for its own termination condition, any holdoff leak from any subsystem converts into a silent, unkillable backend. Detecting "waiting on ParallelFinish with INTERRUPTS_CAN_BE_PROCESSED() false and ParallelMessagePending set" and at least logging (or treating it as an invariant violation) would make this failure mode visible instead of silent.