Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzb1I-0046i7-2V for pgsql-bugs@arkaria.postgresql.org; Thu, 27 Aug 2026 14:23:56 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1wzb1H-002RlX-2T for pgsql-bugs@arkaria.postgresql.org; Thu, 27 Aug 2026 14:23:55 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzCFh-00Cgwk-3B for pgsql-bugs@lists.postgresql.org; Wed, 26 Aug 2026 11:57:10 +0000 Received: from mahout.postgresql.org ([2001:4800:3e1:1::227]) by magus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1wzCFe-00000001MPa-47Eu for pgsql-bugs@lists.postgresql.org; Wed, 26 Aug 2026 11:57:09 +0000 DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=postgresql.org; s=20171124; h=Message-ID:Date:Reply-To:Cc:From:To:Subject: Content-Transfer-Encoding:MIME-Version:Content-Type:Sender:Content-ID: Content-Description:In-Reply-To:References; bh=bEIdqhwaottvqIbI6OeY21ZIWqx0d+txwtSXreeUnQU=; b=l2XRATwzNUjZcMK73B8BbQnrlg u61oblJlBBsV2ZdjucX2NJHEvWOWWO9XvyDQupxkF/ZhXjpYtwqBpJOdkjvaqVsHdufsHq//mjh7E ALH9VrkubQ1F8feqYVQJgBr2KwUufi0ym2lby853dtuFX0uUm+yKQ02ZKbzafnRQvxaYMU1zgZ7EX mlcdCfjKY4TMIL5zf59Z1+eaQ1fuq42K9aBTUr7SPg3aAfuqb/tR98a1eaw0KkWrfIzt6/tF1mvbH Xdf94BQdbIPdKOlvVXe6e1j9zkkuoBvXnWp/x15iP2rkY8ZUfTtMuCAQ+qNGJnbgShvrC45tqKctC TGG89kbQ==; Received: from wrigleys.postgresql.org ([2a02:16a8:dc51::60]) by mahout.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzCFd-007d9E-1q for pgsql-bugs@lists.postgresql.org; Wed, 26 Aug 2026 11:57:05 +0000 Received: from localhost ([127.0.0.1] helo=wrigleys.postgresql.org) by wrigleys.postgresql.org with esmtp (Exim 4.98.2) (envelope-from ) id 1wzCFb-00000002sXd-3V8X for pgsql-bugs@lists.postgresql.org; Wed, 26 Aug 2026 11:57:03 +0000 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Subject: BUG #19640: Standby permanently stuck re-requesting old timeline after promotion, never switches to new timeline To: pgsql-bugs@lists.postgresql.org From: PG Bug reporting form Cc: harshit.singh817775@gmail.com Reply-To: harshit.singh817775@gmail.com, pgsql-bugs@lists.postgresql.org Date: Wed, 26 Aug 2026 11:56:11 +0000 Message-ID: <19640-3003103a974cd7dc@postgresql.org> X-Auto-Response-Suppress: All Auto-Submitted: auto-generated List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk The following bug has been logged on the website: Bug reference: 19640 Logged by: Harshit Singh Email address: harshit.singh817775@gmail.com PostgreSQL version: 18.0 Operating system: rhel 10 Description: =20 PostgreSQL version: Reproduced on both 17.6 and 18.0 =E2=80=94 present in t= he current release, not something already fixed upstream. Operating system: Linux x86_64 (RHEL-family, built with Red Hat gcc 14.3.1) Description: A standby configured with recovery_target_timeline =3D 'latest' (the default under streaming-replication HA managers like Patroni) can get permanently stuck after a timeline promotion elsewhere in the cluster: it endlessly re-requests WAL on its own old, now-superseded timeline, is told "end of WAL reached" by the primary each time, disconnects, and immediately reconnects requesting the same old timeline again =E2=80=94 never advancing to the new= one. This repeats indefinitely (observed for over an hour in one case, tight ~10-20ms reconnect loop), consuming CPU, with no error surfaced to indicate the process needs manual intervention =E2=80=94 patronictl/monitoring tooli= ng on top of it just reports the node as "starting" forever. Steps to reproduce: 1. 3-node streaming replication cluster (repro used Patroni-managed Postgres, but the core issue appears to be in core recovery logic, not Patroni). 2. Node A is primary; other nodes stream from it. 3. Node A is stopped/killed (or even just a plain systemctl restart of the current leader =E2=80=94 no exotic failure needed). Another node is promoted (timeline N =E2=86=92 N+1). 4. Node A is later restarted and attempts to rejoin as a replica of the new leader. 5. Node A's local timeline is N; the new leader is on N+1. 6. Node A's log shows, repeating forever: LOG: started streaming WAL from primary at on timeline N LOG: replication terminated by primary server DETAIL: End of WAL reached on timeline N at . FATAL: terminating walreceiver process due to administrator command LOG: waiting for WAL to become available at =E2=80=94 new walreceiver PID each cycle, always requesting timeline N, = never N+1. Reproduced 5 times across different sessions/timelines (N=3D1 through N=3D4= ) and both PostgreSQL 17.6 and 18.0, including once under active write load (real WAL divergence existed, not just an idle-DB edge case), and via completely ordinary systemctl stop/start of the current leader =E2=80=94 not an exotic scenario. The fact that it reproduces identically on 18.0 indicates this isn't a regression already fixed in the latest release. Expected behavior: The standby should detect, via rescanLatestTimeLine(), that a newer timeline (N+1) now exists and switch its target to follow it, per recovery_target_timeline =3D 'latest' semantics. Suspected root cause / prior art: This looks closely related to the mechanism described by Dilip Kumar on -hackers in "Race condition in recovery?" (https://postgrespro.com/list/thread-id/2526828): WaitForWALToBecomeAvailable() initializes expectedTLEs from receiveTLI rather than recoveryTargetTLI. When rescanLatestTimeLine() finds the newest TLE already matches recoveryTargetTLI, it concludes "nothing to change" =E2= =80=94 but expectedTLEs is left referencing the old timeline regardless, so every subsequent WAL request keeps using it. A patch was proposed there (initializing from recoveryTargetTLI instead) but the thread doesn't show it as committed, and given this is still reproducible on 18.0, it appears that patch =E2=80=94 or an equivalent fix =E2=80=94 never landed. Additional context: A near-identical symptom was reported independently against CloudNativePG (https://github.com/cloudnative-pg/cloudnative-pg/issues/10419), consistent with this being a core recovery-logic issue rather than something specific to any one HA orchestration layer.