Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1vxh9P-00H2tZ-1b for pgsql-hackers@arkaria.postgresql.org; Wed, 04 Mar 2026 08:00:11 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1vxh9M-00BKgT-2j for pgsql-hackers@arkaria.postgresql.org; Wed, 04 Mar 2026 08:00:09 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1vxh9M-00BKgL-1P for pgsql-hackers@lists.postgresql.org; Wed, 04 Mar 2026 08:00:08 +0000 Received: from mail-wm1-x334.google.com ([2a00:1450:4864:20::334]) by makus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.98.2) (envelope-from ) id 1vxh9K-00000000KzF-0WeB for pgsql-hackers@postgresql.org; Wed, 04 Mar 2026 08:00:07 +0000 Received: by mail-wm1-x334.google.com with SMTP id 5b1f17b1804b1-48371bb515eso98034345e9.1 for ; Wed, 04 Mar 2026 00:00:06 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1772611203; x=1773216003; darn=postgresql.org; h=in-reply-to:from:content-language:references:cc:to:subject :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=wlZngAg1tCcBu59s6ZjyHn8+dGjXgB10dpCtUw1t91I=; b=JfAvv/eGzqwawjDANi6JuI0YudhnZ8P0VpWU2oUG+b/hIX/53h3caPAhOB7PJ5FzIe EuSaWxB9FWI56Hp1Cez6qkdF+97KVhBBwG40c3q57Z6GnCSIUMmlAo2erzicA882fPzQ hR9SRNzwwov2LutX6yuJKhCSohbrIaEm1yFnm890E0v4Tv9KS6/u3n12Tf/zCje5GOrr ixAso7gGGcaITC3hMc8qNQx8JNtocYlCq/8Kq69PEEkO6d4tqJPOLCDoB0NqqUOeMt+L BiPmLuYs9A11ygiCC5+/C/uyy8IpbAmMMaDtGk3k0iu6d35xUPo49tXr25HIzNxZZCEA gdXg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1772611203; x=1773216003; h=in-reply-to:from:content-language:references:cc:to:subject :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=wlZngAg1tCcBu59s6ZjyHn8+dGjXgB10dpCtUw1t91I=; b=PiKT0nswdu7h226a25M01m4MZlN37ROFVLNi6Yab91INDSGOz9RPEr3nFPHUMAYH6q 5ZNDe7qfQFyR+cp2NK+71ZfzJezV0FhqT9VwieodycxBxELV1EIrtSdJstpRbHgJiFlC GODyNBdmmdZpKWW8aiAhVlpOcXDdSLvWjLTHqTukLcM0alTOXi13KIJkZO9HUcR1bgwl JbRO4VLdPYx2m9cuzDCxgg4RRxvJdLQWjpeUM2TSPsQZCL3DG5lSMuh0Zu6aOdZxV8SN TVLEejV60A6m7+LYWSNAO8dXp6PqJ+HCT87h3MlOaxlwaUqqlto55Q3q9ZTFRVj/N3op V+Fw== X-Gm-Message-State: AOJu0YxxoJUoYVxzsqX8sbYSye9t3CVwpV6UEranKCMjvii1epfGjw+W BX1PsNo63Q+spM3LNDi63TxSulu6SBYrsl71zRIQoI8uKNhZbLLO8wyg X-Gm-Gg: ATEYQzzT4PWusgvWgFpkJSpG8XNz3Zr1OuH5scOhN69axzzJXWGhDJueZk9O2oZJDFF riq6g8vzj0Us6UytIUZxLa0/fRWUq+3Bd78FmDPdKljkUqTRaqQLmCdsiEzZozrYpE1SuX7FA9x o2nWcY/ICiZCoaFAUQJjjVcTHGaQ26UBcxQzhWZPmqSgBW1ooDdwGn4UsR5Deov49PivLpRMLjQ pc6e+LeBlTNLh0sStkeYMXVFoJ3W3mvPyXN0N178QEQex5qALGl5EN75U9ElW62DhpZtR2QvxnW LNXPFgQDBU+UiVxpi9YjMxwQH1DMdZzo/8QIsiEIeIeo2O6pRo8D2nBWzn5klYVGkoY1EGiQLHb v1AGwiSyRacX21WW4/fq3G0ftawrT0mI5fVnssdgsmo+EuocGFUR5mntiIFoyip45oO7NlZi+Lc cf3SnobUIkeTmIj5yNottxMzXf X-Received: by 2002:a05:600c:8b53:b0:483:498f:7963 with SMTP id 5b1f17b1804b1-4851989024emr15522885e9.26.1772611202754; Wed, 04 Mar 2026 00:00:02 -0800 (PST) Received: from [192.168.0.50] ([89.149.93.164]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4851ad0ad2fsm4270265e9.5.2026.03.04.00.00.01 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Wed, 04 Mar 2026 00:00:02 -0800 (PST) Content-Type: multipart/alternative; boundary="------------1mXXB5OE4ZMBBfPuHgGe00Xs" Message-ID: <9d4abbe2-95aa-47e0-9ce2-842196a662a7@gmail.com> Date: Wed, 4 Mar 2026 10:00:00 +0200 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: BUG: Former primary node might stuck when started as a standby To: Michael Paquier , "Hayato Kuroda (Fujitsu)" Cc: pgsql-hackers , Aleksander Alekseev References: <63e55743-669e-4300-a561-7b7ff63723b6@gmail.com> <045cab6f-4738-417e-b551-01adba44d6c3@gmail.com> <493401a8-063f-436a-8287-a235d9e065fc@gmail.com> Content-Language: en-US From: Alexander Lakhin In-Reply-To: List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk This is a multi-part message in MIME format. --------------1mXXB5OE4ZMBBfPuHgGe00Xs Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit Hello Michael, 04.03.2026 07:31, Michael Paquier wrote >> I guess so. cluster::stop does the `pg_ctl stop -m fast` command. In this case >> the walsender waits till there are nothing to be sent, see WalSndLoop(). >> Do let me know if you have observed the similar failure here. > Exactly. Doing a clean stop of the primary offers a strong guarantee > here. We are sure that the standby will have received all the records > from the primary. Timeline forking is an impossible thing in > 012_subtransactions.pl based on how the switchover from the primary to > the standby happens. I don't see a need for tweaking this test at > all. Or perhaps you did see a failure of some kind in this test, > Alexander? Yes, 012_subtransactions doesn't fail with aggressive bgwriter, as I noted before. I mentioned it exactly to show that stop does matter here. But if we recognize teardown_node in this context as risky, maybe it would make sense to review also other tests in recovery/. I already wrote about 004_timeline_switch, but probably there are more. E.g., 028_pitr_timelines (I haven't tested it intensively yet) does: $node_primary->stop('immediate'); # Promote the standby, and switch WAL so that it archives a WAL segment # that contains all the INSERTs, on a new timeline. $node_standby->promote; Best regards, Alexander --------------1mXXB5OE4ZMBBfPuHgGe00Xs Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: 7bit
Hello Michael,

04.03.2026 07:31, Michael Paquier wrote
I guess so. cluster::stop does the `pg_ctl stop -m fast` command. In this case
the walsender waits till there are nothing to be sent, see WalSndLoop().
Do let me know if you have observed the similar failure here.
Exactly.  Doing a clean stop of the primary offers a strong guarantee
here.  We are sure that the standby will have received all the records
from the primary.  Timeline forking is an impossible thing in
012_subtransactions.pl based on how the switchover from the primary to
the standby happens.  I don't see a need for tweaking this test at
all.  Or perhaps you did see a failure of some kind in this test,
Alexander?


Yes, 012_subtransactions doesn't fail with aggressive bgwriter, as I noted
before. I mentioned it exactly to show that stop does matter here. But if
we recognize teardown_node in this context as risky, maybe it would make
sense to review also other tests in recovery/. I already wrote about
004_timeline_switch, but probably there are more. E.g., 028_pitr_timelines
(I haven't tested it intensively yet) does:
$node_primary->stop('immediate');

# Promote the standby, and switch WAL so that it archives a WAL segment
# that contains all the INSERTs, on a new timeline.
$node_standby->promote;

Best regards,
Alexander
--------------1mXXB5OE4ZMBBfPuHgGe00Xs--