pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Nathan Bossart <nathandbossart@gmail.com>
To: Robert Haas <robertmhaas@gmail.com>
Cc: Michael Paquier <michael@paquier.xyz>
Cc: Tom Lane <tgl@sss.pgh.pa.us>
Cc: Andres Freund <andres@anarazel.de>
Cc: Thomas Munro <thomas.munro@gmail.com>
Cc: Fujii Masao <fujii@postgresql.org>
Cc: Postgres hackers <pgsql-hackers@lists.postgresql.org>
Subject: Re: Weird failure with latches in curculio on v15
Date: Thu, 2 Feb 2023 14:39:19 -0800
Message-ID: <20230202223919.GA3947443@nathanxps13> (raw)
In-Reply-To: <20230202220113.GA3945808@nathanxps13>
References: <1369666.1675264346@sss.pgh.pa.us>
	<20230201165801.33ydbxvjdbomjqa7@alap3.anarazel.de>
	<20230201175806.GA3199959@nathanxps13>
	<20230201223555.GA3721373@nathanxps13>
	<Y9sam108o4mxZFiS@paquier.xyz>
	<1449633.1675305284@sss.pgh.pa.us>
	<Y9s678gkiX0pEb5C@paquier.xyz>
	<20230202200957.GA3944544@nathanxps13>
	<CA+Tgmob+KZQn_EfVOp9umWc6iJmuzn5oH9hF_9+Cdu=wPgiEZg@mail.gmail.com>
	<20230202220113.GA3945808@nathanxps13>

On Thu, Feb 02, 2023 at 02:01:13PM -0800, Nathan Bossart wrote:
> I've been digging into the history here.  This e-mail seems to have the
> most context [0].  IIUC this was intended to prevent "fast" shutdowns from
> escalating to "immediate" shutdowns because the restore command died
> unexpectedly.  This doesn't apply to archive_cleanup_command because we
> don't FATAL if it dies unexpectedly.  It seems like this idea should apply
> to recovery_end_command, too, but AFAICT it doesn't use the same approach.
> My guess is that this hasn't come up because it's less likely that both 1)
> recovery_end_command is used and 2) someone initiates shutdown while it is
> running.

Actually, this still doesn't really explain why we need to exit immediately
in the SIGTERM handler for restore_command.  We already have handling for
when the command indicates it exited due to SIGTERM, so it should be no
problem if the command receives it before the startup process.  And
HandleStartupProcInterrupts() should exit at an appropriate time after the
startup process receives SIGTERM.

My guess was that this is meant to allow breaking out of the system() call,
but I don't understand why that's important here.  Maybe we could just
remove this exit-in-SIGTERM-handler business...

-- 
Nathan Bossart
Amazon Web Services: https://aws.amazon.com





view thread (78+ messages)  latest in thread

Message-ID: <20230202223919.GA3947443@nathanxps13>
Permalink:  ../20230202223919.GA3947443@nathanxps13/
Also on:    postgresql.org/message-id/20230202223919.GA3947443@nathanxps13

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: nathandbossart@gmail.com, robertmhaas@gmail.com, michael@paquier.xyz, tgl@sss.pgh.pa.us, andres@anarazel.de, thomas.munro@gmail.com, fujii@postgresql.org, pgsql-hackers@lists.postgresql.org
  Subject: Re: Weird failure with latches in curculio on v15
  In-Reply-To: <20230202223919.GA3947443@nathanxps13>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox