From: Nathan Bossart <nathandbossart@gmail.com>
To: Christoph Berg <myon@debian.org>
Cc: Laurenz Albe <laurenz.albe@cybertec.at>
Cc: Fujii Masao <masao.fujii@oss.nttdata.com>
Cc: Andres Freund <andres@anarazel.de>
Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>
Subject: Re: CHECKPOINT unlogged data
Date: Mon, 16 Jun 2025 10:18:41 -0500
Message-ID: <aFA10XSdJ9hPGfD1@nathan> (raw)
In-Reply-To: <aFAsCyMj-mybBlv3@msg.df7cb.de>
References: <fc1ed66f-e95a-4d10-a4f7-7fa4bb9a7084@oss.nttdata.com>
<aEL6sDr9RxdSkPm-@msg.df7cb.de>
<aEMIrLUDOKgmE_0P@nathan>
<aEMVRbmqqg-aaxAN@msg.df7cb.de>
<aEMYYEkQb2VeQoBo@nathan>
<aEmIijETyhj9DsWG@msg.df7cb.de>
<aEmLzinvZ-W8EPiW@nathan>
<aEmma1MPm_UgmLP3@msg.df7cb.de>
<086ab9ad5a2c06eb16f6f4e50a04764ba6fafd1b.camel@cybertec.at>
<aFAsCyMj-mybBlv3@msg.df7cb.de>
On Mon, Jun 16, 2025 at 04:36:59PM +0200, Christoph Berg wrote:
> I spent some time digging through the code, but I'm still not entirely
> sure what's happening. There are several parts to it:
>
> 1) the list of buffers to flush is determined at the beginning of the
> checkpoint, so running a 2nd FLUSH_UNLOGGED checkpoint will not make
> the running checkpoint write these
>
> 2) running CHECKPOINT updates the checkpoint flags in shared memory so
> I think the currently running checkpoint picks "MODE FAST" up and
> speeds up. (But I'm not entirely sure, the call stack is quite deep
> there.)
>
> 3) running CHECKPOINT (at least when waiting for it) seems to actually
> start a new checkpoint, so FLUSH_UNLOGGED should still be effective.
> (See the code arount "start_cv" in checkpointer.c)
>
> Admittedly, adding these points together raises some question marks
> about the flag handling, so I would welcome clarification by someone
> more knowledgeable in this area.
I think you've got it right. With CHECKPOINT_WAIT set, RequestCheckpoint()
will wait for a new checkpoint to start, at which point we know that the
new flags have been seen by the checkpointer. If an immediate checkpoint
is pending, CheckpointWriteDelay() will skip sleeping in the
currently-running one, so the current checkpoint will be "upgraded" to
immediate in some sense, but IIUC there will still be another immediate
checkpoint after it completes. But AFAICT it doesn't pick up
FLUSH_UNLOGGED until the next checkpoint begins.
Another thing to note is what I mentioned earlier:
+ Note that the server may consolidate concurrently requested checkpoints or
+ restartpoints. Such consolidated requests will contain a combined set of
+ options. For example, if one session requested an immediate checkpoint and
+ another session requested a non-immediate checkpoint, the server may combine
+ these requests and perform one immediate checkpoint.
--
nathan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: nathandbossart@gmail.com, myon@debian.org, laurenz.albe@cybertec.at, masao.fujii@oss.nttdata.com, andres@anarazel.de, pgsql-hackers@lists.postgresql.org
Subject: Re: CHECKPOINT unlogged data
In-Reply-To: <aFA10XSdJ9hPGfD1@nathan>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox