agora inbox for [email protected]
help / color / mirror / Atom feedFrom: Dmitrii Dolgov <[email protected]>
Subject: [PATCH v5 07/10] Allow to resize shared memory without restart
Date: Tue, 17 Jun 2025 14:16:55 +0200
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers, adjusts its size using ftruncate and adjust reservation
permissions with mprotect. One elected process signals the postmaster
to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f87909fc000-7f8798248000 rw-s /memfd:strategy (deleted)
7f8798248000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a4e84000 rw-s /memfd:checkpoint (deleted)
7f87a4e84000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1b42000 rw-s /memfd:iocv (deleted)
7f87b1b42000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cb59c000 rw-s /memfd:descriptors (deleted)
7f87cb59c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f87ece38000 rw-s /memfd:buffers (deleted)
^ buffers content, ~247 MB
7f87ece38000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~2210 MB
7f8877066000-7f887e7d0000 rw-s /memfd:main (deleted)
7f887e7d0000-7f8890a00000 ---s /memfd:main (deleted)
-- 512 MB
7f87909fc000-7f879866a000 rw-s /memfd:strategy (deleted)
7f879866a000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a50f4000 rw-s /memfd:checkpoint (deleted)
7f87a50f4000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1d82000 rw-s /memfd:iocv (deleted)
7f87b1d82000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cba1c000 rw-s /memfd:descriptors (deleted)
7f87cba1c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f8804fb8000 rw-s /memfd:buffers (deleted)
^ buffers content, ~632 MB
7f8804fb8000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~1824 MB
7f8877066000-7f887e950000 rw-s /memfd:main (deleted)
7f887e950000-7f8890a00000 ---s /memfd:main (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
view thread (301+ messages) latest in thread
reply
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: [email protected]
Cc: [email protected]
Subject: Re: [PATCH v5 07/10] Allow to resize shared memory without restart
In-Reply-To: <no-message-id-1009043@localhost>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox