agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedChanging shared_buffers without restart
167+ messages / 22 participants
[nested] [flat]
* Changing shared_buffers without restart
@ 2024-10-18 19:21 Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-01 15:27 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-07 01:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2024-11-19 12:57 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 5 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2024-10-18 19:21 UTC (permalink / raw)
To: pgsql-hackers
TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
changing shared memory mapping layout. Any feedback is appreciated.
Hi,
Being able to change PostgreSQL configuration on the fly is an important
property for performance tuning, since it reduces the feedback time and
invasiveness of the process. In certain cases it even becomes highly desired,
e.g. when doing automatic tuning. But there are couple of important
configuration options that could not be modified without a restart, the most
notorious example is shared_buffers.
I've been working recently on an idea how to change that, allowing to modify
shared_buffers without a restart. To demonstrate the approach, I've prepared a
PoC that ignores lots of stuff, but works in a limited set of use cases I was
testing. I would like to discuss the idea and get some feedback.
Patches 1-3 prepare the infrastructure and shared memory layout. They could be
useful even with multithreaded PostgreSQL, when there will be no need for
shared memory. I assume, in the multithreaded world there still will be need
for a contiguous chunk of memory to share between threads, and its layout would
be similar to the one with shared memory mappings.
Patch 4 actually does resizing. It's shared memory specific of course, and
utilized Linux specific mremap, meaning open portability questions.
Patch 5 is somewhat independent, but quite convenient to have. It also utilizes
Linux specific call memfd_create.
The patch set still doesn't address lots of things, e.g. shared memory segment
detach/reattach, portability questions, it doesn't touch EXEC_BACKEND code and
huge pages.
So far I was doing some rudimentary testing: spinning up PostgreSQL, then
increasing shared_buffers and running pgbench with the scale factor large
enough to extend the data set into newly allocated buffers:
-- shared_buffers 128 MB
=# SELECT * FROM pg_buffercache_summary();
buffers_used | buffers_unused | buffers_dirty | buffers_pinned
--------------+----------------+---------------+----------------
134 | 16250 | 1 | 0
-- change shared_buffers to 512 MB
=# select pg_reload_conf();
=# SELECT * FROM pg_buffercache_summary();
buffers_used | buffers_unused | buffers_dirty | buffers_pinned
--------------+----------------+---------------+---------------
221 | 65315 | 1 | 0
-- round of pgbench read-only load
=# SELECT * FROM pg_buffercache_summary();
buffers_used | buffers_unused | buffers_dirty | buffers_pinned
--------------+----------------+---------------+---------------
41757 | 23779 | 216 | 0
Here is the breakdown:
v1-0001-Allow-to-use-multiple-shared-memory-mappings.patch
Preparation, introduces the possibility to work with many shmem mappings. To
make it less invasive, I've duplicated the shmem API to extend it with the
shmem_slot argument, while redirecting the original API to it. There are
probably better ways of doing that, I'm open for suggestions.
v1-0002-Allow-placing-shared-memory-mapping-with-an-offse.patch
Implements a new layout of shared memory mappings to include room for resizing.
I've done a couple of tests to verify that such space in between doesn't affect
how the kernel calculates actual used memory, to make sure that e.g. cgroup
will not trigger OOM. The only change seems to be in VmPeak, which is total
mapped pages.
v1-0003-Introduce-multiple-shmem-slots-for-shared-buffers.patch
Splits shared_buffers into multiple slots, moving out structures that depend on
NBuffers into separate mappings. There are two large gaps here:
* Shmem size calculation for those mappings is not correct yet, it includes too
many other things (no particular issues here, just haven't had time).
* It makes hardcoded assumptions about what is the upper limit for resizing,
which is currently low purely for experiments. Ideally there should be a new
configuration option to specify the total available memory, which would be a
base for subsequent calculations.
v1-0004-Allow-to-resize-shared-memory-without-restart.patch
Do shared_buffers change without a restart. Current approach is clumsy, it adds
an assign hook for shared_buffers and goes from there using mremap to resize
mappings. But I haven't immediately found any better approach. Currently it
supports only an increase of shared_buffers.
v1-0005-Use-anonymous-files-to-back-shared-memory-segment.patch
Allows an anonyous file to back a shared mapping. This makes certain things
easier, e.g. mappings visual representation, and gives an fd for possible
future customizations.
In this thread I'm hoping to answer following questions:
* Are there any concerns about this approach?
* What would be a better mechanism to handle resizing than an assign hook?
* Assuming I'll be able to address already known missing bits, what are the
chances the patch series could be accepted?
From 954613a63cb1102d7eb88f92e7ff561828bbb5c9 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 9 Oct 2024 15:41:32 +0200
Subject: [PATCH v1 1/5] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory slot.
There is only fixed amount of available slots, currently only one main
shared memory slot is allocated. A new shared memory API is introduces,
extended with a slot as a new parameter. As a path of least resistance,
the original API is kept in place, utilizing the main shared memory slot.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 +++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 2 +-
src/backend/storage/ipc/ipci.c | 61 ++++++------
src/backend/storage/ipc/shmem.c | 133 ++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 5 +-
src/include/storage/buf_internals.h | 1 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 10 ++
13 files changed, 258 insertions(+), 124 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 64186ec0a7..b97723d2ed 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSlot(PGSemaphoreShmemSize(maxSemas), shmem_slot);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 5b88a92bc9..8ef95b12c9 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -307,7 +307,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
struct stat statbuf;
@@ -328,7 +328,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSlot(PGSemaphoreShmemSize(maxSemas), shmem_slot);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 362a37d3b3..065a5b63ac 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_slot;
+ Size shmem_size; /* Size of the mapping */
+ void *shmem; /* Pointer to the start of the mapped memory */
+ void *seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping slots */
+static int next_free_slot = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_slot)
+{
+ switch (shmem_slot)
+ {
+ case MAIN_SHMEM_SLOT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_slot; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_slot), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("slot[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_slot)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_slot; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_slot];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_slot = next_free_slot;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_slot++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_slot;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_slot; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index f2b54bdfda..d62084cc0d 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index b06e4b8452..2aabd4a77f 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -68,7 +68,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 35fa2e1dda..8224015b53 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -88,7 +88,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_slot)
{
Size size;
int numSemas;
@@ -202,33 +202,36 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int slot = 0; slot < ANON_MAPPINGS; slot++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, slot);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSlot(seghdr, slot);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, slot);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSlot(slot);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -359,7 +362,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SLOT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 6d5f083986..c670b9cf43 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -75,17 +75,12 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSlot(Size size, Size *allocated_size,
+ int shmem_slot);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
-
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
+ShmemSegment Segments[ANON_MAPPINGS];
static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
@@ -99,11 +94,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
void
InitShmemAccess(void *seghdr)
{
- PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ InitShmemAccessInSlot(seghdr, MAIN_SHMEM_SLOT);
+}
- ShmemSegHdr = shmhdr;
- ShmemBase = (void *) shmhdr;
- ShmemEnd = (char *) ShmemBase + shmhdr->totalsize;
+void
+InitShmemAccessInSlot(void *seghdr, int shmem_slot)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_slot];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -114,7 +115,13 @@ InitShmemAccess(void *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSlot(MAIN_SHMEM_SLOT);
+}
+
+void
+InitShmemAllocationInSlot(int shmem_slot)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_slot].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -123,9 +130,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_slot].ShmemLock = (slock_t *) ShmemAllocUnlockedInSlot(sizeof(slock_t), shmem_slot);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_slot].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -150,11 +157,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSlot(size, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemAllocInSlot(Size size, int shmem_slot)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSlot(size, &allocated_size, shmem_slot);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -184,6 +197,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSlot(size, allocated_size, MAIN_SHMEM_SLOT);
+}
+
+static void *
+ShmemAllocRawInSlot(Size size, Size *allocated_size, int shmem_slot)
{
Size newStart;
Size newFree;
@@ -203,22 +222,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_slot].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_slot].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_slot].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_slot].ShmemSegHdr->totalsize)
{
- newSpace = (void *) ((char *) ShmemBase + newStart);
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (void *) ((char *) Segments[shmem_slot].ShmemBase + newStart);
+ Segments[shmem_slot].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_slot].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -236,6 +255,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSlot(size, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemAllocUnlockedInSlot(Size size, int shmem_slot)
{
Size newStart;
Size newFree;
@@ -246,19 +271,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_slot].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_slot].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_slot].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_slot].ShmemSegHdr->freeoffset = newFree;
- newSpace = (void *) ((char *) ShmemBase + newStart);
+ newSpace = (void *) ((char *) Segments[shmem_slot].ShmemBase + newStart);
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -273,7 +298,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSlot(addr, MAIN_SHMEM_SLOT);
+}
+
+bool
+ShmemAddrIsValidInSlot(const void *addr, int shmem_slot)
+{
+ return (addr >= Segments[shmem_slot].ShmemBase) && (addr < Segments[shmem_slot].ShmemEnd);
}
/*
@@ -334,6 +365,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSlot(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SLOT);
+}
+
+HTAB *
+ShmemInitHashInSlot(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_slot) /* in which slot to keep the table */
{
bool found;
void *location;
@@ -350,9 +393,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSlot(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_slot);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -385,6 +428,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSlot(name, size, foundPtr, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
+ int shmem_slot)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -393,7 +443,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_slot].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -416,7 +466,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSlot(size, shmem_slot);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -433,8 +483,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in slot %d",
+ name, shmem_slot)));
}
if (*foundPtr)
@@ -459,7 +509,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSlot(size, &allocated_size, shmem_slot);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -478,14 +528,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSlot(structPtr, shmem_slot));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -545,7 +594,7 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SLOT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -557,15 +606,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index e765754d80..fb0c33bf17 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -81,6 +81,7 @@
#include "pgstat.h"
#include "port/pg_bitutils.h"
#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/spin.h"
@@ -607,9 +608,9 @@ LWLockNewTrancheId(void)
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
/* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SLOT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SLOT].ShmemLock);
return result;
}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index f190e6e5e4..aef80e049b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -23,6 +23,7 @@
#include "storage/latch.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
+#include "storage/pg_shmem.h"
#include "storage/smgr.h"
#include "storage/spin.h"
#include "utils/relcache.h"
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index b2d062781e..be4b131288 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_slot);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index dfef79ac96..081fffaf16 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_slot);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 3065ff5be7..e968deeef7 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+// Number of available slots for anonymous memory mappings
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main slot, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SLOT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 842989111c..d3e9cc721d 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -28,15 +28,25 @@
/* shmem.c */
extern PGDLLIMPORT slock_t *ShmemLock;
extern void InitShmemAccess(void *seghdr);
+extern void InitShmemAccessInSlot(void *seghdr, int shmem_slot);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSlot(int shmem_slot);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSlot(Size size, int shmem_slot);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSlot(Size size, int shmem_slot);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSlot(const void *addr, int shmem_slot);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSlot(const char *name, long init_size, long max_size,
+ HASHCTL *infoP, int hash_flags, int shmem_slot);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
+ int shmem_slot);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: 2488058dc356a43455b21a099ea879fff9266634
--
2.45.1
From e9980f76cbd1ea6f6d732e2a27dd1342258d26e5 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH v1 2/5] Allow placing shared memory mapping with an offset
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is first to get
the address chosen by the kernel, then apply some offset derived from
the expected upper limit. Because we base the layout on the address
chosen by the kernel, things like address space randomization should not
be a problem, since the randomization is applied to the mmap base, which
is one per process. The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
[...free space...]
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
This approach do not impact the actual memory usage as reported by the kernel.
Here is the output of /proc/$PID/status for the master version with
shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages mapped in mm_struct
VmPeak: 422780 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 21248 kB
// Size of resident anonymous memory
RssAnon: 640 kB
// Size of resident file mappings
RssFile: 9728 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 10880 kB
Here is the same for the patch with the shared mapping placed at
an offset 10 GB:
VmPeak: 1102844 kB
VmRSS: 21376 kB
RssAnon: 640 kB
RssFile: 9856 kB
RssShmem: 10880 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
18219008 (~17 MB)
Note that currently the implementation makes assumptions about the upper limit.
Ideally it should be based on the maximum available memory.
---
src/backend/port/sysv_shmem.c | 120 +++++++++++++++++++++++++++++++++-
1 file changed, 119 insertions(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 065a5b63ac..7e6c8bb78d 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,63 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping slots */
static int next_free_slot = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * By letting Linux to chose a mapping address, it will pick up the lowest
+ * possible address that do not clash with any other mappings, which will be
+ * right before locales in the example above. This information (maximum allowed
+ * size of mappings and the lowest mapping address) is enough to place every
+ * mapping as follow:
+ *
+ * - Take the lowest mapping address, which we call later the probe address.
+ * - Substract the offset of the previous mapping.
+ * - Substract the maximum allowed size for the current mapping from the
+ * address.
+ * - Place the mapping by the resulting address.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ */
+Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+ 0, /* MAIN_SHMEM_SLOT */
+};
+
+/* Remembers offset of the last mapping from the probe address */
+static Size last_offset = 0;
+
+/*
+ * Size of the mapping, which will be used to calculate anonymous mapping
+ * address. It should not be too small, otherwise there is a chance the probe
+ * mapping will be created between other mappings, leaving no room extending
+ * it. But it should not be too large either, in case if there are limitations
+ * on the mapping size. Current value is the default shared_buffers.
+ */
+#define PROBE_MAPPING_SIZE (Size) 128 * 1024 * 1024
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -673,13 +730,74 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
+ void *probe = NULL;
+
/*
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+
+ /*
+ * Try to create mapping at an address, which will allow to extend it
+ * later:
+ *
+ * - First create the temporary probe mapping of a fixed size and let
+ * kernel to place it at address of its choice. By the virtue of the
+ * probe mapping size we expect it to be located at the lowest
+ * possible address, expecting some non mapped space above.
+ *
+ * - Unmap the probe mapping, remember the address.
+ *
+ * - Create an actual anonymous mapping at that address with the
+ * offset. The offset is calculated in such a way to allow growing
+ * the mapping withing certain boundaries. For this mapping we use
+ * MAP_FIXED_NOREPLACE, which will error out with EEXIST if there is
+ * any mapping clash.
+ *
+ * - If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized
+ * without a restart.
+ */
+ probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
+
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: probe mmap(%zu) failed: %m",
+ MappingName(mapping->shmem_slot), allocsize);
+ }
+ else
+ {
+ Size offset = last_offset + SHMEM_EXTRA_SIZE_LIMIT[next_free_slot] + allocsize;
+ last_offset = offset;
+
+ munmap(probe, PROBE_MAPPING_SIZE);
+
+ ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ mmap_errno = errno;
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: mmap(%zu) at address %p failed: %m",
+ MappingName(mapping->shmem_slot), allocsize, probe - offset);
+ }
+
+ }
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ /*
+ * Fallback to the portable way of creating a mapping.
+ */
+ allocsize = mapping->shmem_size;
+
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
}
--
2.45.1
From 62ae567f1c7a56c32722508b60251f9cec245ea3 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:24:04 +0200
Subject: [PATCH v1 3/5] Introduce multiple shmem slots for shared buffers
Add more shmem slots to split shared buffers into following chunks:
* BUFFERS_SHMEM_SLOT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SLOT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SLOT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SLOT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SLOT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers, meaning
that if we would like to change NBuffers, they have to be resized
correspondingly. Placing each of them in a separate shmem slot allows to
achieve that.
There are some asumptions made about each of shmem slots upper size limit. The
buffer blocks have the largest, while the rest claim less extra room for
resize. Ideally those limits have to be deduced from the maximum allowed shared
memory.
---
src/backend/port/sysv_shmem.c | 17 +++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 5 +-
src/backend/storage/buffer/freelist.c | 4 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 23 +++++++-
7 files changed, 97 insertions(+), 35 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 7e6c8bb78d..beebd4d85e 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -149,8 +149,13 @@ static int next_free_slot = 0;
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
*/
-Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+Size SHMEM_EXTRA_SIZE_LIMIT[6] = {
0, /* MAIN_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 1024 * 10, /* BUFFERS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 1024 * 1, /* BUFFER_DESCRIPTORS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* BUFFER_IOCV_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* CHECKPOINT_BUFFERS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* STRATEGY_SHMEM_SLOT */
};
/* Remembers offset of the last mapping from the probe address */
@@ -179,6 +184,16 @@ MappingName(int shmem_slot)
{
case MAIN_SHMEM_SLOT:
return "main";
+ case BUFFERS_SHMEM_SLOT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SLOT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SLOT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SLOT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SLOT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 46116a1f64..6bca286bef 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory slot, which
+ * could be resized on demand.
*/
void
InitBufferPool(void)
@@ -74,22 +77,22 @@ InitBufferPool(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSlot("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SLOT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSlot("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SLOT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSlot("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SLOT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ InitBufferPool(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSlot("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SLOT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -154,33 +158,54 @@ InitBufferPool(void)
* BufferShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory slot. The main slot must not allocate anything
+ * related to buffers, every other slot will receive part of the
+ * data.
*/
Size
-BufferShmemSize(void)
+BufferShmemSize(int shmem_slot)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_slot == MAIN_SHMEM_SLOT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_slot == BUFFER_DESCRIPTORS_SHMEM_SLOT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_slot == BUFFERS_SHMEM_SLOT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_slot == STRATEGY_SHMEM_SLOT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_slot == BUFFER_IOCV_SHMEM_SLOT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_slot == CHECKPOINT_BUFFERS_SHMEM_SLOT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 0fa5468930..ccbaed8010 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -59,10 +59,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSlot("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SLOT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 19797de31a..8ce1611db2 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -491,9 +491,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSlot("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SLOT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8224015b53..fbaddba396 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -115,7 +115,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_slot)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferShmemSize());
+ size = add_size(size, BufferShmemSize(shmem_slot));
size = add_size(size, LockShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index c8422571b7..4c09d270c9 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -301,7 +301,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void InitBufferPool(void);
-extern Size BufferShmemSize(void);
+extern Size BufferShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index e968deeef7..c0143e3899 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
// Number of available slots for anonymous memory mappings
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -105,7 +105,28 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings slots. Each slot
+ * contains only certain part of the data, which size depends on NBuffers.
+ */
+
/* The main slot, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SLOT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SLOT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SLOT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SLOT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SLOT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SLOT 5
+
#endif /* PG_SHMEM_H */
--
2.45.1
From 7183999bba1cbeebd059d18e5a590cbef7aff2d1 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:24:58 +0200
Subject: [PATCH v1 4/5] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Size for every shared memory slot is recalculated based on the new
NBuffers, and extended using mremap. After allocating new space, new
shared structures (buffer blocks, descriptors, etc) are allocated as
needed. Here is how it looks like after raising shared_buffers from 128
MB to 512 MB and calling pg_reload_conf():
-- 128 MB
7f5a2bd04000-7f5a32e52000 /dev/zero (deleted)
7f5a39252000-7f5a4030e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d7ba000 /dev/zero (deleted)
7f5a53bba000-7f5a5ad26000 /dev/zero (deleted)
7f5a9ad26000-7f5aa9d94000 /dev/zero (deleted)
^ buffers mapping, ~240 MB
7f5d29d94000-7f5d30e00000 /dev/zero (deleted)
-- 512 MB
7f5a2bd04000-7f5a33274000 /dev/zero (deleted)
7f5a39252000-7f5a4057e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d9fa000 /dev/zero (deleted)
7f5a53bba000-7f5a5b1a6000 /dev/zero (deleted)
7f5a9ad26000-7f5ac1f14000 /dev/zero (deleted)
^ buffers mapping, ~625 MB
7f5d29d94000-7f5d30f80000 /dev/zero (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend, executing pg_reload_conf interactively, will not resize
mappings immediately, for some reason it will require another command.
Note, that mremap is Linux specific, thus the implementation not very
portable.
---
src/backend/port/sysv_shmem.c | 62 +++++++++++++
src/backend/storage/buffer/buf_init.c | 86 +++++++++++++++++++
src/backend/storage/ipc/ipci.c | 11 +++
src/backend/storage/ipc/shmem.c | 14 ++-
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/storage/bufmgr.h | 1 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 2 +
9 files changed, 171 insertions(+), 11 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index beebd4d85e..4bdadbb0e2 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,9 +30,11 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
@@ -859,6 +861,66 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * An assign callback for shared_buffers GUC -- a somewhat clumsy way of
+ * resizing shared memory without a restart. On NBuffers change use the new
+ * value to recalculate required size for every shmem slot, then base on the
+ * new and old values initialize new buffer blocks.
+ *
+ * The actual slot resizing is done via mremap, which will fail if is not
+ * sufficient space to expand the mapping.
+ *
+ * XXX: For some readon in the current implementation the change is applied to
+ * the backend calling pg_reload_conf only at the backend exit.
+ */
+void
+AnonymousShmemResize(int newval, void *extra)
+{
+ int numSemas;
+ bool reinit = false;
+ int NBuffersOld = NBuffers;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffers > newval)
+ return;
+
+ /* XXX: Hack, NBuffers has to be exposed in the the interface for
+ * memory calculation and buffer blocks reinitialization instead. */
+ NBuffers = newval;
+
+ for(int i = 0; i < next_free_slot; i++)
+ {
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
+ elog(LOG, "mremap(%p, %zu) failed: %m",
+ m->shmem, m->shmem_size);
+ else
+ {
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+ }
+
+ if (reinit)
+ {
+ LWLockAcquire(ShmemResizeLock, LW_EXCLUSIVE);
+ ResizeBufferPool(NBuffersOld);
+ LWLockRelease(ShmemResizeLock);
+ }
+}
+
/*
* PGSharedMemoryCreate
*
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6bca286bef..4054abf0e8 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -154,6 +154,92 @@ InitBufferPool(void)
&backend_flush_after);
}
+/*
+ * Reinitialize shared memory structures, which size depends on NBuffers. It's
+ * similar to InitBufferPool, but applied only to the buffers in the range
+ * between NBuffersOld and NBuffers.
+ */
+void
+ResizeBufferPool(int NBuffersOld)
+{
+ bool foundBufs,
+ foundDescs,
+ foundIOCV,
+ foundBufCkpt;
+ int i;
+
+ /* XXX: Only increasing of shared_buffers is supported in this function */
+ if(NBuffersOld > NBuffers)
+ return;
+
+ /* Align descriptors to a cacheline boundary. */
+ BufferDescriptors = (BufferDescPadded *)
+ ShmemInitStructInSlot("Buffer Descriptors",
+ NBuffers * sizeof(BufferDescPadded),
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SLOT);
+
+ /* Align condition variables to cacheline boundary. */
+ BufferIOCVArray = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSlot("Buffer IO Condition Variables",
+ NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &foundIOCV, BUFFER_IOCV_SHMEM_SLOT);
+
+ /*
+ * The array used to sort to-be-checkpointed buffer ids is located in
+ * shared memory, to avoid having to allocate significant amounts of
+ * memory at runtime. As that'd be in the middle of a checkpoint, or when
+ * the checkpointer is restarted, memory allocation failures would be
+ * painful.
+ */
+ CkptBufferIds = (CkptSortItem *)
+ ShmemInitStructInSlot("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SLOT);
+
+ /* Align buffer pool on IO page size boundary. */
+ BufferBlocks = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSlot("Buffer Blocks",
+ NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &foundBufs, BUFFERS_SHMEM_SLOT));
+
+ /*
+ * Initialize the headers for new buffers.
+ */
+ for (i = NBuffersOld - 1; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
+
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+
+ buf->buf_id = i;
+
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
+
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
+
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ }
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+
+ /* Init other shared buffer-management stuff */
+ StrategyInitialize(!foundDescs);
+
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+}
+
/*
* BufferShmemSize
*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index fbaddba396..56fa339f55 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,6 +86,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory slots are incorrect, it includes
+ * more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_slot)
@@ -153,6 +156,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_slot)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, WaitLSNShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index c670b9cf43..20c4b1d5ad 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -491,17 +491,13 @@ ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
*/
if (result->size != size)
- {
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
- }
+ result->size = size;
+
structPtr = result->location;
}
else
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index d10ca723dc..42296d950e 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -347,6 +347,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
WaitLSN "Waiting to read or update shared Wait-for-LSN state."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 636780673b..7f2c45b7f9 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2301,14 +2301,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, AnonymousShmemResize, NULL
},
{
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 4c09d270c9..ff75c46307 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -302,6 +302,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void InitBufferPool(void);
extern Size BufferShmemSize(int);
+extern void ResizeBufferPool(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index 88dc79b2bd..fb310e8b9d 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, WaitLSN)
+PG_LWLOCK(54, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index c0143e3899..ff4736c6c8 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -105,6 +105,8 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+void AnonymousShmemResize(int newval, void *extra);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings slots. Each slot
--
2.45.1
From 6df85a35e8f6cca94a963d516f1b6974850ba05b Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 15 Oct 2024 16:18:45 +0200
Subject: [PATCH v1 5/5] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f5a2bd04000-7f5a32e52000 rw-s 00000000 00:01 1845 /memfd:strategy (deleted)
7f5a39252000-7f5a4030e000 rw-s 00000000 00:01 1842 /memfd:checkpoint (deleted)
7f5a4670e000-7f5a4d7ba000 rw-s 00000000 00:01 1839 /memfd:iocv (deleted)
7f5a53bba000-7f5a5ad26000 rw-s 00000000 00:01 1836 /memfd:descriptors (deleted)
7f5a9ad26000-7f5aa9d94000 rw-s 00000000 00:01 1833 /memfd:buffers (deleted)
7f5d29d94000-7f5d30e00000 rw-s 00000000 00:01 1830 /memfd:main (deleted)
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 47 +++++++++++++++++++++++++++++------
src/include/portability/mem.h | 2 +-
2 files changed, 40 insertions(+), 9 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 4bdadbb0e2..a01c3e4789 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -103,6 +103,7 @@ typedef struct AnonymousMapping
void *shmem; /* Pointer to the start of the mapped memory */
void *seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -116,7 +117,7 @@ static int next_free_slot = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -143,9 +144,9 @@ static int next_free_slot = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* [...free space...]
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* [...free space...]
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -708,6 +709,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_slot), 0);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
@@ -725,8 +738,13 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ /*
+ * Do not use an anonymous file here yet. When adding it, do not forget
+ * to use ftruncate and flags MFD_HUGETLB & MFD_HUGE_2MB/MFD_HUGE_1GB
+ * in memfd_create.
+ */
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
@@ -762,7 +780,8 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* - First create the temporary probe mapping of a fixed size and let
* kernel to place it at address of its choice. By the virtue of the
* probe mapping size we expect it to be located at the lowest
- * possible address, expecting some non mapped space above.
+ * possible address, expecting some non mapped space above. The probe
+ * is does not need to be backed by an anonymous file.
*
* - Unmap the probe mapping, remember the address.
*
@@ -777,7 +796,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* without a restart.
*/
probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS, -1, 0);
if (probe == MAP_FAILED)
{
@@ -793,8 +812,14 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
munmap(probe, PROBE_MAPPING_SIZE);
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
{
@@ -813,8 +838,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
*/
allocsize = mapping->shmem_size;
+ /* Specify the segment file size using allocsize. */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
@@ -903,6 +931,9 @@ AnonymousShmemResize(int newval, void *extra)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ ftruncate(m->segment_fd, new_size);
+
if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
elog(LOG, "mremap(%p, %zu) failed: %m",
m->shmem, m->shmem_size);
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index 2cd05313b8..50db0da28d 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
--
2.45.1
Attachments:
[text/plain] v1-0001-Allow-to-use-multiple-shared-memory-mappings.patch (28.7K, ../../cnthxg2eekacrejyeonuhiaezc7vd7o2uowlsbenxqfkjwgvwj@qgzu6eoqrglb/2-v1-0001-Allow-to-use-multiple-shared-memory-mappings.patch)
download | inline diff:
From 954613a63cb1102d7eb88f92e7ff561828bbb5c9 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 9 Oct 2024 15:41:32 +0200
Subject: [PATCH v1 1/5] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory slot.
There is only fixed amount of available slots, currently only one main
shared memory slot is allocated. A new shared memory API is introduces,
extended with a slot as a new parameter. As a path of least resistance,
the original API is kept in place, utilizing the main shared memory slot.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 +++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 2 +-
src/backend/storage/ipc/ipci.c | 61 ++++++------
src/backend/storage/ipc/shmem.c | 133 ++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 5 +-
src/include/storage/buf_internals.h | 1 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 10 ++
13 files changed, 258 insertions(+), 124 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 64186ec0a7..b97723d2ed 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSlot(PGSemaphoreShmemSize(maxSemas), shmem_slot);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 5b88a92bc9..8ef95b12c9 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -307,7 +307,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
struct stat statbuf;
@@ -328,7 +328,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSlot(PGSemaphoreShmemSize(maxSemas), shmem_slot);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 362a37d3b3..065a5b63ac 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_slot;
+ Size shmem_size; /* Size of the mapping */
+ void *shmem; /* Pointer to the start of the mapped memory */
+ void *seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping slots */
+static int next_free_slot = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_slot)
+{
+ switch (shmem_slot)
+ {
+ case MAIN_SHMEM_SLOT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_slot; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_slot), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("slot[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_slot)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_slot; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_slot];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_slot = next_free_slot;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_slot++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_slot;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_slot; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index f2b54bdfda..d62084cc0d 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index b06e4b8452..2aabd4a77f 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -68,7 +68,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 35fa2e1dda..8224015b53 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -88,7 +88,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_slot)
{
Size size;
int numSemas;
@@ -202,33 +202,36 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int slot = 0; slot < ANON_MAPPINGS; slot++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, slot);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSlot(seghdr, slot);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, slot);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSlot(slot);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -359,7 +362,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SLOT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 6d5f083986..c670b9cf43 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -75,17 +75,12 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSlot(Size size, Size *allocated_size,
+ int shmem_slot);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
-
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
+ShmemSegment Segments[ANON_MAPPINGS];
static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
@@ -99,11 +94,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
void
InitShmemAccess(void *seghdr)
{
- PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ InitShmemAccessInSlot(seghdr, MAIN_SHMEM_SLOT);
+}
- ShmemSegHdr = shmhdr;
- ShmemBase = (void *) shmhdr;
- ShmemEnd = (char *) ShmemBase + shmhdr->totalsize;
+void
+InitShmemAccessInSlot(void *seghdr, int shmem_slot)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_slot];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -114,7 +115,13 @@ InitShmemAccess(void *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSlot(MAIN_SHMEM_SLOT);
+}
+
+void
+InitShmemAllocationInSlot(int shmem_slot)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_slot].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -123,9 +130,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_slot].ShmemLock = (slock_t *) ShmemAllocUnlockedInSlot(sizeof(slock_t), shmem_slot);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_slot].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -150,11 +157,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSlot(size, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemAllocInSlot(Size size, int shmem_slot)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSlot(size, &allocated_size, shmem_slot);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -184,6 +197,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSlot(size, allocated_size, MAIN_SHMEM_SLOT);
+}
+
+static void *
+ShmemAllocRawInSlot(Size size, Size *allocated_size, int shmem_slot)
{
Size newStart;
Size newFree;
@@ -203,22 +222,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_slot].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_slot].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_slot].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_slot].ShmemSegHdr->totalsize)
{
- newSpace = (void *) ((char *) ShmemBase + newStart);
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (void *) ((char *) Segments[shmem_slot].ShmemBase + newStart);
+ Segments[shmem_slot].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_slot].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -236,6 +255,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSlot(size, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemAllocUnlockedInSlot(Size size, int shmem_slot)
{
Size newStart;
Size newFree;
@@ -246,19 +271,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_slot].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_slot].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_slot].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_slot].ShmemSegHdr->freeoffset = newFree;
- newSpace = (void *) ((char *) ShmemBase + newStart);
+ newSpace = (void *) ((char *) Segments[shmem_slot].ShmemBase + newStart);
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -273,7 +298,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSlot(addr, MAIN_SHMEM_SLOT);
+}
+
+bool
+ShmemAddrIsValidInSlot(const void *addr, int shmem_slot)
+{
+ return (addr >= Segments[shmem_slot].ShmemBase) && (addr < Segments[shmem_slot].ShmemEnd);
}
/*
@@ -334,6 +365,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSlot(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SLOT);
+}
+
+HTAB *
+ShmemInitHashInSlot(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_slot) /* in which slot to keep the table */
{
bool found;
void *location;
@@ -350,9 +393,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSlot(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_slot);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -385,6 +428,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSlot(name, size, foundPtr, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
+ int shmem_slot)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -393,7 +443,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_slot].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -416,7 +466,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSlot(size, shmem_slot);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -433,8 +483,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in slot %d",
+ name, shmem_slot)));
}
if (*foundPtr)
@@ -459,7 +509,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSlot(size, &allocated_size, shmem_slot);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -478,14 +528,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSlot(structPtr, shmem_slot));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -545,7 +594,7 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SLOT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -557,15 +606,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index e765754d80..fb0c33bf17 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -81,6 +81,7 @@
#include "pgstat.h"
#include "port/pg_bitutils.h"
#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/spin.h"
@@ -607,9 +608,9 @@ LWLockNewTrancheId(void)
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
/* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SLOT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SLOT].ShmemLock);
return result;
}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index f190e6e5e4..aef80e049b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -23,6 +23,7 @@
#include "storage/latch.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
+#include "storage/pg_shmem.h"
#include "storage/smgr.h"
#include "storage/spin.h"
#include "utils/relcache.h"
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index b2d062781e..be4b131288 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_slot);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index dfef79ac96..081fffaf16 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_slot);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 3065ff5be7..e968deeef7 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+// Number of available slots for anonymous memory mappings
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main slot, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SLOT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 842989111c..d3e9cc721d 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -28,15 +28,25 @@
/* shmem.c */
extern PGDLLIMPORT slock_t *ShmemLock;
extern void InitShmemAccess(void *seghdr);
+extern void InitShmemAccessInSlot(void *seghdr, int shmem_slot);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSlot(int shmem_slot);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSlot(Size size, int shmem_slot);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSlot(Size size, int shmem_slot);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSlot(const void *addr, int shmem_slot);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSlot(const char *name, long init_size, long max_size,
+ HASHCTL *infoP, int hash_flags, int shmem_slot);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
+ int shmem_slot);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: 2488058dc356a43455b21a099ea879fff9266634
--
2.45.1
[text/plain] v1-0002-Allow-placing-shared-memory-mapping-with-an-offse.patch (8.6K, ../../cnthxg2eekacrejyeonuhiaezc7vd7o2uowlsbenxqfkjwgvwj@qgzu6eoqrglb/3-v1-0002-Allow-placing-shared-memory-mapping-with-an-offse.patch)
download | inline diff:
From e9980f76cbd1ea6f6d732e2a27dd1342258d26e5 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH v1 2/5] Allow placing shared memory mapping with an offset
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is first to get
the address chosen by the kernel, then apply some offset derived from
the expected upper limit. Because we base the layout on the address
chosen by the kernel, things like address space randomization should not
be a problem, since the randomization is applied to the mmap base, which
is one per process. The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
[...free space...]
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
This approach do not impact the actual memory usage as reported by the kernel.
Here is the output of /proc/$PID/status for the master version with
shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages mapped in mm_struct
VmPeak: 422780 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 21248 kB
// Size of resident anonymous memory
RssAnon: 640 kB
// Size of resident file mappings
RssFile: 9728 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 10880 kB
Here is the same for the patch with the shared mapping placed at
an offset 10 GB:
VmPeak: 1102844 kB
VmRSS: 21376 kB
RssAnon: 640 kB
RssFile: 9856 kB
RssShmem: 10880 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
18219008 (~17 MB)
Note that currently the implementation makes assumptions about the upper limit.
Ideally it should be based on the maximum available memory.
---
src/backend/port/sysv_shmem.c | 120 +++++++++++++++++++++++++++++++++-
1 file changed, 119 insertions(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 065a5b63ac..7e6c8bb78d 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,63 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping slots */
static int next_free_slot = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * By letting Linux to chose a mapping address, it will pick up the lowest
+ * possible address that do not clash with any other mappings, which will be
+ * right before locales in the example above. This information (maximum allowed
+ * size of mappings and the lowest mapping address) is enough to place every
+ * mapping as follow:
+ *
+ * - Take the lowest mapping address, which we call later the probe address.
+ * - Substract the offset of the previous mapping.
+ * - Substract the maximum allowed size for the current mapping from the
+ * address.
+ * - Place the mapping by the resulting address.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ */
+Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+ 0, /* MAIN_SHMEM_SLOT */
+};
+
+/* Remembers offset of the last mapping from the probe address */
+static Size last_offset = 0;
+
+/*
+ * Size of the mapping, which will be used to calculate anonymous mapping
+ * address. It should not be too small, otherwise there is a chance the probe
+ * mapping will be created between other mappings, leaving no room extending
+ * it. But it should not be too large either, in case if there are limitations
+ * on the mapping size. Current value is the default shared_buffers.
+ */
+#define PROBE_MAPPING_SIZE (Size) 128 * 1024 * 1024
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -673,13 +730,74 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
+ void *probe = NULL;
+
/*
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+
+ /*
+ * Try to create mapping at an address, which will allow to extend it
+ * later:
+ *
+ * - First create the temporary probe mapping of a fixed size and let
+ * kernel to place it at address of its choice. By the virtue of the
+ * probe mapping size we expect it to be located at the lowest
+ * possible address, expecting some non mapped space above.
+ *
+ * - Unmap the probe mapping, remember the address.
+ *
+ * - Create an actual anonymous mapping at that address with the
+ * offset. The offset is calculated in such a way to allow growing
+ * the mapping withing certain boundaries. For this mapping we use
+ * MAP_FIXED_NOREPLACE, which will error out with EEXIST if there is
+ * any mapping clash.
+ *
+ * - If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized
+ * without a restart.
+ */
+ probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
+
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: probe mmap(%zu) failed: %m",
+ MappingName(mapping->shmem_slot), allocsize);
+ }
+ else
+ {
+ Size offset = last_offset + SHMEM_EXTRA_SIZE_LIMIT[next_free_slot] + allocsize;
+ last_offset = offset;
+
+ munmap(probe, PROBE_MAPPING_SIZE);
+
+ ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ mmap_errno = errno;
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: mmap(%zu) at address %p failed: %m",
+ MappingName(mapping->shmem_slot), allocsize, probe - offset);
+ }
+
+ }
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ /*
+ * Fallback to the portable way of creating a mapping.
+ */
+ allocsize = mapping->shmem_size;
+
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
}
--
2.45.1
[text/plain] v1-0003-Introduce-multiple-shmem-slots-for-shared-buffers.patch (10.7K, ../../cnthxg2eekacrejyeonuhiaezc7vd7o2uowlsbenxqfkjwgvwj@qgzu6eoqrglb/4-v1-0003-Introduce-multiple-shmem-slots-for-shared-buffers.patch)
download | inline diff:
From 62ae567f1c7a56c32722508b60251f9cec245ea3 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:24:04 +0200
Subject: [PATCH v1 3/5] Introduce multiple shmem slots for shared buffers
Add more shmem slots to split shared buffers into following chunks:
* BUFFERS_SHMEM_SLOT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SLOT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SLOT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SLOT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SLOT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers, meaning
that if we would like to change NBuffers, they have to be resized
correspondingly. Placing each of them in a separate shmem slot allows to
achieve that.
There are some asumptions made about each of shmem slots upper size limit. The
buffer blocks have the largest, while the rest claim less extra room for
resize. Ideally those limits have to be deduced from the maximum allowed shared
memory.
---
src/backend/port/sysv_shmem.c | 17 +++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 5 +-
src/backend/storage/buffer/freelist.c | 4 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 23 +++++++-
7 files changed, 97 insertions(+), 35 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 7e6c8bb78d..beebd4d85e 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -149,8 +149,13 @@ static int next_free_slot = 0;
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
*/
-Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+Size SHMEM_EXTRA_SIZE_LIMIT[6] = {
0, /* MAIN_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 1024 * 10, /* BUFFERS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 1024 * 1, /* BUFFER_DESCRIPTORS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* BUFFER_IOCV_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* CHECKPOINT_BUFFERS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* STRATEGY_SHMEM_SLOT */
};
/* Remembers offset of the last mapping from the probe address */
@@ -179,6 +184,16 @@ MappingName(int shmem_slot)
{
case MAIN_SHMEM_SLOT:
return "main";
+ case BUFFERS_SHMEM_SLOT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SLOT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SLOT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SLOT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SLOT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 46116a1f64..6bca286bef 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory slot, which
+ * could be resized on demand.
*/
void
InitBufferPool(void)
@@ -74,22 +77,22 @@ InitBufferPool(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSlot("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SLOT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSlot("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SLOT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSlot("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SLOT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ InitBufferPool(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSlot("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SLOT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -154,33 +158,54 @@ InitBufferPool(void)
* BufferShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory slot. The main slot must not allocate anything
+ * related to buffers, every other slot will receive part of the
+ * data.
*/
Size
-BufferShmemSize(void)
+BufferShmemSize(int shmem_slot)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_slot == MAIN_SHMEM_SLOT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_slot == BUFFER_DESCRIPTORS_SHMEM_SLOT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_slot == BUFFERS_SHMEM_SLOT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_slot == STRATEGY_SHMEM_SLOT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_slot == BUFFER_IOCV_SHMEM_SLOT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_slot == CHECKPOINT_BUFFERS_SHMEM_SLOT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 0fa5468930..ccbaed8010 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -59,10 +59,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSlot("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SLOT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 19797de31a..8ce1611db2 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -491,9 +491,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSlot("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SLOT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8224015b53..fbaddba396 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -115,7 +115,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_slot)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferShmemSize());
+ size = add_size(size, BufferShmemSize(shmem_slot));
size = add_size(size, LockShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index c8422571b7..4c09d270c9 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -301,7 +301,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void InitBufferPool(void);
-extern Size BufferShmemSize(void);
+extern Size BufferShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index e968deeef7..c0143e3899 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
// Number of available slots for anonymous memory mappings
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -105,7 +105,28 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings slots. Each slot
+ * contains only certain part of the data, which size depends on NBuffers.
+ */
+
/* The main slot, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SLOT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SLOT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SLOT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SLOT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SLOT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SLOT 5
+
#endif /* PG_SHMEM_H */
--
2.45.1
[text/plain] v1-0004-Allow-to-resize-shared-memory-without-restart.patch (12.4K, ../../cnthxg2eekacrejyeonuhiaezc7vd7o2uowlsbenxqfkjwgvwj@qgzu6eoqrglb/5-v1-0004-Allow-to-resize-shared-memory-without-restart.patch)
download | inline diff:
From 7183999bba1cbeebd059d18e5a590cbef7aff2d1 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:24:58 +0200
Subject: [PATCH v1 4/5] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Size for every shared memory slot is recalculated based on the new
NBuffers, and extended using mremap. After allocating new space, new
shared structures (buffer blocks, descriptors, etc) are allocated as
needed. Here is how it looks like after raising shared_buffers from 128
MB to 512 MB and calling pg_reload_conf():
-- 128 MB
7f5a2bd04000-7f5a32e52000 /dev/zero (deleted)
7f5a39252000-7f5a4030e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d7ba000 /dev/zero (deleted)
7f5a53bba000-7f5a5ad26000 /dev/zero (deleted)
7f5a9ad26000-7f5aa9d94000 /dev/zero (deleted)
^ buffers mapping, ~240 MB
7f5d29d94000-7f5d30e00000 /dev/zero (deleted)
-- 512 MB
7f5a2bd04000-7f5a33274000 /dev/zero (deleted)
7f5a39252000-7f5a4057e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d9fa000 /dev/zero (deleted)
7f5a53bba000-7f5a5b1a6000 /dev/zero (deleted)
7f5a9ad26000-7f5ac1f14000 /dev/zero (deleted)
^ buffers mapping, ~625 MB
7f5d29d94000-7f5d30f80000 /dev/zero (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend, executing pg_reload_conf interactively, will not resize
mappings immediately, for some reason it will require another command.
Note, that mremap is Linux specific, thus the implementation not very
portable.
---
src/backend/port/sysv_shmem.c | 62 +++++++++++++
src/backend/storage/buffer/buf_init.c | 86 +++++++++++++++++++
src/backend/storage/ipc/ipci.c | 11 +++
src/backend/storage/ipc/shmem.c | 14 ++-
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/storage/bufmgr.h | 1 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 2 +
9 files changed, 171 insertions(+), 11 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index beebd4d85e..4bdadbb0e2 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,9 +30,11 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
@@ -859,6 +861,66 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * An assign callback for shared_buffers GUC -- a somewhat clumsy way of
+ * resizing shared memory without a restart. On NBuffers change use the new
+ * value to recalculate required size for every shmem slot, then base on the
+ * new and old values initialize new buffer blocks.
+ *
+ * The actual slot resizing is done via mremap, which will fail if is not
+ * sufficient space to expand the mapping.
+ *
+ * XXX: For some readon in the current implementation the change is applied to
+ * the backend calling pg_reload_conf only at the backend exit.
+ */
+void
+AnonymousShmemResize(int newval, void *extra)
+{
+ int numSemas;
+ bool reinit = false;
+ int NBuffersOld = NBuffers;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffers > newval)
+ return;
+
+ /* XXX: Hack, NBuffers has to be exposed in the the interface for
+ * memory calculation and buffer blocks reinitialization instead. */
+ NBuffers = newval;
+
+ for(int i = 0; i < next_free_slot; i++)
+ {
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
+ elog(LOG, "mremap(%p, %zu) failed: %m",
+ m->shmem, m->shmem_size);
+ else
+ {
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+ }
+
+ if (reinit)
+ {
+ LWLockAcquire(ShmemResizeLock, LW_EXCLUSIVE);
+ ResizeBufferPool(NBuffersOld);
+ LWLockRelease(ShmemResizeLock);
+ }
+}
+
/*
* PGSharedMemoryCreate
*
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6bca286bef..4054abf0e8 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -154,6 +154,92 @@ InitBufferPool(void)
&backend_flush_after);
}
+/*
+ * Reinitialize shared memory structures, which size depends on NBuffers. It's
+ * similar to InitBufferPool, but applied only to the buffers in the range
+ * between NBuffersOld and NBuffers.
+ */
+void
+ResizeBufferPool(int NBuffersOld)
+{
+ bool foundBufs,
+ foundDescs,
+ foundIOCV,
+ foundBufCkpt;
+ int i;
+
+ /* XXX: Only increasing of shared_buffers is supported in this function */
+ if(NBuffersOld > NBuffers)
+ return;
+
+ /* Align descriptors to a cacheline boundary. */
+ BufferDescriptors = (BufferDescPadded *)
+ ShmemInitStructInSlot("Buffer Descriptors",
+ NBuffers * sizeof(BufferDescPadded),
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SLOT);
+
+ /* Align condition variables to cacheline boundary. */
+ BufferIOCVArray = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSlot("Buffer IO Condition Variables",
+ NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &foundIOCV, BUFFER_IOCV_SHMEM_SLOT);
+
+ /*
+ * The array used to sort to-be-checkpointed buffer ids is located in
+ * shared memory, to avoid having to allocate significant amounts of
+ * memory at runtime. As that'd be in the middle of a checkpoint, or when
+ * the checkpointer is restarted, memory allocation failures would be
+ * painful.
+ */
+ CkptBufferIds = (CkptSortItem *)
+ ShmemInitStructInSlot("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SLOT);
+
+ /* Align buffer pool on IO page size boundary. */
+ BufferBlocks = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSlot("Buffer Blocks",
+ NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &foundBufs, BUFFERS_SHMEM_SLOT));
+
+ /*
+ * Initialize the headers for new buffers.
+ */
+ for (i = NBuffersOld - 1; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
+
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+
+ buf->buf_id = i;
+
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
+
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
+
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ }
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+
+ /* Init other shared buffer-management stuff */
+ StrategyInitialize(!foundDescs);
+
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+}
+
/*
* BufferShmemSize
*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index fbaddba396..56fa339f55 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,6 +86,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory slots are incorrect, it includes
+ * more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_slot)
@@ -153,6 +156,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_slot)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, WaitLSNShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index c670b9cf43..20c4b1d5ad 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -491,17 +491,13 @@ ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
*/
if (result->size != size)
- {
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
- }
+ result->size = size;
+
structPtr = result->location;
}
else
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index d10ca723dc..42296d950e 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -347,6 +347,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
WaitLSN "Waiting to read or update shared Wait-for-LSN state."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 636780673b..7f2c45b7f9 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2301,14 +2301,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, AnonymousShmemResize, NULL
},
{
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 4c09d270c9..ff75c46307 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -302,6 +302,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void InitBufferPool(void);
extern Size BufferShmemSize(int);
+extern void ResizeBufferPool(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index 88dc79b2bd..fb310e8b9d 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, WaitLSN)
+PG_LWLOCK(54, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index c0143e3899..ff4736c6c8 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -105,6 +105,8 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+void AnonymousShmemResize(int newval, void *extra);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings slots. Each slot
--
2.45.1
[text/plain] v1-0005-Use-anonymous-files-to-back-shared-memory-segment.patch (6.8K, ../../cnthxg2eekacrejyeonuhiaezc7vd7o2uowlsbenxqfkjwgvwj@qgzu6eoqrglb/6-v1-0005-Use-anonymous-files-to-back-shared-memory-segment.patch)
download | inline diff:
From 6df85a35e8f6cca94a963d516f1b6974850ba05b Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 15 Oct 2024 16:18:45 +0200
Subject: [PATCH v1 5/5] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f5a2bd04000-7f5a32e52000 rw-s 00000000 00:01 1845 /memfd:strategy (deleted)
7f5a39252000-7f5a4030e000 rw-s 00000000 00:01 1842 /memfd:checkpoint (deleted)
7f5a4670e000-7f5a4d7ba000 rw-s 00000000 00:01 1839 /memfd:iocv (deleted)
7f5a53bba000-7f5a5ad26000 rw-s 00000000 00:01 1836 /memfd:descriptors (deleted)
7f5a9ad26000-7f5aa9d94000 rw-s 00000000 00:01 1833 /memfd:buffers (deleted)
7f5d29d94000-7f5d30e00000 rw-s 00000000 00:01 1830 /memfd:main (deleted)
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 47 +++++++++++++++++++++++++++++------
src/include/portability/mem.h | 2 +-
2 files changed, 40 insertions(+), 9 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 4bdadbb0e2..a01c3e4789 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -103,6 +103,7 @@ typedef struct AnonymousMapping
void *shmem; /* Pointer to the start of the mapped memory */
void *seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -116,7 +117,7 @@ static int next_free_slot = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -143,9 +144,9 @@ static int next_free_slot = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* [...free space...]
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* [...free space...]
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -708,6 +709,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_slot), 0);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
@@ -725,8 +738,13 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ /*
+ * Do not use an anonymous file here yet. When adding it, do not forget
+ * to use ftruncate and flags MFD_HUGETLB & MFD_HUGE_2MB/MFD_HUGE_1GB
+ * in memfd_create.
+ */
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
@@ -762,7 +780,8 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* - First create the temporary probe mapping of a fixed size and let
* kernel to place it at address of its choice. By the virtue of the
* probe mapping size we expect it to be located at the lowest
- * possible address, expecting some non mapped space above.
+ * possible address, expecting some non mapped space above. The probe
+ * is does not need to be backed by an anonymous file.
*
* - Unmap the probe mapping, remember the address.
*
@@ -777,7 +796,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* without a restart.
*/
probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS, -1, 0);
if (probe == MAP_FAILED)
{
@@ -793,8 +812,14 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
munmap(probe, PROBE_MAPPING_SIZE);
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
{
@@ -813,8 +838,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
*/
allocsize = mapping->shmem_size;
+ /* Specify the segment file size using allocsize. */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
@@ -903,6 +931,9 @@ AnonymousShmemResize(int newval, void *extra)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ ftruncate(m->segment_fd, new_size);
+
if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
elog(LOG, "mremap(%p, %zu) failed: %m",
m->shmem, m->shmem_size);
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index 2cd05313b8..50db0da28d 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
--
2.45.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-01 15:27 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-06 19:10 ` Re: Changing shared_buffers without restart Vladlen Popolitov <v.popolitov@postgrespro.ru>
4 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-01 15:27 UTC (permalink / raw)
To: pgsql-hackers
> On Fri, Oct 18, 2024 at 09:21:19PM GMT, Dmitry Dolgov wrote:
>
> TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
> changing shared memory mapping layout. Any feedback is appreciated.
It was pointed out to me, that earlier this year there was a useful
discussion about similar matters "PGC_SIGHUP shared_buffers?" [1]. From
what I see the patch series falls into the "re-map" category in that
thread.
[1]: https://www.postgresql.org/message-id/flat/CA%2BTgmoaGCFPhMjz7veJOeef30%3DKdpOxgywcLwNbr-Gny-mXwcg%4...
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-01 15:27 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-06 19:10 ` Vladlen Popolitov <v.popolitov@postgrespro.ru>
2024-11-08 16:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Vladlen Popolitov @ 2024-11-06 19:10 UTC (permalink / raw)
To: pgsql-hackers@lists.postgresql.org; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>
Hi
I tried to apply patches, but failed. I suppose the problem with CRLF in the end of lines in the patch files. At least, after manual change of v1-0001 and v1-0002 from CRLF to LF patches applied, but it was not helped for v1-0003 - v1.0005 - they have also other mistakes during patch process. Could you check patch files and place them in correct format?
The new status of this patch is: Waiting on Author
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-01 15:27 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-06 19:10 ` Re: Changing shared_buffers without restart Vladlen Popolitov <v.popolitov@postgrespro.ru>
@ 2024-11-08 16:43 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-08 16:43 UTC (permalink / raw)
To: Vladlen Popolitov <v.popolitov@postgrespro.ru>; +Cc: pgsql-hackers@lists.postgresql.org
> On Wed, Nov 06, 2024 at 07:10:06PM GMT, Vladlen Popolitov wrote:
> Hi
>
> I tried to apply patches, but failed. I suppose the problem with CRLF in the end of lines in the patch files. At least, after manual change of v1-0001 and v1-0002 from CRLF to LF patches applied, but it was not helped for v1-0003 - v1.0005 - they have also other mistakes during patch process. Could you check patch files and place them in correct format?
>
> The new status of this patch is: Waiting on Author
Well, I'm going to rebase the patch if that's what you mean. But just
FYI -- it could be applied without any issues to the base commit
mentioned in the series.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-07 01:05 ` Thomas Munro <thomas.munro@gmail.com>
2024-11-08 16:40 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
4 siblings, 1 reply; 167+ messages in thread
From: Thomas Munro @ 2024-11-07 01:05 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers
On Sat, Oct 19, 2024 at 8:21 AM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> Currently it
> supports only an increase of shared_buffers.
Just BTW in case it is interesting, Palak and I experimented with how
to shrink the buffer pool while PostgreSQL is running, while we were
talking about 13453ee (which it shares infrastructure with). This
version fails if something is pinned and in the way of the shrink
operation, but you could imagine other policies (wait, cancel it,
...):
https://github.com/macdice/postgres/commit/db26fe0c98476cdbbd1bcf553f3b7864cb142247
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-07 01:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
@ 2024-11-08 16:40 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-08 16:40 UTC (permalink / raw)
To: Thomas Munro <thomas.munro@gmail.com>; +Cc: pgsql-hackers
> On Thu, Nov 07, 2024 at 02:05:52PM GMT, Thomas Munro wrote:
> On Sat, Oct 19, 2024 at 8:21 AM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > Currently it
> > supports only an increase of shared_buffers.
>
> Just BTW in case it is interesting, Palak and I experimented with how
> to shrink the buffer pool while PostgreSQL is running, while we were
> talking about 13453ee (which it shares infrastructure with). This
> version fails if something is pinned and in the way of the shrink
> operation, but you could imagine other policies (wait, cancel it,
> ...):
>
> https://github.com/macdice/postgres/commit/db26fe0c98476cdbbd1bcf553f3b7864cb142247
Thanks, looks interesting. I'll try to experiment with that in the next
version of the patch.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-19 12:57 ` Peter Eisentraut <peter@eisentraut.org>
2024-11-19 13:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
4 siblings, 1 reply; 167+ messages in thread
From: Peter Eisentraut @ 2024-11-19 12:57 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
On 18.10.24 21:21, Dmitry Dolgov wrote:
> v1-0001-Allow-to-use-multiple-shared-memory-mappings.patch
>
> Preparation, introduces the possibility to work with many shmem mappings. To
> make it less invasive, I've duplicated the shmem API to extend it with the
> shmem_slot argument, while redirecting the original API to it. There are
> probably better ways of doing that, I'm open for suggestions.
After studying this a bit, I tend to think you should just change the
existing APIs in place. So for example,
void *ShmemAlloc(Size size);
becomes
void *ShmemAlloc(int shmem_slot, Size size);
There aren't that many callers, and all these duplicated interfaces
almost add more new code than they save.
It might be worth making exceptions for interfaces that are likely to be
used by extensions. For example, I see pg_stat_statements using
ShmemInitStruct() and ShmemInitHash(). But that seems to be it. Are
there any other examples out there? Maybe there are many more that I
don't see right now. But at least for the initialization functions, it
doesn't seem worth it to preserve the existing interfaces exactly.
In any case, I think the slot number should be the first argument. This
matches how MemoryContextAlloc() or also talloc() work.
(Now here is an idea: Could these just be memory contexts? Instead of
making six shared memory slots, could you make six memory contexts with
a special shared memory type. And ShmemAlloc becomes the allocation
function, etc.?)
I noticed the existing code made inconsistent use of PGShmemHeader * vs.
void *, which also bled into your patch. I made the attached little
patch to clean that up a bit.
I suggest splitting the struct ShmemSegment into one struct for the
three memory addresses and a separate array just for the slock_t's. The
former struct can then stay private in storage/ipc/shmem.c, only the
locks need to be exported.
Maybe rename ANON_MAPPINGS to something like NUM_ANON_MAPPINGS.
Also, maybe some of this should be declared in storage/shmem.h rather
than in storage/pg_shmem.h. We have the existing ShmemLock in there, so
it would be a bit confusing to have the per-segment locks elsewhere.
> v1-0003-Introduce-multiple-shmem-slots-for-shared-buffers.patch
>
> Splits shared_buffers into multiple slots, moving out structures that depend on
> NBuffers into separate mappings. There are two large gaps here:
>
> * Shmem size calculation for those mappings is not correct yet, it includes too
> many other things (no particular issues here, just haven't had time).
> * It makes hardcoded assumptions about what is the upper limit for resizing,
> which is currently low purely for experiments. Ideally there should be a new
> configuration option to specify the total available memory, which would be a
> base for subsequent calculations.
Yes, I imagine a shared_buffers_hard_limit setting. We could maybe
default that to the total available memory, but it would also be good to
be able to specify it directly, for testing.
> v1-0005-Use-anonymous-files-to-back-shared-memory-segment.patch
>
> Allows an anonyous file to back a shared mapping. This makes certain things
> easier, e.g. mappings visual representation, and gives an fd for possible
> future customizations.
I think this could be a useful patch just by itself, without the rest of
the series, because of
> * By default, Linux will not add file-backed shared mappings into a
> core dump, making it more convenient to work with them in PostgreSQL:
> no more huge dumps to process.
This could be significant operational benefit.
When you say "by default", is this adjustable? Does someone actually
want the whole shared memory in their core file? (If it's adjustable,
is it also adjustable for anonymous mappings?)
I'm wondering about this change:
-#define PG_MMAP_FLAGS
(MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
It looks like this would affect all mmap() calls, not only the one
you're changing. But that's the only one that uses this macro! I don't
understand why we need this; I don't see anything in the commit log
about this ever being used for any portability. I think we should just
get rid of it and have mmap() use the right flags directly.
I see that FreeBSD has a memfd_create() function. Might be worth a try.
Obviously, this whole thing needs a configure test for memfd_create()
anyway.
I see that memfd_create() has a MFD_HUGETLB flag. It's not very clear
how that interacts with the MAP_HUGETLB flag for mmap(). Do you need to
specify both of them if you want huge pages?
From 78562cb315da1cc5b35c07aba5a3fd7faacdad48 Mon Sep 17 00:00:00 2001
From: Peter Eisentraut <peter@eisentraut.org>
Date: Tue, 19 Nov 2024 13:15:15 +0100
Subject: [PATCH] More thorough use of PGShmemHeader * instead of void *
---
src/backend/port/sysv_shmem.c | 4 ++--
src/backend/port/win32_shmem.c | 4 ++--
src/backend/postmaster/launch_backend.c | 2 +-
src/backend/storage/ipc/shmem.c | 13 ++++---------
src/include/storage/pg_shmem.h | 2 +-
src/include/storage/shmem.h | 3 ++-
6 files changed, 12 insertions(+), 16 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 362a37d3b3a..fa6ee15ce56 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -92,7 +92,7 @@ typedef enum
unsigned long UsedShmemSegID = 0;
-void *UsedShmemSegAddr = NULL;
+PGShmemHeader *UsedShmemSegAddr = NULL;
static Size AnonymousShmemSize;
static void *AnonymousShmem = NULL;
@@ -892,7 +892,7 @@ PGSharedMemoryReAttach(void)
IpcMemoryId shmid;
PGShmemHeader *hdr;
IpcMemoryState state;
- void *origUsedShmemSegAddr = UsedShmemSegAddr;
+ PGShmemHeader *origUsedShmemSegAddr = UsedShmemSegAddr;
Assert(UsedShmemSegAddr != NULL);
Assert(IsUnderPostmaster);
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 3bcce9d3b63..827f9cd79b4 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -42,7 +42,7 @@
void *ShmemProtectiveRegion = NULL;
HANDLE UsedShmemSegID = INVALID_HANDLE_VALUE;
-void *UsedShmemSegAddr = NULL;
+PGShmemHeader *UsedShmemSegAddr = NULL;
static Size UsedShmemSegSize = 0;
static bool EnableLockPagesPrivilege(int elevel);
@@ -424,7 +424,7 @@ void
PGSharedMemoryReAttach(void)
{
PGShmemHeader *hdr;
- void *origUsedShmemSegAddr = UsedShmemSegAddr;
+ PGShmemHeader *origUsedShmemSegAddr = UsedShmemSegAddr;
Assert(ShmemProtectiveRegion != NULL);
Assert(UsedShmemSegAddr != NULL);
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index 1f2d829ec5a..8f48c938968 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -94,7 +94,7 @@ typedef struct
void *ShmemProtectiveRegion;
HANDLE UsedShmemSegID;
#endif
- void *UsedShmemSegAddr;
+ PGShmemHeader *UsedShmemSegAddr;
slock_t *ShmemLock;
#ifdef USE_INJECTION_POINTS
struct InjectionPointsCtl *ActiveInjectionPoints;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 6d5f0839864..50f987ae240 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -92,18 +92,13 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
/*
* InitShmemAccess() --- set up basic pointers to shared memory.
- *
- * Note: the argument should be declared "PGShmemHeader *seghdr",
- * but we use void to avoid having to include ipc.h in shmem.h.
*/
void
-InitShmemAccess(void *seghdr)
+InitShmemAccess(PGShmemHeader *seghdr)
{
- PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
-
- ShmemSegHdr = shmhdr;
- ShmemBase = (void *) shmhdr;
- ShmemEnd = (char *) ShmemBase + shmhdr->totalsize;
+ ShmemSegHdr = seghdr;
+ ShmemBase = seghdr;
+ ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
}
/*
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 3065ff5be71..7a07c5807ac 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -69,7 +69,7 @@ extern PGDLLIMPORT unsigned long UsedShmemSegID;
extern PGDLLIMPORT HANDLE UsedShmemSegID;
extern PGDLLIMPORT void *ShmemProtectiveRegion;
#endif
-extern PGDLLIMPORT void *UsedShmemSegAddr;
+extern PGDLLIMPORT PGShmemHeader *UsedShmemSegAddr;
#if !defined(WIN32) && !defined(EXEC_BACKEND)
#define DEFAULT_SHARED_MEMORY_TYPE SHMEM_TYPE_MMAP
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 842989111c3..8cdbe7a89c8 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -27,7 +27,8 @@
/* shmem.c */
extern PGDLLIMPORT slock_t *ShmemLock;
-extern void InitShmemAccess(void *seghdr);
+struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
+extern void InitShmemAccess(struct PGShmemHeader *seghdr);
extern void InitShmemAllocation(void);
extern void *ShmemAlloc(Size size);
extern void *ShmemAllocNoError(Size size);
--
2.47.0
Attachments:
[text/plain] 0001-More-thorough-use-of-PGShmemHeader-instead-of-void.patch.nocfbot (4.4K, ../../12add41a-7625-4639-a394-a5563e349322@eisentraut.org/2-0001-More-thorough-use-of-PGShmemHeader-instead-of-void.patch.nocfbot)
download | inline diff:
From 78562cb315da1cc5b35c07aba5a3fd7faacdad48 Mon Sep 17 00:00:00 2001
From: Peter Eisentraut <peter@eisentraut.org>
Date: Tue, 19 Nov 2024 13:15:15 +0100
Subject: [PATCH] More thorough use of PGShmemHeader * instead of void *
---
src/backend/port/sysv_shmem.c | 4 ++--
src/backend/port/win32_shmem.c | 4 ++--
src/backend/postmaster/launch_backend.c | 2 +-
src/backend/storage/ipc/shmem.c | 13 ++++---------
src/include/storage/pg_shmem.h | 2 +-
src/include/storage/shmem.h | 3 ++-
6 files changed, 12 insertions(+), 16 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 362a37d3b3a..fa6ee15ce56 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -92,7 +92,7 @@ typedef enum
unsigned long UsedShmemSegID = 0;
-void *UsedShmemSegAddr = NULL;
+PGShmemHeader *UsedShmemSegAddr = NULL;
static Size AnonymousShmemSize;
static void *AnonymousShmem = NULL;
@@ -892,7 +892,7 @@ PGSharedMemoryReAttach(void)
IpcMemoryId shmid;
PGShmemHeader *hdr;
IpcMemoryState state;
- void *origUsedShmemSegAddr = UsedShmemSegAddr;
+ PGShmemHeader *origUsedShmemSegAddr = UsedShmemSegAddr;
Assert(UsedShmemSegAddr != NULL);
Assert(IsUnderPostmaster);
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 3bcce9d3b63..827f9cd79b4 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -42,7 +42,7 @@
void *ShmemProtectiveRegion = NULL;
HANDLE UsedShmemSegID = INVALID_HANDLE_VALUE;
-void *UsedShmemSegAddr = NULL;
+PGShmemHeader *UsedShmemSegAddr = NULL;
static Size UsedShmemSegSize = 0;
static bool EnableLockPagesPrivilege(int elevel);
@@ -424,7 +424,7 @@ void
PGSharedMemoryReAttach(void)
{
PGShmemHeader *hdr;
- void *origUsedShmemSegAddr = UsedShmemSegAddr;
+ PGShmemHeader *origUsedShmemSegAddr = UsedShmemSegAddr;
Assert(ShmemProtectiveRegion != NULL);
Assert(UsedShmemSegAddr != NULL);
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index 1f2d829ec5a..8f48c938968 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -94,7 +94,7 @@ typedef struct
void *ShmemProtectiveRegion;
HANDLE UsedShmemSegID;
#endif
- void *UsedShmemSegAddr;
+ PGShmemHeader *UsedShmemSegAddr;
slock_t *ShmemLock;
#ifdef USE_INJECTION_POINTS
struct InjectionPointsCtl *ActiveInjectionPoints;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 6d5f0839864..50f987ae240 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -92,18 +92,13 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
/*
* InitShmemAccess() --- set up basic pointers to shared memory.
- *
- * Note: the argument should be declared "PGShmemHeader *seghdr",
- * but we use void to avoid having to include ipc.h in shmem.h.
*/
void
-InitShmemAccess(void *seghdr)
+InitShmemAccess(PGShmemHeader *seghdr)
{
- PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
-
- ShmemSegHdr = shmhdr;
- ShmemBase = (void *) shmhdr;
- ShmemEnd = (char *) ShmemBase + shmhdr->totalsize;
+ ShmemSegHdr = seghdr;
+ ShmemBase = seghdr;
+ ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
}
/*
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 3065ff5be71..7a07c5807ac 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -69,7 +69,7 @@ extern PGDLLIMPORT unsigned long UsedShmemSegID;
extern PGDLLIMPORT HANDLE UsedShmemSegID;
extern PGDLLIMPORT void *ShmemProtectiveRegion;
#endif
-extern PGDLLIMPORT void *UsedShmemSegAddr;
+extern PGDLLIMPORT PGShmemHeader *UsedShmemSegAddr;
#if !defined(WIN32) && !defined(EXEC_BACKEND)
#define DEFAULT_SHARED_MEMORY_TYPE SHMEM_TYPE_MMAP
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 842989111c3..8cdbe7a89c8 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -27,7 +27,8 @@
/* shmem.c */
extern PGDLLIMPORT slock_t *ShmemLock;
-extern void InitShmemAccess(void *seghdr);
+struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
+extern void InitShmemAccess(struct PGShmemHeader *seghdr);
extern void InitShmemAllocation(void);
extern void *ShmemAlloc(Size size);
extern void *ShmemAllocNoError(Size size);
--
2.47.0
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-19 12:57 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
@ 2024-11-19 13:29 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-21 07:55 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
2024-11-26 07:53 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
0 siblings, 2 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-19 13:29 UTC (permalink / raw)
To: Peter Eisentraut <peter@eisentraut.org>; +Cc: pgsql-hackers
> On Tue, Nov 19, 2024 at 01:57:00PM GMT, Peter Eisentraut wrote:
> On 18.10.24 21:21, Dmitry Dolgov wrote:
> > v1-0001-Allow-to-use-multiple-shared-memory-mappings.patch
> >
> > Preparation, introduces the possibility to work with many shmem mappings. To
> > make it less invasive, I've duplicated the shmem API to extend it with the
> > shmem_slot argument, while redirecting the original API to it. There are
> > probably better ways of doing that, I'm open for suggestions.
>
> After studying this a bit, I tend to think you should just change the
> existing APIs in place. So for example,
>
> void *ShmemAlloc(Size size);
>
> becomes
>
> void *ShmemAlloc(int shmem_slot, Size size);
>
> There aren't that many callers, and all these duplicated interfaces almost
> add more new code than they save.
>
> It might be worth making exceptions for interfaces that are likely to be
> used by extensions. For example, I see pg_stat_statements using
> ShmemInitStruct() and ShmemInitHash(). But that seems to be it. Are there
> any other examples out there? Maybe there are many more that I don't see
> right now. But at least for the initialization functions, it doesn't seem
> worth it to preserve the existing interfaces exactly.
>
> In any case, I think the slot number should be the first argument. This
> matches how MemoryContextAlloc() or also talloc() work.
Yeah, agree. I'll reshape this part, thanks.
> (Now here is an idea: Could these just be memory contexts? Instead of
> making six shared memory slots, could you make six memory contexts with a
> special shared memory type. And ShmemAlloc becomes the allocation function,
> etc.?)
Sound interesting. I don't know how good the memory context interface
would fit here, but I'll do some investigation.
> I noticed the existing code made inconsistent use of PGShmemHeader * vs.
> void *, which also bled into your patch. I made the attached little patch
> to clean that up a bit.
Right, it was bothering me the whole time, but not strong enough to make
me fix this in the PoC just yet.
> I suggest splitting the struct ShmemSegment into one struct for the three
> memory addresses and a separate array just for the slock_t's. The former
> struct can then stay private in storage/ipc/shmem.c, only the locks need to
> be exported.
>
> Maybe rename ANON_MAPPINGS to something like NUM_ANON_MAPPINGS.
>
> Also, maybe some of this should be declared in storage/shmem.h rather than
> in storage/pg_shmem.h. We have the existing ShmemLock in there, so it would
> be a bit confusing to have the per-segment locks elsewhere.
>
> [...]
>
> I'm wondering about this change:
>
> -#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
> +#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
>
> It looks like this would affect all mmap() calls, not only the one you're
> changing. But that's the only one that uses this macro! I don't understand
> why we need this; I don't see anything in the commit log about this ever
> being used for any portability. I think we should just get rid of it and
> have mmap() use the right flags directly.
>
> I see that FreeBSD has a memfd_create() function. Might be worth a try.
> Obviously, this whole thing needs a configure test for memfd_create()
> anyway.
Yep, those points make sense to me.
> > v1-0005-Use-anonymous-files-to-back-shared-memory-segment.patch
> >
> > Allows an anonyous file to back a shared mapping. This makes certain things
> > easier, e.g. mappings visual representation, and gives an fd for possible
> > future customizations.
>
> I think this could be a useful patch just by itself, without the rest of the
> series, because of
>
> > * By default, Linux will not add file-backed shared mappings into a
> > core dump, making it more convenient to work with them in PostgreSQL:
> > no more huge dumps to process.
>
> This could be significant operational benefit.
>
> When you say "by default", is this adjustable? Does someone actually want
> the whole shared memory in their core file? (If it's adjustable, is it also
> adjustable for anonymous mappings?)
Yes, there is /proc/<pid>/coredump_filter [1], that allows to specify
what to include. One can ask to exclude anon, file-backed and hugetlb
shared memory, with the only caveat that it's per process. I guess
normally no one wants to have a full shared memory in the coredump, but
there could be exceptions.
> I see that memfd_create() has a MFD_HUGETLB flag. It's not very clear how
> that interacts with the MAP_HUGETLB flag for mmap(). Do you need to specify
> both of them if you want huge pages?
Correct, both (one flag in memfd_create and one for mmap) are needed to
use huge pages.
[1]: https://www.kernel.org/doc/html/latest/filesystems/proc.html#proc-pid-coredump-filter-core-dump-filt...
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-19 12:57 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
2024-11-19 13:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-21 07:55 ` Peter Eisentraut <peter@eisentraut.org>
2025-04-17 15:54 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Peter Eisentraut @ 2024-11-21 07:55 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers
On 19.11.24 14:29, Dmitry Dolgov wrote:
>> I see that memfd_create() has a MFD_HUGETLB flag. It's not very clear how
>> that interacts with the MAP_HUGETLB flag for mmap(). Do you need to specify
>> both of them if you want huge pages?
> Correct, both (one flag in memfd_create and one for mmap) are needed to
> use huge pages.
I was worried because the FreeBSD man page says
MFD_HUGETLB This flag is currently unsupported.
It looks like FreeBSD doesn't have MAP_HUGETLB, so maybe this is irrelevant.
But you should make sure in your patch that the right set of flags for
huge pages is passed.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-19 12:57 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
2024-11-19 13:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-21 07:55 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
@ 2025-04-17 15:54 ` Thomas Munro <thomas.munro@gmail.com>
2025-04-18 01:27 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Thomas Munro @ 2025-04-17 15:54 UTC (permalink / raw)
To: Peter Eisentraut <peter@eisentraut.org>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
On Thu, Nov 21, 2024 at 8:55 PM Peter Eisentraut <peter@eisentraut.org> wrote:
> On 19.11.24 14:29, Dmitry Dolgov wrote:
> >> I see that memfd_create() has a MFD_HUGETLB flag. It's not very clear how
> >> that interacts with the MAP_HUGETLB flag for mmap(). Do you need to specify
> >> both of them if you want huge pages?
> > Correct, both (one flag in memfd_create and one for mmap) are needed to
> > use huge pages.
>
> I was worried because the FreeBSD man page says
>
> MFD_HUGETLB This flag is currently unsupported.
>
> It looks like FreeBSD doesn't have MAP_HUGETLB, so maybe this is irrelevant.
>
> But you should make sure in your patch that the right set of flags for
> huge pages is passed.
MFD_HUGETLB does actually work on FreeBSD, but the man page doesn't
admit it (guessing an oversight, not sure, will see). And you don't
need the corresponding (non-existent) mmap flag. You also have to
specify a size eg MFD_HUGETLB | MFD_HUGE_2MB or you get ENOTSUPP, but
other than that quirk I see it definitely working with eg procstat -v.
That might be because FreeBSD doesn't have a default huge page size
concept? On Linux that's a boot time setting, I guess rarely changed.
I contemplated that once before, when I wrote a quick demo patch[1] to
implement huge_pages=on for FreeBSD (ie explicit rather than
transparent). I used a different function, not the Linuxoid one but
it's the same under the covers, and I wrote:
+ /*
+ * Find the matching page size index, or if huge_page_size wasn't set,
+ * then skip the smallest size and take the next one after that.
+ */
Swapping that topic back in, I was left wondering: (1) how to choose
between SHM_LARGEPAGE_ALLOC_DEFAULT, a policy that will cause
ftruncate() to try to defragment physical memory to fulfil your
request and can eat some serious CPU, and SHM_LARGEPAGE_ALLOC_NOWAIT,
and (2) if it's the second thing, well Linux is like that in respect
of failing fast, but for it to succeed you have to configure
nr_hugepages in the OS as a separate administrative step and *that's*
when it does any defragmentation required, and that's another concept
FreeBSD doesn't have. It's a bit of a weird concept too, I mean those
pages are not reserved for you in any way and anyone could nab them,
which is undeniably practical but it lacks a few qualities one might
hope for in a kernel facility... IDK. Anyway, the Linux-like
memfd_create() always does it the _DEFAULT way. EIther way, we can't
have identical "try" semantics: it'll actually put some effort into
trying, perhaps burning many seconds of CPU.
I took a peek at what we're doing for Windows and the man pages tell
me that it's like that too. I don't recall hearing any complaints
about that, but it's gated on a Windows permission that I assume very
few enabled, so "try" probably isn't trying for most systems.
Quoting:
"Large-page memory regions may be difficult to obtain after the system
has been running for a long time because the physical space for each
large page must be contiguous, but the memory may have become
fragmented. Allocating large pages under these conditions can
significantly affect system performance. Therefore, applications
should avoid making repeated large-page allocations and instead
allocate all large pages one time, at startup."
For Windows we also interpret "on" with GetLargePageMinimum(), which
sounds like my "second known page size" idea.
To make Windows do the thing that this thread wants, I found a thread
saying that calling VirtualAlloc(..., MEM_RESET) and then convincing
every process to call VirtualUnlock(...) might work:
https://groups.google.com/g/microsoft.public.win32.programmer.kernel/c/3SvznY38SSc/m/4Sx_xwon1vsJ
I'm not sure what to do about the other Unixen. One option is
nothing, no feature, patches welcome. Another is to use
shm_open(<made up name>), like DSM segments, except we never need to
reopen these ones so we could immediately call shm_unlink() to leave
only a very short window to crash and leak a name. It'd be low risk
name pollution in a name space that POSIX forgot to provide any way to
list. The other idea is non-standard madvise tricks but they seem
far too squishy to be part of a "portable" fallback if they even work
at all, so it might be better not to have the feature than that I
think.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-19 12:57 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
2024-11-19 13:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-21 07:55 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
2025-04-17 15:54 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
@ 2025-04-18 01:27 ` Thomas Munro <thomas.munro@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Thomas Munro @ 2025-04-18 01:27 UTC (permalink / raw)
To: Peter Eisentraut <peter@eisentraut.org>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
On Fri, Apr 18, 2025 at 3:54 AM Thomas Munro <thomas.munro@gmail.com> wrote:
> I contemplated that once before, when I wrote a quick demo patch[1] to
> implement huge_pages=on for FreeBSD (ie explicit rather than
> transparent). I used a different function, not the Linuxoid one but
Oops, I forgot to supply that link[1]. And by the way all that
technical mumbo jumbo about FreeBSD was just me writing up why I
didn't pull the trigger and add explicit huge_pages support for it.
The short version is: you shouldn't try to use that flag at all on
FreeBSD yet, as it's a separate research project to add that feature.
I care about PostgreSQL/FreeBSD personally and may consider that again
as I learn more about virtual memory topics, but actually its
transparent super pages seem to do a pretty decent job already and
people don't seem to want to turn them off.
For an actionable plan that should be portable everywhere, how about
this: use shm_open(<tempname>, O_CREAT | O_EXCL, S_IRUSR | S_IWUSR)
followed by shm_unlink(<tempname>) to make this work on every Unix
(FreeBSD could use its slightly better SHM_ANON as the name and skip
the unlink), and redirect to memfd inside #ifdef __linux__. One thing
to consider is that shm_open() descriptors are implicitly set to
FD_CLOEXEC per POSIX, so I think you need to clear that flag with
fcntl() in EXEC_BACKEND builds, and then also set it again in children
so that they don't pass the descriptor to subprograms they run with
system() etc. memfd_create() needs the same consideration, except its
default is the other way: I think you need to supply the MFD_CLOEXEC
flag explicitly, unless it's an EXEC_BACKEND build, and use the same
fnctl() to clear it in children if it is. To restate that the other
way around, in non-EXEC_BACKEND builds shm_open() already does the
right thing and memfd_create() needs MFD_CLOEXEC, with no extra steps
after that.
The only systems I'm aware of that *don't* have shm_open() are (1)
Android, but it's Linux so I assume it has memfd_create() (just for
fun: you can run PostgreSQL on a phone with termux[2], and you can see
that their package supplies a fake shm_open() that redirects to plain
open(); I guess didn't realise they could have supplied an ENOSYS
dummy and just set dynamic_shared_memory_type=mmap instead, and we'd
have done that for them!), and (2) the capability-based research OS
projects like Capsicum (and probably the others like it) that rip out
all the global namespace Unix APIs for approximately the same reason
as Android (PostgreSQL can't run under those yet, but just for fun: I
had PostgreSQL mostly working under Capsicum once, and noticed that
the problems to be solved had significant overlap with the
multithreading project: the global namespace stuff like signals/PIDs
and onymous IPC go away, and the only other major thing is absolute
paths, many of which are easily made relative to a pgdata fd and
handled with openat() in fd.c, but I digress...).
[1] https://www.postgresql.org/message-id/CA%2BhUKGLmBWHF6gusP55R7jVS1%3D6T%3DGphbZpUXiOgMMHDUkVCgw%40ma...
[2] https://github.com/termux/termux-packages/tree/master/packages/postgresql
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-19 12:57 ` Re: Changing shared_buffers without restart Peter Eisentraut <peter@eisentraut.org>
2024-11-19 13:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-26 07:53 ` Peter Eisentraut <peter@eisentraut.org>
1 sibling, 0 replies; 167+ messages in thread
From: Peter Eisentraut @ 2024-11-26 07:53 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers
On 19.11.24 14:29, Dmitry Dolgov wrote:
>> I noticed the existing code made inconsistent use of PGShmemHeader * vs.
>> void *, which also bled into your patch. I made the attached little patch
>> to clean that up a bit.
> Right, it was bothering me the whole time, but not strong enough to make
> me fix this in the PoC just yet.
I committed a bit of this, so check that when you're rebasing your patch
set.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-25 19:33 ` Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
4 siblings, 1 reply; 167+ messages in thread
From: Robert Haas @ 2024-11-25 19:33 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers
On Fri, Oct 18, 2024 at 3:21 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
> changing shared memory mapping layout. Any feedback is appreciated.
A lot of people would like to have this feature, so I hope this
proposal works out. Thanks for working on it.
I think the idea of having multiple shared memory segments is
interesting and makes sense, but I would prefer to see them called
"segments" rather than "slots" just as do we do for DSMs. The name
"slot" is somewhat overused, and invites confusion with replication
slots, inter alia. I think it's possible that having multiple fixed
shared memory segments will spell trouble on Windows, where we already
need to use a retry loop to try to get the main shared memory segment
mapped at the correct address. If there are multiple segments and we
need whatever ASLR stuff happens on Windows to not place anything else
overlapping with any of them, that means there's more chances for
stuff to fail than if we just need one address range to be free.
Granted, the individual ranges are smaller, so maybe it's fine? But I
don't know.
The big thing that worries me is synchronization, and while I've only
looked at the patch set briefly, it doesn't look to me as though
there's enough machinery here to make that work correctly. Suppose
that shared_buffers=8GB (a million buffers) and I change it to
shared_buffers=16GB (2 million buffers). As soon as any one backend
has seen that changed and expanded shared_buffers, there's a
possibility that some other backend which has not yet seen the change
might see a buffer number greater than a million. If it tries to use
that buffer number before it absorbs the change, something bad will
happen. The most obvious way for it to see such a buffer number - and
possibly the only one - is to do a lookup in the buffer mapping table
and find a buffer ID there that was inserted by some other backend
that has already seen the change.
Fixing this seems tricky. My understanding is that BufferGetBlock() is
extremely performance-critical, so having to do a bounds check there
to make sure that a given buffer number is in range would probably be
bad for performance. Also, even if the overhead weren't prohibitive, I
don't think we can safely stick code that unmaps and remaps shared
memory segments into a function that currently just does math, because
we've probably got places where we assume this operation can't fail --
as well as places where we assume that if we call BufferGetBlock(i)
and then BufferGetBlock(j), the second call won't change the answer to
the first.
It seems to me that it's probably only safe to swap out a backend's
notion of where shared_buffers is located when the backend holds on
buffer pins, and maybe not even all such places, because it would be a
problem if a backend looks up the address of a buffer before actually
pinning it, on the assumption that the answer can't change. I don't
know if that ever happens, but it would be a legal coding pattern
today. Doing it between statements seems safe as long as there are no
cursors holding pins. Doing it in the middle of a statement is
probably possible if we can verify that we're at a "safe" point in the
code, but I'm not sure exactly which points are safe. If we have no
code anywhere that assumes the address of an unpinned buffer can't
change before we pin it, then I guess the check for pins is the only
thing we need, but I don't know that to be the case.
I guess I would have imagined that a change like this would have to be
done in phases. In phase 1, we'd tell all of the backends that
shared_buffers had expanded to some new, larger value; but the new
buffers wouldn't be usable for anything yet. Then, once we confirmed
that everyone had the memo, we'd tell all the backends that those
buffers are now available for use. If shared_buffers were contracted,
phase 1 would tell all of the backends that shared_buffers had
contracted to some new, smaller value. Once a particular backend
learns about that, they will refuse to put any new pages into those
high-numbered buffers, but the existing contents would still be valid.
Once everyone has been told about this, we can go through and evict
all of those buffers, and then let everyone know that's done. Then
they shrink their mappings.
It looks to me like the patch doesn't expand the buffer mapping table,
which seems essential. But maybe I missed that.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
@ 2024-11-26 19:17 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-26 19:17 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: pgsql-hackers
> On Mon, Nov 25, 2024 at 02:33:48PM GMT, Robert Haas wrote:
>
> I think the idea of having multiple shared memory segments is
> interesting and makes sense, but I would prefer to see them called
> "segments" rather than "slots" just as do we do for DSMs. The name
> "slot" is somewhat overused, and invites confusion with replication
> slots, inter alia. I think it's possible that having multiple fixed
> shared memory segments will spell trouble on Windows, where we already
> need to use a retry loop to try to get the main shared memory segment
> mapped at the correct address. If there are multiple segments and we
> need whatever ASLR stuff happens on Windows to not place anything else
> overlapping with any of them, that means there's more chances for
> stuff to fail than if we just need one address range to be free.
> Granted, the individual ranges are smaller, so maybe it's fine? But I
> don't know.
I haven't had a chance to experiment with that on Windows, but I'm
hoping that in the worst case fallback to a single mapping via proposed
infrastructure (and the consequent limitations) would be acceptable.
> The big thing that worries me is synchronization, and while I've only
> looked at the patch set briefly, it doesn't look to me as though
> there's enough machinery here to make that work correctly. Suppose
> that shared_buffers=8GB (a million buffers) and I change it to
> shared_buffers=16GB (2 million buffers). As soon as any one backend
> has seen that changed and expanded shared_buffers, there's a
> possibility that some other backend which has not yet seen the change
> might see a buffer number greater than a million. If it tries to use
> that buffer number before it absorbs the change, something bad will
> happen. The most obvious way for it to see such a buffer number - and
> possibly the only one - is to do a lookup in the buffer mapping table
> and find a buffer ID there that was inserted by some other backend
> that has already seen the change.
Right, I haven't put much efforts into synchronization yet. It's in my
bucket list for the next iteration of the patch.
> code, but I'm not sure exactly which points are safe. If we have no
> code anywhere that assumes the address of an unpinned buffer can't
> change before we pin it, then I guess the check for pins is the only
> thing we need, but I don't know that to be the case.
Probably I'm missing something here. What scenario do you have in mind,
when the address of a buffer is changing?
> I guess I would have imagined that a change like this would have to be
> done in phases. In phase 1, we'd tell all of the backends that
> shared_buffers had expanded to some new, larger value; but the new
> buffers wouldn't be usable for anything yet. Then, once we confirmed
> that everyone had the memo, we'd tell all the backends that those
> buffers are now available for use. If shared_buffers were contracted,
> phase 1 would tell all of the backends that shared_buffers had
> contracted to some new, smaller value. Once a particular backend
> learns about that, they will refuse to put any new pages into those
> high-numbered buffers, but the existing contents would still be valid.
> Once everyone has been told about this, we can go through and evict
> all of those buffers, and then let everyone know that's done. Then
> they shrink their mappings.
Yep, sounds good. I was pondering about more crude approach, but doing
this in phases seems to be a way to go.
> It looks to me like the patch doesn't expand the buffer mapping table,
> which seems essential. But maybe I missed that.
Do you mean the "Shared Buffer Lookup Table"? It does expand it, but
under somewhat unfitting name STRATEGY_SHMEM_SLOT. But now that I look
at the code, I see a few issues around that -- so I would have to
improve it anyway, thanks for pointing that out.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-27 15:20 ` Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Robert Haas @ 2024-11-27 15:20 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers
On Tue, Nov 26, 2024 at 2:18 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> I haven't had a chance to experiment with that on Windows, but I'm
> hoping that in the worst case fallback to a single mapping via proposed
> infrastructure (and the consequent limitations) would be acceptable.
Yeah, if you can still fall back to a single mapping, I think that's
OK. It would be nicer if it could work on every platform in the same
way, but half a loaf is better than none.
> > code, but I'm not sure exactly which points are safe. If we have no
> > code anywhere that assumes the address of an unpinned buffer can't
> > change before we pin it, then I guess the check for pins is the only
> > thing we need, but I don't know that to be the case.
>
> Probably I'm missing something here. What scenario do you have in mind,
> when the address of a buffer is changing?
I was assuming that if you expand the mapping for shared_buffers, you
can't count on the new mapping being at the same address as the old
mapping. If you can, that makes things simpler, but what if the OS has
mapped something else just afterward, in the address space that you're
hoping to use when you expand the mapping?
> > It looks to me like the patch doesn't expand the buffer mapping table,
> > which seems essential. But maybe I missed that.
>
> Do you mean the "Shared Buffer Lookup Table"? It does expand it, but
> under somewhat unfitting name STRATEGY_SHMEM_SLOT. But now that I look
> at the code, I see a few issues around that -- so I would have to
> improve it anyway, thanks for pointing that out.
Yeah, we -- or at least I -- usually call that the buffer mapping
table. There are identifiers like BufMappingPartitionLock, for
example. I'm slightly surprised that the ShmemInitHash() call uses
something else as the identifier, but I guess that's how it is.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
@ 2024-11-27 20:48 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-27 20:48 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: pgsql-hackers
> On Wed, Nov 27, 2024 at 10:20:27AM GMT, Robert Haas wrote:
> > >
> > > code, but I'm not sure exactly which points are safe. If we have no
> > > code anywhere that assumes the address of an unpinned buffer can't
> > > change before we pin it, then I guess the check for pins is the only
> > > thing we need, but I don't know that to be the case.
> >
> > Probably I'm missing something here. What scenario do you have in mind,
> > when the address of a buffer is changing?
>
> I was assuming that if you expand the mapping for shared_buffers, you
> can't count on the new mapping being at the same address as the old
> mapping. If you can, that makes things simpler, but what if the OS has
> mapped something else just afterward, in the address space that you're
> hoping to use when you expand the mapping?
Yes, that's the whole point of the exercise with remap -- to keep
addresses unchanged, making buffer management simpler and allowing
resize mappings quicker. The trade off is that we would need to take
care of shared mapping placing.
My understanding is that clashing of mappings (either at creation time
or when resizing) could happen only withing the process address space,
and the assumption is that by the time we prepare the mapping layout all
the rest of mappings for the process are already done. But I agree, it's
an interesting question -- I'm going to investigate if those assumptions
could be wrong under certain conditions. Currently if something else is
mapped at the same address where we want to expand the mapping, we will
get an error and can decide how to proceed (e.g. if it happens at
creation time, proceed with a single mapping, otherwise ignore mapping
resize).
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-27 21:05 ` Robert Haas <robertmhaas@gmail.com>
2024-11-27 21:28 ` Re: Changing shared_buffers without restart Jelte Fennema-Nio <postgres@jeltef.nl>
2024-11-27 21:41 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 3 replies; 167+ messages in thread
From: Robert Haas @ 2024-11-27 21:05 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers
On Wed, Nov 27, 2024 at 3:48 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> My understanding is that clashing of mappings (either at creation time
> or when resizing) could happen only withing the process address space,
> and the assumption is that by the time we prepare the mapping layout all
> the rest of mappings for the process are already done.
I don't think that's correct at all. First, the user could type LOAD
'whatever' at any time. But second, even if they don't or you prohibit
them from doing so, the process could allocate memory for any of a
million different things, and that could require mapping a new region
of memory, and the OS could choose to place that just after an
existing mapping, or at least close enough that we can't expand the
object size as much as desired.
If we had an upper bound on the size of shared_buffers and could
reserve that amount of address space at startup time but only actually
map a portion of it, then we could later remap and expand into the
reserved space. Without that, I think there's absolutely no guarantee
that the amount of address space that we need is available when we
want to extend a mapping.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
@ 2024-11-27 21:28 ` Jelte Fennema-Nio <postgres@jeltef.nl>
2024-11-28 01:26 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2 siblings, 1 reply; 167+ messages in thread
From: Jelte Fennema-Nio @ 2024-11-27 21:28 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
On Wed, 27 Nov 2024 at 22:06, Robert Haas <robertmhaas@gmail.com> wrote:
> If we had an upper bound on the size of shared_buffers
I think a fairly reliable upper bound is the amount of physical memory
on the system at time of postmaster start. We could make it a GUC to
set the upper bound for the rare cases where people do stuff like
adding swap space later or doing online VM growth. We could even have
the default be something like 4x the physical memory to accommodate
those people by default.
> reserve that amount of address space at startup time but only actually
> map a portion of it
Or is this the difficult part?
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 21:28 ` Re: Changing shared_buffers without restart Jelte Fennema-Nio <postgres@jeltef.nl>
@ 2024-11-28 01:26 ` Robert Haas <robertmhaas@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Robert Haas @ 2024-11-28 01:26 UTC (permalink / raw)
To: Jelte Fennema-Nio <postgres@jeltef.nl>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
On Wed, Nov 27, 2024 at 4:28 PM Jelte Fennema-Nio <postgres@jeltef.nl> wrote:
> On Wed, 27 Nov 2024 at 22:06, Robert Haas <robertmhaas@gmail.com> wrote:
> > If we had an upper bound on the size of shared_buffers
>
> I think a fairly reliable upper bound is the amount of physical memory
> on the system at time of postmaster start. We could make it a GUC to
> set the upper bound for the rare cases where people do stuff like
> adding swap space later or doing online VM growth. We could even have
> the default be something like 4x the physical memory to accommodate
> those people by default.
Yes, Peter mentioned similar ideas on this thread last week.
> > reserve that amount of address space at startup time but only actually
> > map a portion of it
>
> Or is this the difficult part?
I'm not sure how difficult this is, although I'm pretty sure that it's
more difficult than adding a GUC. My point wasn't so much whether this
is easy or hard but rather that it's essential if you want to avoid
having addresses change when the resizing happens.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
@ 2024-11-27 21:41 ` Andres Freund <andres@anarazel.de>
2024-11-28 01:28 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2 siblings, 1 reply; 167+ messages in thread
From: Andres Freund @ 2024-11-27 21:41 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
Hi,
On 2024-11-27 16:05:47 -0500, Robert Haas wrote:
> On Wed, Nov 27, 2024 at 3:48 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > My understanding is that clashing of mappings (either at creation time
> > or when resizing) could happen only withing the process address space,
> > and the assumption is that by the time we prepare the mapping layout all
> > the rest of mappings for the process are already done.
>
> I don't think that's correct at all. First, the user could type LOAD
> 'whatever' at any time. But second, even if they don't or you prohibit
> them from doing so, the process could allocate memory for any of a
> million different things, and that could require mapping a new region
> of memory, and the OS could choose to place that just after an
> existing mapping, or at least close enough that we can't expand the
> object size as much as desired.
>
> If we had an upper bound on the size of shared_buffers and could
> reserve that amount of address space at startup time but only actually
> map a portion of it, then we could later remap and expand into the
> reserved space. Without that, I think there's absolutely no guarantee
> that the amount of address space that we need is available when we
> want to extend a mapping.
Strictly speaking we don't actually need to map shared buffers to the same
location in each process... We do need that for most other uses of shared
memory, including the buffer mapping table, but not for the buffer data
itself.
Whether it's worth the complexity of dealing with differing locations is
another matter.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 21:41 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2024-11-28 01:28 ` Robert Haas <robertmhaas@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Robert Haas @ 2024-11-28 01:28 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
On Wed, Nov 27, 2024 at 4:41 PM Andres Freund <andres@anarazel.de> wrote:
> Strictly speaking we don't actually need to map shared buffers to the same
> location in each process... We do need that for most other uses of shared
> memory, including the buffer mapping table, but not for the buffer data
> itself.
Well, if it can move, then you have to make sure it doesn't move while
someone's holding onto a pointer into it. I'm not exactly sure how
hard it is to guarantee that, but we certainly do construct pointers
into shared_buffers and use them at least for short periods of time,
so it's not a purely academic concern.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
@ 2024-11-28 16:30 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-29 18:17 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2 siblings, 2 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-28 16:30 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: pgsql-hackers
> On Wed, Nov 27, 2024 at 04:05:47PM GMT, Robert Haas wrote:
> On Wed, Nov 27, 2024 at 3:48 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > My understanding is that clashing of mappings (either at creation time
> > or when resizing) could happen only withing the process address space,
> > and the assumption is that by the time we prepare the mapping layout all
> > the rest of mappings for the process are already done.
>
> I don't think that's correct at all. First, the user could type LOAD
> 'whatever' at any time. But second, even if they don't or you prohibit
> them from doing so, the process could allocate memory for any of a
> million different things, and that could require mapping a new region
> of memory, and the OS could choose to place that just after an
> existing mapping, or at least close enough that we can't expand the
> object size as much as desired.
>
> If we had an upper bound on the size of shared_buffers and could
> reserve that amount of address space at startup time but only actually
> map a portion of it, then we could later remap and expand into the
> reserved space. Without that, I think there's absolutely no guarantee
> that the amount of address space that we need is available when we
> want to extend a mapping.
Just done a couple of experiments, and I think this could be addressed by
careful placing of mappings as well, based on two assumptions: for a new
mapping the kernel always picks up a lowest address that allows enough space,
and the maximum amount of allocable memory for other mappings could be derived
from total available memory. With that in mind the shared mapping layout will
have to have a large gap at the start, between the lowest address and the
shared mappings used for buffers and rest -- the gap where all the other
mapping (allocations, libraries, madvise, etc) will land. It's similar to
address space reserving you mentioned above, will reduce possibility of
clashing significantly, and looks something like this:
01339000-0139e000 [heap]
0139e000-014aa000 [heap]
7f2dd72f6000-7f2dfbc9c000 /memfd:strategy (deleted)
7f2e0209c000-7f2e269b0000 /memfd:checkpoint (deleted)
7f2e2cdb0000-7f2e516b4000 /memfd:iocv (deleted)
7f2e57ab4000-7f2e7c478000 /memfd:descriptors (deleted)
7f2ebc478000-7f2ee8d3c000 /memfd:buffers (deleted)
^ note the distance between two mappings,
which is intended for resize
7f3168d3c000-7f318d600000 /memfd:main (deleted)
^ here is where the gap starts
7f4194c00000-7f4194e7d000
^ this one is an anonymous maping created due to large
memory allocation after shared mappings were created
7f4195000000-7f419527d000
7f41952dc000-7f4195416000
7f4195416000-7f4195600000 /dev/shm/PostgreSQL.2529797530
7f4195600000-7f41a311d000 /usr/lib/locale/locale-archive
7f41a317f000-7f41a3200000
7f41a3200000-7f41a3201000 /usr/lib64/libicudata.so.74.2
The assumption about picking up a lowest address is just how it works right now
on Linux, this fact is already used in the patch. The idea that we could put
upper boundary on the size of other mappings based on total available memory
comes from the fact that anonymous mappings, that are much larger than memory,
will fail without overcommit. With overcommit it becomes different, but if
allocations are hitting that limit I can imagine there are bigger problems than
shared buffer resize.
This approach follows the same ideas already used in the patch, and have the
same trade offs: no address changes, but questions about portability.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-28 17:18 ` Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:45 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 2 replies; 167+ messages in thread
From: Robert Haas @ 2024-11-28 17:18 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers
On Thu, Nov 28, 2024 at 11:30 AM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> on Linux, this fact is already used in the patch. The idea that we could put
> upper boundary on the size of other mappings based on total available memory
> comes from the fact that anonymous mappings, that are much larger than memory,
> will fail without overcommit. With overcommit it becomes different, but if
> allocations are hitting that limit I can imagine there are bigger problems than
> shared buffer resize.
>
> This approach follows the same ideas already used in the patch, and have the
> same trade offs: no address changes, but questions about portability.
I definitely welcome the fact that you have some platform-specific
knowledge of the Linux behavior, because that's expertise that is
obviously quite useful here and which I lack. I'm personally not
overly concerned about whether it works on every other platform -- I
would prefer an implementation that works everywhere, but I'd rather
have one that works on Linux than have nothing. It's unclear to me why
operating systems don't offer better primitives for this sort of thing
-- in theory there could be a system call that sets aside a pool of
address space and then other system calls that let you allocate
shared/unshared memory within that space or even at specific
addresses, but actually such things don't exist.
All that having been said, what does concern me a bit is our ability
to predict what Linux will do well enough to keep what we're doing
safe; and also whether the Linux behavior might abruptly change in the
future. Users would be sad if we released this feature and then a
future kernel upgrade causes PostgreSQL to completely stop working. I
don't know how the Linux kernel developers actually feel about this
sort of thing, but if I imagine myself as a kernel developer, I can
totally see myself saying "well, we never promised that this would
work in any particular way, so we're free to change it whenever we
like." We've certainly used that argument here countless times.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
@ 2024-11-28 18:13 ` Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
1 sibling, 1 reply; 167+ messages in thread
From: Matthias van de Meent @ 2024-11-28 18:13 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
On Thu, 28 Nov 2024 at 18:19, Robert Haas <robertmhaas@gmail.com> wrote:
>
> [...] It's unclear to me why
> operating systems don't offer better primitives for this sort of thing
> -- in theory there could be a system call that sets aside a pool of
> address space and then other system calls that let you allocate
> shared/unshared memory within that space or even at specific
> addresses, but actually such things don't exist.
Isn't that more a stdlib/malloc issue? AFAIK, Linux's mmap(2) syscall
allows you to request memory from the OS at arbitrary addresses - it's
just that stdlib's malloc doens't expose the 'alloc at this address'
part of that API.
Windows seems to have an equivalent API in VirtualAlloc*. Both the
Windows API and Linux's mmap have an optional address argument, which
(when not NULL) is where the allocation will be placed (some
conditions apply, based on flags and specific API used), so, assuming
we have some control on where to allocate memory, we should be able to
reserve enough memory by using these APIs.
Kind regards,
Matthias van de Meent
Neon (https://neon.tech)
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
@ 2024-11-28 18:57 ` Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 00:56 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Tom Lane @ 2024-11-28 18:57 UTC (permalink / raw)
To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: Robert Haas <robertmhaas@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
Matthias van de Meent <boekewurm+postgres@gmail.com> writes:
> On Thu, 28 Nov 2024 at 18:19, Robert Haas <robertmhaas@gmail.com> wrote:
>> [...] It's unclear to me why
>> operating systems don't offer better primitives for this sort of thing
>> -- in theory there could be a system call that sets aside a pool of
>> address space and then other system calls that let you allocate
>> shared/unshared memory within that space or even at specific
>> addresses, but actually such things don't exist.
> Isn't that more a stdlib/malloc issue? AFAIK, Linux's mmap(2) syscall
> allows you to request memory from the OS at arbitrary addresses - it's
> just that stdlib's malloc doens't expose the 'alloc at this address'
> part of that API.
I think what Robert is concerned about is that there is exactly 0
guarantee that that will succeed, because you have no control over
system-driven allocations of address space (for example, loading
of extensions or JIT code). In fact, given things like ASLR, there
is pressure on the kernel crew to make that *less* predictable not
more so. So even if we devise a method that seems to work reliably
today, we could have little faith that it would work with next year's
kernels.
regards, tom lane
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
@ 2024-11-29 00:56 ` Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-29 01:42 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 16:47 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Matthias van de Meent @ 2024-11-29 00:56 UTC (permalink / raw)
To: Tom Lane <tgl@sss.pgh.pa.us>; +Cc: Robert Haas <robertmhaas@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
On Thu, 28 Nov 2024 at 19:57, Tom Lane <tgl@sss.pgh.pa.us> wrote:
>
> Matthias van de Meent <boekewurm+postgres@gmail.com> writes:
> > On Thu, 28 Nov 2024 at 18:19, Robert Haas <robertmhaas@gmail.com> wrote:
> >> [...] It's unclear to me why
> >> operating systems don't offer better primitives for this sort of thing
> >> -- in theory there could be a system call that sets aside a pool of
> >> address space and then other system calls that let you allocate
> >> shared/unshared memory within that space or even at specific
> >> addresses, but actually such things don't exist.
>
> > Isn't that more a stdlib/malloc issue? AFAIK, Linux's mmap(2) syscall
> > allows you to request memory from the OS at arbitrary addresses - it's
> > just that stdlib's malloc doens't expose the 'alloc at this address'
> > part of that API.
>
> I think what Robert is concerned about is that there is exactly 0
> guarantee that that will succeed, because you have no control over
> system-driven allocations of address space (for example, loading
> of extensions or JIT code). In fact, given things like ASLR, there
> is pressure on the kernel crew to make that *less* predictable not
> more so.
I see what you mean, but I think that shouldn't be much of an issue.
I'm not a kernel hacker, but I've never heard about anyone arguing to
remove mmap's mapping-overwriting behavior for user-controlled
mappings - it seems too useful as a way to guarantee relative memory
addresses (agreed, there is now mseal(2), but that is the user asking
for security on their own mapping, this isn't applied to arbitrary
mappings).
I mean, we can do the following to get a nice contiguous empty address
space no other mmap(NULL)s will get put into:
/* reserve size bytes of memory */
base = mmap(NULL, size, PROT_NONE, ...flags, ...);
/* use the first small_size bytes of that reservation */
allocated_in_reserved = mmap(base, small_size, PROT_READ |
PROT_WRITE, MAP_FIXED, ...);
With the PROT_NONE protection option the OS doesn't actually allocate
any backing memory, but guarantees no other mmap(NULL, ...) will get
placed in that area such that it overlaps with that allocation until
the area is munmap-ed, thus allowing us to reserve a chunk of address
space without actually using (much) memory. Deallocations have to go
through mmap(... PROT_NONE, ...) instead of munmap if we'd want to
keep the full area reserved, but I think that's not that much of an
issue.
I also highly doubt Linux will remove or otherwise limit the PROT_NONE
option to such a degree that we won't be able to "balloon" the memory
address space for (e.g.) dynamic shared buffer resizing.
See also: FreeBSD's MAP_GUARD mmap flag, Window's MEM_RESERVE and
MEM_RESERVE_PLACEHOLDER flags for VirtualAlloc[2][Ex].
See also [0] where PROT_NONE is explicitly called out as a tool for
reserving memory address space.
> So even if we devise a method that seems to work reliably
> today, we could have little faith that it would work with next year's
> kernels.
I really don't think that userspace memory address space reservations
through e.g. PROT_NONE or MEM_RESERVE[_PLACEHOLDER] will be retired
anytime soon, at least not without the relevant kernels also providing
effective alternatives.
Kind regards,
Matthias van de Meent
Neon (https://neon.tech)
[0] https://www.gnu.org/software/libc/manual/html_node/Memory-Protection.html
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 00:56 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
@ 2024-11-29 01:42 ` Tom Lane <tgl@sss.pgh.pa.us>
1 sibling, 0 replies; 167+ messages in thread
From: Tom Lane @ 2024-11-29 01:42 UTC (permalink / raw)
To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: Robert Haas <robertmhaas@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers
Matthias van de Meent <boekewurm+postgres@gmail.com> writes:
> I mean, we can do the following to get a nice contiguous empty address
> space no other mmap(NULL)s will get put into:
> /* reserve size bytes of memory */
> base = mmap(NULL, size, PROT_NONE, ...flags, ...);
> /* use the first small_size bytes of that reservation */
> allocated_in_reserved = mmap(base, small_size, PROT_READ |
> PROT_WRITE, MAP_FIXED, ...);
> With the PROT_NONE protection option the OS doesn't actually allocate
> any backing memory, but guarantees no other mmap(NULL, ...) will get
> placed in that area such that it overlaps with that allocation until
> the area is munmap-ed, thus allowing us to reserve a chunk of address
> space without actually using (much) memory.
Well, that's all great if it works portably. But I don't see one word
in either POSIX or the Linux mmap(2) man page that promises those
semantics for PROT_NONE. I also wonder how well a giant chunk of
"unbacked" address space will interoperate with the OOM killer,
top(1)'s display of used memory, and other things that have caused us
headaches with large shared-memory arenas.
Maybe those issues are all in the past and this'll work great.
I'm not holding my breath though.
regards, tom lane
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 00:56 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
@ 2024-11-29 16:47 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-02 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-29 16:47 UTC (permalink / raw)
To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Robert Haas <robertmhaas@gmail.com>; pgsql-hackers
> On Fri, Nov 29, 2024 at 01:56:30AM GMT, Matthias van de Meent wrote:
>
> I mean, we can do the following to get a nice contiguous empty address
> space no other mmap(NULL)s will get put into:
>
> /* reserve size bytes of memory */
> base = mmap(NULL, size, PROT_NONE, ...flags, ...);
> /* use the first small_size bytes of that reservation */
> allocated_in_reserved = mmap(base, small_size, PROT_READ |
> PROT_WRITE, MAP_FIXED, ...);
>
> With the PROT_NONE protection option the OS doesn't actually allocate
> any backing memory, but guarantees no other mmap(NULL, ...) will get
> placed in that area such that it overlaps with that allocation until
> the area is munmap-ed, thus allowing us to reserve a chunk of address
> space without actually using (much) memory.
From what I understand it's not much different from the scenario when we
just map as much as we want in advance. The actual memory will not be
allocated in both cases due to CoW, oom_score seems to be the same. I
agree it sounds attractive, but after some experimenting it looks like
it won't work with huge pages insige a cgroup v2 (=container).
The reason is Linux has recently learned to apply memory reservation
limits on hugetlb inside a cgroup, which are applied to mmap. Nowadays
this feature is often configured out of the box in various container
orchestrators, meaning that a scenario "set hugetlb=1GB on a container,
reserve 32GB with PROT_NONE" will fail. I've also tried to mix and
match, reserve some address space via non-hugetlb mapping, and allocate
a hugetlb out of it, but it doesn't work either (the smaller mmap
complains about MAP_HUGETLB with EINVAL).
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 00:56 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-29 16:47 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-12-02 19:17 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-03 14:31 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2024-12-02 19:17 UTC (permalink / raw)
To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Robert Haas <robertmhaas@gmail.com>; pgsql-hackers
> On Fri, Nov 29, 2024 at 05:47:27PM GMT, Dmitry Dolgov wrote:
> > On Fri, Nov 29, 2024 at 01:56:30AM GMT, Matthias van de Meent wrote:
> >
> > I mean, we can do the following to get a nice contiguous empty address
> > space no other mmap(NULL)s will get put into:
> >
> > /* reserve size bytes of memory */
> > base = mmap(NULL, size, PROT_NONE, ...flags, ...);
> > /* use the first small_size bytes of that reservation */
> > allocated_in_reserved = mmap(base, small_size, PROT_READ |
> > PROT_WRITE, MAP_FIXED, ...);
> >
> > With the PROT_NONE protection option the OS doesn't actually allocate
> > any backing memory, but guarantees no other mmap(NULL, ...) will get
> > placed in that area such that it overlaps with that allocation until
> > the area is munmap-ed, thus allowing us to reserve a chunk of address
> > space without actually using (much) memory.
>
> From what I understand it's not much different from the scenario when we
> just map as much as we want in advance. The actual memory will not be
> allocated in both cases due to CoW, oom_score seems to be the same. I
> agree it sounds attractive, but after some experimenting it looks like
> it won't work with huge pages insige a cgroup v2 (=container).
>
> The reason is Linux has recently learned to apply memory reservation
> limits on hugetlb inside a cgroup, which are applied to mmap. Nowadays
> this feature is often configured out of the box in various container
> orchestrators, meaning that a scenario "set hugetlb=1GB on a container,
> reserve 32GB with PROT_NONE" will fail. I've also tried to mix and
> match, reserve some address space via non-hugetlb mapping, and allocate
> a hugetlb out of it, but it doesn't work either (the smaller mmap
> complains about MAP_HUGETLB with EINVAL).
I've asked about that in linux-mm [1]. To my surprise, the
recommendations were to stick to creating a large mapping in advance,
and slice smaller mappings out of that, which could be resized later.
The OOM score should not be affected, and hugetlb could be avoided using
MAP_NORESERVE flag for the initial mapping (I've experimented with that,
seems to be working just fine, even if the slices are not using
MAP_NORESERVE).
I guess that would mean I'll try to experiment with this approach as
well. But what others think? How much research do we need to do, to gain
some confidence about large shared mappings and make it realistically
acceptable?
[1]: https://lore.kernel.org/linux-mm/pr7zggtdgjqjwyrfqzusih2suofszxvlfxdptbo2smneixkp7i@nrmtbhemy3is/t/
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 00:56 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-29 16:47 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-02 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-12-03 14:31 ` Robert Haas <robertmhaas@gmail.com>
2024-12-17 14:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Robert Haas @ 2024-12-03 14:31 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; Tom Lane <tgl@sss.pgh.pa.us>; pgsql-hackers
On Mon, Dec 2, 2024 at 2:18 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> I've asked about that in linux-mm [1]. To my surprise, the
> recommendations were to stick to creating a large mapping in advance,
> and slice smaller mappings out of that, which could be resized later.
> The OOM score should not be affected, and hugetlb could be avoided using
> MAP_NORESERVE flag for the initial mapping (I've experimented with that,
> seems to be working just fine, even if the slices are not using
> MAP_NORESERVE).
>
> I guess that would mean I'll try to experiment with this approach as
> well. But what others think? How much research do we need to do, to gain
> some confidence about large shared mappings and make it realistically
> acceptable?
Personally, I like this approach. It seems to me that this opens up
the possibility of a system where the virtual addresses of data
structures in shared memory never change, which I think will avoid an
absolutely massive amount of implementation complexity. It's obviously
not ideal that we have to specify in advance an upper limit on the
potential size of shared_buffers, but we can live with it. It's better
than what we have today; and certainly cloud providers will have no
issue with pre-setting that to a reasonable value. I don't know if we
can port it to other operating systems, but it seems at least possible
that they offer similar primitives, or will in the future; if not, we
can disable the feature on those platforms.
I still think the synchronization is going to be tricky. For example
when you go to shrink a mapping, you need to make sure that it's free
of buffers that anyone might touch; and when you grow a mapping, you
need to make sure that nobody tries to touch that address space before
they grow the mapping, which goes back to my earlier point about
someone doing a lookup into the buffer mapping table and finding a
buffer number that is beyond the end of what they've already mapped.
But I think it may be doable with sufficient cleverness.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 00:56 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-29 16:47 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-02 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-03 14:31 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
@ 2024-12-17 14:10 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-01-13 08:11 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2024-12-17 14:10 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; Tom Lane <tgl@sss.pgh.pa.us>; pgsql-hackers
On Tue, Dec 3, 2024 at 8:01 PM Robert Haas <robertmhaas@gmail.com> wrote:
>
> On Mon, Dec 2, 2024 at 2:18 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > I've asked about that in linux-mm [1]. To my surprise, the
> > recommendations were to stick to creating a large mapping in advance,
> > and slice smaller mappings out of that, which could be resized later.
> > The OOM score should not be affected, and hugetlb could be avoided using
> > MAP_NORESERVE flag for the initial mapping (I've experimented with that,
> > seems to be working just fine, even if the slices are not using
> > MAP_NORESERVE).
> >
> > I guess that would mean I'll try to experiment with this approach as
> > well. But what others think? How much research do we need to do, to gain
> > some confidence about large shared mappings and make it realistically
> > acceptable?
>
> Personally, I like this approach. It seems to me that this opens up
> the possibility of a system where the virtual addresses of data
> structures in shared memory never change, which I think will avoid an
> absolutely massive amount of implementation complexity. It's obviously
> not ideal that we have to specify in advance an upper limit on the
> potential size of shared_buffers, but we can live with it. It's better
> than what we have today; and certainly cloud providers will have no
> issue with pre-setting that to a reasonable value. I don't know if we
> can port it to other operating systems, but it seems at least possible
> that they offer similar primitives, or will in the future; if not, we
> can disable the feature on those platforms.
>
> I still think the synchronization is going to be tricky. For example
> when you go to shrink a mapping, you need to make sure that it's free
> of buffers that anyone might touch; and when you grow a mapping, you
> need to make sure that nobody tries to touch that address space before
> they grow the mapping, which goes back to my earlier point about
> someone doing a lookup into the buffer mapping table and finding a
> buffer number that is beyond the end of what they've already mapped.
> But I think it may be doable with sufficient cleverness.
>
From the discussion so far, the protocol for each shared memory slot
(or segment as suggested by Robert) seems to be the following.
1. At the start create a memory mapping using mmap with maximum
allocation (maxsize) with PROT_READ/PROT_WRITE and MAP_NORESERVE to
reserve address space. Assume this is created at virtual address
maddr.
2. Resize it to the required size (size) using mremap() - this will be
used to create shared memory objects
3. Map a segment with PROT_NONE and MAP_NORESERVE at maddr + size.
This segment would not allow any other mapping to be added in the
required space. PROT_NONE will protect from unintentional writes/reads
from this space.
4. When resizing the segment remove the mapping created in step 3 and
execute step 2 and 3 again. Synchronization, mentioned by Robert,
should be carried out somewhere in this step.
Note that the addresses need to be aligned as per mmap and mremap requirements.
Please correct me if I am wrong.
I wrote the attached simple program simulating this protocol. It seems
to work as expected. However, mmap'ing with MAP_FIXED would still be
able to dislodge the reserved memory. But that's true with any mapped
segment; not just with reserved memory.
A bit about the program: It reserves a 3MB memory segment and resizes
it to 1MB, 2MB and back to 3MB, thus exercising both shrinking and
enlarging the memory. It forks a child process after resizing the the
memory segment first time. At every step it makes sure that the parent
and child programs can write and read at the boundaries of the resized
memory segment. The program waits for getchar() at these steps. So in
case the program seems to be stuck, try pressing Enter once or twice.
I could verify the memory mappings, their sizes etc. by looking at
/proc/PID/maps and /proc/PID/status but I did not find a way to verify
the amount of memory actually allocated and verify that it's actually
shrinking and expanding. Please let me know how to verify that.
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-csrc] mmap_exp.c (7.4K, ../../CAExHW5uyuHc3SkQhb8P_SKjMMPwY_Jp3=NCDneUpQWQuvq_ZUA@mail.gmail.com/2-mmap_exp.c)
download | inline:
#define _GNU_SOURCE 1 /* See feature_test_macros(7) */
#include <errno.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mman.h>
#include <unistd.h>
#include <stdbool.h>
void
unmap_memory(void *memaddr, size_t size, const char *tag)
{
if (munmap(memaddr, size) < 0)
{
printf("%s: shared memory unmapping failed with eno %d\n", tag, errno);
exit(__LINE__);
}
printf("%s: unmapped memory of size %lu from %p.\n", tag, size, memaddr);
}
void
unmap_on_exit(void *memaddr, size_t size, int exit_code, const char *tag)
{
unmap_memory(memaddr, size, tag);
exit(exit_code);
}
void *
map_memory(void *addr, size_t size, bool noreserve, bool fixed, int protection)
{
void *memaddr;
int flags = MAP_SHARED|MAP_ANONYMOUS;
if (noreserve)
flags = flags | MAP_NORESERVE;
if (fixed)
flags = flags | MAP_FIXED_NOREPLACE;
memaddr = mmap(addr, size, protection, flags, -1, 0);
if (memaddr == MAP_FAILED )
{
printf("shared memory mapping at %p of size %lu with %s failed with eno %d\n",
addr, size, noreserve ? "no reservation" : "reservation", errno);
return NULL;
}
if (addr != NULL && memaddr != addr)
{
printf("Expected memory of size %lu with %s to be mapped at %p but got mapped at %p",
size, noreserve ? "no reservation" : "reservation", addr, memaddr);
}
printf("mapped memory of size %lu with %s at %p.\n",
size, noreserve ? "no reservation" : "reservation", memaddr);
return memaddr;
}
void
p_write_and_readwait(int *addr, int wsign, int readsign)
{
*addr = wsign;
printf("parent wrote value %d at %p\n", wsign, addr);
while (*addr != readsign)
{
printf("parent is sleeping for child to increment signature value to %d at %p\n", readsign, addr);
sleep(1);
}
printf("parent found signature value of %d at %p.\n", *addr, addr);
}
void
c_readwait_and_write(int *addr, int readsign, int wsign)
{
while (*addr != readsign)
{
printf("child is sleeping for parent to write signature value of %d at %p\n", readsign, addr);
sleep(1);
}
printf("child found signature value of %d at %p.\n", *addr, addr);
*addr = wsign;
printf("child wrote value %d at %p\n", wsign, addr);
}
/*
* Unmap reserved space, resize memory and add reserved space.
*/
void
resize_memory(void *addr, size_t oldsize, size_t newsize, size_t maxsize, int sign, const char *tag)
{
int *signaddr = addr;
/* Unmap existing unreserved memory first */
if (oldsize != maxsize)
unmap_memory(addr + oldsize, maxsize - oldsize, tag);
/* Resize memory and check sanity */
void *oldaddr = addr;
addr = mremap(addr, oldsize, newsize, 0);
if (addr == MAP_FAILED)
{
printf("%s: resizing memory at %p from %lu to %lu failed with errno %d\n",
tag, oldaddr, oldsize, newsize, errno);
return;
}
if (*signaddr != sign)
{
printf("%s: didn't find expected value %d at %p after resizing, Instead found %d\n", tag, sign, signaddr, *signaddr);
return;
}
if (addr != oldaddr)
{
printf("%s: remapped to %p instead of %p", tag, addr, oldaddr);
return;
}
if (newsize != maxsize)
map_memory(addr + newsize, maxsize - newsize, true, false, PROT_NONE);
printf("%s: resized memory at %p from %lu to %lu successfully retaining old value %d at %p. Press Enter\n",
tag, addr, oldsize, newsize, sign, signaddr);
getchar();
}
void
parent_process(void *memaddr, size_t *sizes, int *signs, int numsizes, size_t maxsize)
{
void *otheraddr;
printf("parent: *** checking consistency before resizing ***\n");
p_write_and_readwait(memaddr, signs[0], signs[0] + 1);
p_write_and_readwait(memaddr + sizes[0] - sizeof(int) - 1, signs[0] + 2, signs[0] + 3);
getchar();
resize_memory(memaddr, sizes[0], sizes[1], maxsize, signs[0] + 1, "parent");
printf("parent: *** checking consistency after resizing ***\n");
p_write_and_readwait(memaddr + sizes[0], signs[1], signs[1] + 2);
p_write_and_readwait(memaddr + sizes[1] - sizeof(int) - 1, signs[1] + 3, signs[1] + 4);
getchar();
/*
* Try adding a mapping between current boundary and max boundary. This
* should not succeed because of reserved space at the end.
*/
otheraddr = map_memory(memaddr + sizes[1], maxsize - sizes[1] - 1024, false, true, PROT_WRITE | PROT_READ);
if (otheraddr != NULL)
{
printf("Extra memory segment mapped in the reserved space from %p to %p.\n", memaddr, memaddr + maxsize);
unmap_on_exit(memaddr, maxsize, __LINE__, "child");
}
resize_memory(memaddr, sizes[1], sizes[2], maxsize, signs[0] + 1, "parent");
printf("parent: ***** checking consistency after 2nd resizing *****\n");
p_write_and_readwait(memaddr + sizes[1], signs[2], signs[2] + 5);
p_write_and_readwait(memaddr + sizes[2] - sizeof(int) - 1, signs[2] + 6, signs[2] + 7);
getchar();
}
void
child_process(void *memaddr, size_t *sizes, int *signs, int numsizes, size_t maxsize)
{
void *otheraddr;
printf("child: check memory mapping /proc/%d/maps and status /proc/%d/status\n", getpid(), getpid());
/* Read and write: at boundaries */
printf("child: *** checking consistency before resizing ***\n");
c_readwait_and_write(memaddr, signs[0], signs[0] + 1);
c_readwait_and_write(memaddr + sizes[0] - sizeof(int) - 1, signs[0] + 2, signs[0] + 3);
getchar();
resize_memory(memaddr, sizes[0], sizes[1], maxsize, signs[0] + 1, "child");
printf("child: *** checking consistency after resizing ***\n");
c_readwait_and_write(memaddr + sizes[0], signs[1], signs[1] + 2);
c_readwait_and_write(memaddr + sizes[1] - sizeof(int) - 1, signs[1] + 3, signs[1] + 4);
getchar();
/*
* Try adding a mapping between current boundary and max boundary. This
* should not succeed because of reserved space at the end.
*/
otheraddr = map_memory(memaddr + sizes[1], maxsize - sizes[1] - 1024, false, true, PROT_WRITE | PROT_READ);
if (otheraddr != NULL)
{
if (otheraddr >= memaddr && otheraddr <= memaddr + maxsize)
printf("Extra memory segment mapped in the reserved space from %p to %p.\n", memaddr, memaddr + maxsize);
unmap_on_exit(memaddr, maxsize, __LINE__, "child");
}
resize_memory(memaddr, sizes[1], sizes[2], maxsize, signs[0] + 1, "child");
printf("child: *** checking consistency after 2nd resizing ***\n");
c_readwait_and_write(memaddr + sizes[1], signs[2], signs[2] + 5);
c_readwait_and_write(memaddr + sizes[2] - sizeof(int) - 1, signs[2] + 6, signs[2] + 7);
getchar();
}
int
main(int argc, char **argv)
{
size_t sizes[] = {100 * 1024 * 1024, 200 * 1024 * 1024, 300 * 1024 * 1024};
int signs[] = {435, 643, 586};
int numsizes = sizeof(sizes)/sizeof(sizes[0]);
void *memaddr;
size_t maxsize = 0;
pid_t chldpid;
char *tag = "parent";
#define FIRST_SIGN 100
if (numsizes != sizeof(signs)/sizeof(signs[0]))
printf("mismatch in number of sizes and number of signs %d vs %ld", numsizes, sizeof(signs)/sizeof(signs[0]));
for (int i = 0; i < numsizes; i++)
{
if (maxsize < sizes[i])
maxsize = sizes[i];
}
printf("parent: check memory mapping /proc/%d/maps and status /proc/%d/status\n", getpid(), getpid());
/* Reserve memory but don't allocate */
memaddr = map_memory(NULL, maxsize, true, false, PROT_WRITE | PROT_READ);
if (memaddr == NULL)
exit(1);
*(int *)memaddr = FIRST_SIGN;
getchar();
resize_memory(memaddr, maxsize, sizes[0], maxsize, FIRST_SIGN, "parent");
chldpid = fork();
if (chldpid < 0)
{
printf("forking a child failed\n");
unmap_on_exit(memaddr, maxsize, __LINE__, tag);
}
else if (chldpid == 0)
{
child_process(memaddr, sizes, signs, numsizes, maxsize);
tag = "child";
}
else
{
parent_process(memaddr, sizes, signs, numsizes, maxsize);
}
unmap_on_exit(memaddr, maxsize, __LINE__, tag);
}
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Re: Changing shared_buffers without restart Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 00:56 ` Re: Changing shared_buffers without restart Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-29 16:47 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-02 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-03 14:31 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-12-17 14:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-01-13 08:11 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-01-13 08:11 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; Tom Lane <tgl@sss.pgh.pa.us>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Hi Dmitry,
On Tue, Dec 17, 2024 at 7:40 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> I could verify the memory mappings, their sizes etc. by looking at
> /proc/PID/maps and /proc/PID/status but I did not find a way to verify
> the amount of memory actually allocated and verify that it's actually
> shrinking and expanding. Please let me know how to verify that.
As somewhere mentioned upthread, the mmap or mremap by themselves do
not allocate any memory. Writing to the mapped region causes memory to
be allocated and shows up in VmRSS and RssShmem. But it does get
resized if mremap() shrinks the mapped region.
Attached are patches rebased on top of commit
2a7b2d97171dd39dca7cefb91008a3c84ec003ba. I have also fixed
compilation errors. Otherwise I haven't changed anything in the
patches. The last patches adds some TODOs and questions, which I think
we need to address while completing this work, just add for as a
reminder later. The TODO in postgres.c is related to your observation
> Another rough edge is that a
> backend, executing pg_reload_conf interactively, will not resize
> mappings immediately, for some reason it will require another command.
I don't have a solution right now, but at least the comment documents
the reason and points to its origin.
I am next looking at the problem of synchronizing the change across
the backends.
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] 0002-Allow-placing-shared-memory-mapping-with-an-20250113.patch (8.7K, ../../CAExHW5tAKdTXeifgfL6zbJAzi0_H_=5ae8r5GTGg9bg8c1xuFQ@mail.gmail.com/2-0002-Allow-placing-shared-memory-mapping-with-an-20250113.patch)
download | inline diff:
From b5d295fc0a58b47228add95b4b8ad00fc0a66c07 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH 2/7] Allow placing shared memory mapping with an offset
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is first to get
the address chosen by the kernel, then apply some offset derived from
the expected upper limit. Because we base the layout on the address
chosen by the kernel, things like address space randomization should not
be a problem, since the randomization is applied to the mmap base, which
is one per process. The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
[...free space...]
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
This approach do not impact the actual memory usage as reported by the kernel.
Here is the output of /proc/$PID/status for the master version with
shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages mapped in mm_struct
VmPeak: 422780 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 21248 kB
// Size of resident anonymous memory
RssAnon: 640 kB
// Size of resident file mappings
RssFile: 9728 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 10880 kB
Here is the same for the patch with the shared mapping placed at
an offset 10 GB:
VmPeak: 1102844 kB
VmRSS: 21376 kB
RssAnon: 640 kB
RssFile: 9856 kB
RssShmem: 10880 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
18219008 (~17 MB)
Note that currently the implementation makes assumptions about the upper limit.
Ideally it should be based on the maximum available memory.
---
src/backend/port/sysv_shmem.c | 122 +++++++++++++++++++++++++++++++++-
1 file changed, 121 insertions(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 475c9c8f1a1..bae8f19a755 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,63 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping slots */
static int next_free_slot = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * By letting Linux to chose a mapping address, it will pick up the lowest
+ * possible address that do not clash with any other mappings, which will be
+ * right before locales in the example above. This information (maximum allowed
+ * size of mappings and the lowest mapping address) is enough to place every
+ * mapping as follow:
+ *
+ * - Take the lowest mapping address, which we call later the probe address.
+ * - Substract the offset of the previous mapping.
+ * - Substract the maximum allowed size for the current mapping from the
+ * address.
+ * - Place the mapping by the resulting address.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ */
+Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+ 0, /* MAIN_SHMEM_SLOT */
+};
+
+/* Remembers offset of the last mapping from the probe address */
+static Size last_offset = 0;
+
+/*
+ * Size of the mapping, which will be used to calculate anonymous mapping
+ * address. It should not be too small, otherwise there is a chance the probe
+ * mapping will be created between other mappings, leaving no room extending
+ * it. But it should not be too large either, in case if there are limitations
+ * on the mapping size. Current value is the default shared_buffers.
+ */
+#define PROBE_MAPPING_SIZE (Size) 128 * 1024 * 1024
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -673,13 +730,76 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
+ void *probe = NULL;
+
/*
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+
+ /*
+ * Try to create mapping at an address, which will allow to extend it
+ * later:
+ *
+ * - First create the temporary probe mapping of a fixed size and let
+ * kernel to place it at address of its choice. By the virtue of the
+ * probe mapping size we expect it to be located at the lowest
+ * possible address, expecting some non mapped space above.
+ *
+ * - Unmap the probe mapping, remember the address.
+ *
+ * - Create an actual anonymous mapping at that address with the
+ * offset. The offset is calculated in such a way to allow growing
+ * the mapping withing certain boundaries. For this mapping we use
+ * MAP_FIXED_NOREPLACE, which will error out with EEXIST if there is
+ * any mapping clash.
+ *
+ * - If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized
+ * without a restart.
+ */
+ probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
+
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: probe mmap(%zu) failed: %m",
+ MappingName(mapping->shmem_slot), allocsize);
+ }
+ else
+ {
+ Size offset = last_offset + SHMEM_EXTRA_SIZE_LIMIT[next_free_slot] + allocsize;
+ void *mapping_addr = (char *) probe - offset;
+
+ last_offset = offset;
+
+ munmap(probe, PROBE_MAPPING_SIZE);
+
+ ptr = mmap(mapping_addr, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ mmap_errno = errno;
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: mmap(%zu) at address %p failed: %m",
+ MappingName(mapping->shmem_slot), allocsize, mapping_addr);
+ }
+
+ }
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ /*
+ * Fallback to the portable way of creating a mapping.
+ */
+ allocsize = mapping->shmem_size;
+
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
}
--
2.34.1
[text/x-patch] 0005-Use-anonymous-files-to-back-shared-memory-s-20250113.patch (7.2K, ../../CAExHW5tAKdTXeifgfL6zbJAzi0_H_=5ae8r5GTGg9bg8c1xuFQ@mail.gmail.com/3-0005-Use-anonymous-files-to-back-shared-memory-s-20250113.patch)
download | inline diff:
From 746970c489f975b0d3add01b8d85d7cdab601b6d Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 15 Oct 2024 16:18:45 +0200
Subject: [PATCH 5/7] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f5a2bd04000-7f5a32e52000 rw-s 00000000 00:01 1845 /memfd:strategy (deleted)
7f5a39252000-7f5a4030e000 rw-s 00000000 00:01 1842 /memfd:checkpoint (deleted)
7f5a4670e000-7f5a4d7ba000 rw-s 00000000 00:01 1839 /memfd:iocv (deleted)
7f5a53bba000-7f5a5ad26000 rw-s 00000000 00:01 1836 /memfd:descriptors (deleted)
7f5a9ad26000-7f5aa9d94000 rw-s 00000000 00:01 1833 /memfd:buffers (deleted)
7f5d29d94000-7f5d30e00000 rw-s 00000000 00:01 1830 /memfd:main (deleted)
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 64 ++++++++++++++++++++++++++++++-----
src/include/portability/mem.h | 2 +-
2 files changed, 57 insertions(+), 9 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 72e823618ef..b2173e1a078 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -103,6 +103,7 @@ typedef struct AnonymousMapping
void *shmem; /* Pointer to the start of the mapped memory */
void *seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -116,7 +117,7 @@ static int next_free_slot = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -143,9 +144,9 @@ static int next_free_slot = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* [...free space...]
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* [...free space...]
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -708,6 +709,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_slot), 0);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
@@ -725,8 +738,13 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ /*
+ * Do not use an anonymous file here yet. When adding it, do not forget
+ * to use ftruncate and flags MFD_HUGETLB & MFD_HUGE_2MB/MFD_HUGE_1GB
+ * in memfd_create.
+ */
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
@@ -762,7 +780,8 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* - First create the temporary probe mapping of a fixed size and let
* kernel to place it at address of its choice. By the virtue of the
* probe mapping size we expect it to be located at the lowest
- * possible address, expecting some non mapped space above.
+ * possible address, expecting some non mapped space above. The probe
+ * is does not need to be backed by an anonymous file.
*
* - Unmap the probe mapping, remember the address.
*
@@ -777,7 +796,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* without a restart.
*/
probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS, -1, 0);
if (probe == MAP_FAILED)
{
@@ -795,8 +814,20 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
munmap(probe, PROBE_MAPPING_SIZE);
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ if (ftruncate(mapping->segment_fd, allocsize) < 0)
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: ftruncate(%zu) failed: %m",
+ MappingName(mapping->shmem_slot), allocsize);
+
+ }
+
ptr = mmap(mapping_addr, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
{
@@ -815,8 +846,17 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
*/
allocsize = mapping->shmem_size;
+ /* Specify the segment file size using allocsize. */
+ if (ftruncate(mapping->segment_fd, allocsize) < 0)
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: ftruncate(%zu) failed: %m",
+ MappingName(mapping->shmem_slot), allocsize);
+
+ }
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
@@ -905,6 +945,14 @@ AnonymousShmemResize(int newval, void *extra)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ if (ftruncate(m->segment_fd, new_size) < 0)
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: ftruncate(%zu) failed: %m",
+ MappingName(m->shmem_slot), new_size);
+ }
+
if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
elog(LOG, "mremap(%p, %zu) failed: %m",
m->shmem, m->shmem_size);
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index 2cd05313b82..50db0da28dc 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
--
2.34.1
[text/x-patch] 0001-Allow-to-use-multiple-shared-memory-mapping-20250113.patch (28.7K, ../../CAExHW5tAKdTXeifgfL6zbJAzi0_H_=5ae8r5GTGg9bg8c1xuFQ@mail.gmail.com/4-0001-Allow-to-use-multiple-shared-memory-mapping-20250113.patch)
download | inline diff:
From e7b3e3690a9b9575394cffbde2f7c5e674d28ed9 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 9 Oct 2024 15:41:32 +0200
Subject: [PATCH 1/7] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory slot.
There is only fixed amount of available slots, currently only one main
shared memory slot is allocated. A new shared memory API is introduces,
extended with a slot as a new parameter. As a path of least resistance,
the original API is kept in place, utilizing the main shared memory slot.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 +++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 2 +-
src/backend/storage/ipc/ipci.c | 61 ++++++------
src/backend/storage/ipc/shmem.c | 135 ++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 5 +-
src/include/storage/buf_internals.h | 1 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 10 ++
13 files changed, 261 insertions(+), 123 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 64186ec0a7e..b97723d2ede 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSlot(PGSemaphoreShmemSize(maxSemas), shmem_slot);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 68835723b90..e6720a6a077 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -313,7 +313,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
struct stat statbuf;
@@ -334,7 +334,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSlot(PGSemaphoreShmemSize(maxSemas), shmem_slot);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a5a4511f66d..475c9c8f1a1 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_slot;
+ Size shmem_size; /* Size of the mapping */
+ void *shmem; /* Pointer to the start of the mapped memory */
+ void *seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping slots */
+static int next_free_slot = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_slot)
+{
+ switch (shmem_slot)
+ {
+ case MAIN_SHMEM_SLOT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_slot; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "slot[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_slot), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("slot[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_slot)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_slot; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_slot];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_slot = next_free_slot;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_slot++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_slot;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_slot; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index f2b54bdfda0..d62084cc0d9 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_slot)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index b06e4b84528..2aabd4a77f3 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -68,7 +68,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 7783ba854fc..c0e1d94d1f7 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -85,7 +85,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_slot)
{
Size size;
int numSemas;
@@ -204,33 +204,36 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int slot = 0; slot < ANON_MAPPINGS; slot++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, slot);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSlot(seghdr, slot);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, slot);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSlot(slot);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -360,7 +363,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SLOT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 6d3074594a6..89d8c7baf16 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -75,17 +75,12 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSlot(Size size, Size *allocated_size,
+ int shmem_slot);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
-
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
+ShmemSegment Segments[ANON_MAPPINGS];
static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
@@ -96,9 +91,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSlot(seghdr, MAIN_SHMEM_SLOT);
+}
+
+void
+InitShmemAccessInSlot(PGShmemHeader *seghdr, int shmem_slot)
+{
+ ShmemSegment *seg = &Segments[shmem_slot];
+
+ seg->ShmemSegHdr = seghdr;
+ seg->ShmemBase = (void *) seghdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + seghdr->totalsize;
}
/*
@@ -109,7 +112,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSlot(MAIN_SHMEM_SLOT);
+}
+
+void
+InitShmemAllocationInSlot(int shmem_slot)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_slot].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -118,9 +127,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_slot].ShmemLock = (slock_t *) ShmemAllocUnlockedInSlot(sizeof(slock_t), shmem_slot);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_slot].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -145,11 +154,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSlot(size, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemAllocInSlot(Size size, int shmem_slot)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSlot(size, &allocated_size, shmem_slot);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -179,10 +194,17 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSlot(size, allocated_size, MAIN_SHMEM_SLOT);
+}
+
+static void *
+ShmemAllocRawInSlot(Size size, Size *allocated_size, int shmem_slot)
{
Size newStart;
Size newFree;
void *newSpace;
+ ShmemSegment *seg = &Segments[shmem_slot];
/*
* Ensure all space is adequately aligned. We used to only MAXALIGN this
@@ -198,22 +220,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(seg->ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(seg->ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = seg->ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= seg->ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) seg->ShmemBase + newStart;
+ seg->ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(seg->ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -231,29 +253,36 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSlot(size, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemAllocUnlockedInSlot(Size size, int shmem_slot)
{
Size newStart;
Size newFree;
void *newSpace;
+ ShmemSegment *seg = &Segments[shmem_slot];
/*
* Ensure allocated space is adequately aligned.
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(seg->ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = seg->ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > seg->ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ seg->ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) seg->ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -268,7 +297,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSlot(addr, MAIN_SHMEM_SLOT);
+}
+
+bool
+ShmemAddrIsValidInSlot(const void *addr, int shmem_slot)
+{
+ return (addr >= Segments[shmem_slot].ShmemBase) && (addr < Segments[shmem_slot].ShmemEnd);
}
/*
@@ -329,6 +364,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSlot(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SLOT);
+}
+
+HTAB *
+ShmemInitHashInSlot(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_slot) /* in which slot to keep the table */
{
bool found;
void *location;
@@ -345,9 +392,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSlot(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_slot);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -380,6 +427,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSlot(name, size, foundPtr, MAIN_SHMEM_SLOT);
+}
+
+void *
+ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
+ int shmem_slot)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -388,7 +442,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_slot].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -411,7 +465,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSlot(size, shmem_slot);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -428,8 +482,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in slot %d",
+ name, shmem_slot)));
}
if (*foundPtr)
@@ -454,7 +508,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSlot(size, &allocated_size, shmem_slot);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -473,14 +527,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSlot(structPtr, shmem_slot));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -540,7 +593,7 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SLOT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -552,15 +605,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SLOT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 9cf3e4f4f3a..cd3237b3736 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -81,6 +81,7 @@
#include "pgstat.h"
#include "port/pg_bitutils.h"
#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/spin.h"
@@ -607,9 +608,9 @@ LWLockNewTrancheId(void)
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
/* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SLOT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SLOT].ShmemLock);
return result;
}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index eda6c699212..b25dc0199b8 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -22,6 +22,7 @@
#include "storage/condition_variable.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
+#include "storage/pg_shmem.h"
#include "storage/smgr.h"
#include "storage/spin.h"
#include "utils/relcache.h"
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index b2d062781ec..be4b1312888 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_slot);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index dfef79ac963..081fffaf165 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_slot);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 3065ff5be71..e968deeef7f 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+// Number of available slots for anonymous memory mappings
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main slot, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SLOT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 8cdbe7a89c8..4261b4039b9 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,25 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSlot(struct PGShmemHeader *seghdr, int shmem_slot);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSlot(int shmem_slot);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSlot(Size size, int shmem_slot);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSlot(Size size, int shmem_slot);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSlot(const void *addr, int shmem_slot);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSlot(const char *name, long init_size, long max_size,
+ HASHCTL *infoP, int hash_flags, int shmem_slot);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
+ int shmem_slot);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
--
2.34.1
[text/x-patch] 0004-Allow-to-resize-shared-memory-without-resta-20250113.patch (12.5K, ../../CAExHW5tAKdTXeifgfL6zbJAzi0_H_=5ae8r5GTGg9bg8c1xuFQ@mail.gmail.com/5-0004-Allow-to-resize-shared-memory-without-resta-20250113.patch)
download | inline diff:
From ba93989e755b23f3cef21ad0d11a14231e866b07 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:24:58 +0200
Subject: [PATCH 4/7] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Size for every shared memory slot is recalculated based on the new
NBuffers, and extended using mremap. After allocating new space, new
shared structures (buffer blocks, descriptors, etc) are allocated as
needed. Here is how it looks like after raising shared_buffers from 128
MB to 512 MB and calling pg_reload_conf():
-- 128 MB
7f5a2bd04000-7f5a32e52000 /dev/zero (deleted)
7f5a39252000-7f5a4030e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d7ba000 /dev/zero (deleted)
7f5a53bba000-7f5a5ad26000 /dev/zero (deleted)
7f5a9ad26000-7f5aa9d94000 /dev/zero (deleted)
^ buffers mapping, ~240 MB
7f5d29d94000-7f5d30e00000 /dev/zero (deleted)
-- 512 MB
7f5a2bd04000-7f5a33274000 /dev/zero (deleted)
7f5a39252000-7f5a4057e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d9fa000 /dev/zero (deleted)
7f5a53bba000-7f5a5b1a6000 /dev/zero (deleted)
7f5a9ad26000-7f5ac1f14000 /dev/zero (deleted)
^ buffers mapping, ~625 MB
7f5d29d94000-7f5d30f80000 /dev/zero (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend, executing pg_reload_conf interactively, will not resize
mappings immediately, for some reason it will require another command.
Note, that mremap is Linux specific, thus the implementation not very
portable.
---
src/backend/port/sysv_shmem.c | 62 +++++++++++++
src/backend/storage/buffer/buf_init.c | 86 +++++++++++++++++++
src/backend/storage/ipc/ipci.c | 11 +++
src/backend/storage/ipc/shmem.c | 14 ++-
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/storage/bufmgr.h | 1 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 2 +
9 files changed, 171 insertions(+), 11 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 7157bf95b1a..72e823618ef 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,9 +30,11 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
@@ -861,6 +863,66 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * An assign callback for shared_buffers GUC -- a somewhat clumsy way of
+ * resizing shared memory without a restart. On NBuffers change use the new
+ * value to recalculate required size for every shmem slot, then base on the
+ * new and old values initialize new buffer blocks.
+ *
+ * The actual slot resizing is done via mremap, which will fail if is not
+ * sufficient space to expand the mapping.
+ *
+ * XXX: For some readon in the current implementation the change is applied to
+ * the backend calling pg_reload_conf only at the backend exit.
+ */
+void
+AnonymousShmemResize(int newval, void *extra)
+{
+ int numSemas;
+ bool reinit = false;
+ int NBuffersOld = NBuffers;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffers > newval)
+ return;
+
+ /* XXX: Hack, NBuffers has to be exposed in the the interface for
+ * memory calculation and buffer blocks reinitialization instead. */
+ NBuffers = newval;
+
+ for(int i = 0; i < next_free_slot; i++)
+ {
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
+ elog(LOG, "mremap(%p, %zu) failed: %m",
+ m->shmem, m->shmem_size);
+ else
+ {
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+ }
+
+ if (reinit)
+ {
+ LWLockAcquire(ShmemResizeLock, LW_EXCLUSIVE);
+ BufferManagerShmemResize(NBuffersOld);
+ LWLockRelease(ShmemResizeLock);
+ }
+}
+
/*
* PGSharedMemoryCreate
*
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index b066e97a0c9..ae58f82937f 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -153,6 +153,92 @@ BufferManagerShmemInit(void)
&backend_flush_after);
}
+/*
+ * Reinitialize shared memory structures, which size depends on NBuffers. It's
+ * similar to BufferManagerShmemInit, but applied only to the buffers in the range
+ * between NBuffersOld and NBuffers.
+ */
+void
+BufferManagerShmemResize(int NBuffersOld)
+{
+ bool foundBufs,
+ foundDescs,
+ foundIOCV,
+ foundBufCkpt;
+ int i;
+
+ /* XXX: Only increasing of shared_buffers is supported in this function */
+ if(NBuffersOld > NBuffers)
+ return;
+
+ /* Align descriptors to a cacheline boundary. */
+ BufferDescriptors = (BufferDescPadded *)
+ ShmemInitStructInSlot("Buffer Descriptors",
+ NBuffers * sizeof(BufferDescPadded),
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SLOT);
+
+ /* Align condition variables to cacheline boundary. */
+ BufferIOCVArray = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSlot("Buffer IO Condition Variables",
+ NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &foundIOCV, BUFFER_IOCV_SHMEM_SLOT);
+
+ /*
+ * The array used to sort to-be-checkpointed buffer ids is located in
+ * shared memory, to avoid having to allocate significant amounts of
+ * memory at runtime. As that'd be in the middle of a checkpoint, or when
+ * the checkpointer is restarted, memory allocation failures would be
+ * painful.
+ */
+ CkptBufferIds = (CkptSortItem *)
+ ShmemInitStructInSlot("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SLOT);
+
+ /* Align buffer pool on IO page size boundary. */
+ BufferBlocks = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSlot("Buffer Blocks",
+ NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &foundBufs, BUFFERS_SHMEM_SLOT));
+
+ /*
+ * Initialize the headers for new buffers.
+ */
+ for (i = NBuffersOld - 1; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
+
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+
+ buf->buf_id = i;
+
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
+
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
+
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ }
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+
+ /* Init other shared buffer-management stuff */
+ StrategyInitialize(!foundDescs);
+
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+}
+
/*
* BufferManagerShmemSize
*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index fd8b44b8161..15d06fd4ca4 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -83,6 +83,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory slots are incorrect, it includes
+ * more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_slot)
@@ -149,6 +152,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_slot)
size = add_size(size, InjectionPointShmemSize());
size = add_size(size, SlotSyncShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 89d8c7baf16..faca7c9a525 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -490,17 +490,13 @@ ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
*/
if (result->size != size)
- {
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
- }
+ result->size = size;
+
structPtr = result->location;
}
else
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 16144c2b72d..e8ecff5f7f0 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -345,6 +345,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 8cf1afbad20..7a12eedbbd3 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2318,14 +2318,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, AnonymousShmemResize, NULL
},
{
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 27c4cac8540..ead69a2974c 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -302,6 +302,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
extern Size BufferManagerShmemSize(int);
+extern void BufferManagerShmemResize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index 6a2f64c54fb..e8d379e4b0b 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -83,3 +83,4 @@ PG_LWLOCK(49, WALSummarizer)
PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
+PG_LWLOCK(53, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index c0143e38995..c1a96240d79 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -105,6 +105,8 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void AnonymousShmemResize(int newval, void *extra);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings slots. Each slot
--
2.34.1
[text/x-patch] 0003-Introduce-multiple-shmem-slots-for-shared-b-20250113.patch (10.8K, ../../CAExHW5tAKdTXeifgfL6zbJAzi0_H_=5ae8r5GTGg9bg8c1xuFQ@mail.gmail.com/6-0003-Introduce-multiple-shmem-slots-for-shared-b-20250113.patch)
download | inline diff:
From 4217a9e2b1922d22cc54b89bcdd6af9159fa9398 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:24:04 +0200
Subject: [PATCH 3/7] Introduce multiple shmem slots for shared buffers
Add more shmem slots to split shared buffers into following chunks:
* BUFFERS_SHMEM_SLOT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SLOT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SLOT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SLOT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SLOT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers, meaning
that if we would like to change NBuffers, they have to be resized
correspondingly. Placing each of them in a separate shmem slot allows to
achieve that.
There are some asumptions made about each of shmem slots upper size limit. The
buffer blocks have the largest, while the rest claim less extra room for
resize. Ideally those limits have to be deduced from the maximum allowed shared
memory.
---
src/backend/port/sysv_shmem.c | 17 +++++-
src/backend/storage/buffer/buf_init.c | 77 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 5 +-
src/backend/storage/buffer/freelist.c | 4 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 23 +++++++-
7 files changed, 96 insertions(+), 34 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index bae8f19a755..7157bf95b1a 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -149,8 +149,13 @@ static int next_free_slot = 0;
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
*/
-Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+Size SHMEM_EXTRA_SIZE_LIMIT[6] = {
0, /* MAIN_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 1024 * 10, /* BUFFERS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 1024 * 1, /* BUFFER_DESCRIPTORS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* BUFFER_IOCV_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* CHECKPOINT_BUFFERS_SHMEM_SLOT */
+ (Size) 1024 * 1024 * 100, /* STRATEGY_SHMEM_SLOT */
};
/* Remembers offset of the last mapping from the probe address */
@@ -179,6 +184,16 @@ MappingName(int shmem_slot)
{
case MAIN_SHMEM_SLOT:
return "main";
+ case BUFFERS_SHMEM_SLOT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SLOT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SLOT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SLOT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SLOT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 56761a8eedc..b066e97a0c9 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -61,7 +61,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory slot, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -73,22 +76,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSlot("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SLOT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSlot("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SLOT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSlot("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SLOT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -98,8 +101,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSlot("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SLOT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -153,33 +157,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory slot. The main slot must not allocate anything
+ * related to buffers, every other slot will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_slot)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_slot == MAIN_SHMEM_SLOT)
+ return size;
+
+ if (shmem_slot == BUFFER_DESCRIPTORS_SHMEM_SLOT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_slot == BUFFERS_SHMEM_SLOT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_slot == STRATEGY_SHMEM_SLOT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
+ if (shmem_slot == BUFFER_IOCV_SHMEM_SLOT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_slot == CHECKPOINT_BUFFERS_SHMEM_SLOT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 141dd724802..ff761574aa4 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -59,10 +59,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSlot("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SLOT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index dffdd57e9b5..325606dae71 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -491,9 +491,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSlot("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SLOT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index c0e1d94d1f7..fd8b44b8161 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -112,7 +112,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_slot)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_slot));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index eb0fba4230b..27c4cac8540 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -301,7 +301,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index e968deeef7f..c0143e38995 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
// Number of available slots for anonymous memory mappings
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -105,7 +105,28 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings slots. Each slot
+ * contains only certain part of the data, which size depends on NBuffers.
+ */
+
/* The main slot, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SLOT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SLOT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SLOT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SLOT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SLOT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SLOT 5
+
#endif /* PG_SHMEM_H */
--
2.34.1
[text/x-patch] 0006-Add-TODOs-and-questions-about-previous-comm-20250113.patch (10.3K, ../../CAExHW5tAKdTXeifgfL6zbJAzi0_H_=5ae8r5GTGg9bg8c1xuFQ@mail.gmail.com/7-0006-Add-TODOs-and-questions-about-previous-comm-20250113.patch)
download | inline diff:
From f33d7888253650c9f10634c8c28ea10c2e3d0fd8 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 6 Jan 2025 14:40:51 +0530
Subject: [PATCH 6/7] Add TODOs and questions about previous commits
The commit just marks the places which need more work or whether the
code raises some questions. This is not an exhaustive list of TODOs.
More TODOs may come up as I work with these patches further.
Ashutosh Bapat
---
src/backend/storage/buffer/buf_init.c | 9 +++++---
src/backend/storage/ipc/ipc.c | 4 ++++
src/backend/storage/ipc/ipci.c | 11 +++++++++-
src/backend/storage/ipc/shmem.c | 22 ++++++++++++++++---
src/backend/storage/lmgr/lwlock.c | 5 +++++
src/backend/tcop/postgres.c | 6 +++++
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/buf_internals.h | 5 +++++
src/include/storage/pg_shmem.h | 18 ++++++++++-----
9 files changed, 69 insertions(+), 12 deletions(-)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ae58f82937f..bf74b4b01a7 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -155,8 +155,11 @@ BufferManagerShmemInit(void)
/*
* Reinitialize shared memory structures, which size depends on NBuffers. It's
- * similar to BufferManagerShmemInit, but applied only to the buffers in the range
- * between NBuffersOld and NBuffers.
+ * similar to BufferManagerShmemInit, but applied only to the buffers in the
+ * range between NBuffersOld and NBuffers.
+ *
+ * TODO: Avoid code duplication with BufferManagerShmemInit() and also assess
+ * which functionality in the latter is required in this function.
*/
void
BufferManagerShmemResize(int NBuffersOld)
@@ -255,7 +258,7 @@ BufferManagerShmemSize(int shmem_slot)
if (shmem_slot == MAIN_SHMEM_SLOT)
return size;
-
+
if (shmem_slot == BUFFER_DESCRIPTORS_SHMEM_SLOT)
{
/* size of buffer descriptors */
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 2aabd4a77f3..556eb469f4f 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -68,6 +68,10 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
+/*
+ * TODO: Why do we need to increase this by 20? I didn't notice any new calls to
+ * on_shmem_exit or on_proc_exit or before_shmem_exit.
+ */
#define MAX_ON_EXITS 40
struct ONEXIT
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 15d06fd4ca4..e076f96ebf2 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -204,6 +204,12 @@ AttachSharedMemoryStructs(void)
/*
* CreateSharedMemoryAndSemaphores
* Creates and initializes shared memory and semaphores.
+ *
+ * TODO: IMO this function should be rewritten to calculate the size of each
+ * shared memory slot or mapping. Instead of passing slot number to
+ * CalculateShmemSize, we should instead let each shared memory module use their
+ * own slot number and update the required sizes in the corresponding mapping.
+ * Then allocate shared memory in each of the mappings.
*/
void
CreateSharedMemoryAndSemaphores(void)
@@ -222,7 +228,10 @@ CreateSharedMemoryAndSemaphores(void)
elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
/*
- * Create the shmem segment
+ * Create the shmem segment.
+ * TODO: while each slot will return a different shim, only the last one
+ * is passed to dsm_postmaster_startup(). Is that right? Shouldn't we
+ * pass all of them or none.
*/
seghdr = PGSharedMemoryCreate(size, &shim);
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index faca7c9a525..c1dde378329 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -82,6 +82,7 @@ static void *ShmemAllocRawInSlot(Size size, Size *allocated_size,
ShmemSegment Segments[ANON_MAPPINGS];
+/*TODO: shouldn't this be part of the ShmemSegment structure? */
static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
@@ -490,9 +491,18 @@ ShmemInitStructInSlot(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. Verify the structure's size:
- * - If it's the same, we've found the expected structure.
- * - If it's different, we're resizing the expected structure.
+ * already. Verify the structure's size: - If it's the same, we've found
+ * the expected structure. - If it's different, we're resizing the
+ * expected structure.
+ *
+ * TODO: This works because every structure that needs to be resized
+ * resides in a shmem slot by itself. But it won't work if a slot
+ * contains more structures, that need to be resized, placed in adjacent
+ * memory. Also we are not updating the Shmem stats like freeoffset. I
+ * think we will keep all resizable structures in a slot for themselves,
+ * and not have a hash table in such slots since resizing the hash table
+ * itself might cause memory to be allocated next to the resizable
+ * structure making it difficult to resize it.
*/
if (result->size != size)
result->size = size;
@@ -584,6 +594,12 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
hash_seq_init(&hstat, ShmemIndex);
+ /*
+ * TODO: For the sake of completeness we should rotate through all the slots
+ * (after saving slotwise ShmemIndex, if any). Do we want to also output
+ * shmem slot name, but that would expose the slotified structure of shared
+ * memory.
+ */
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index cd3237b3736..0be59074709 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -608,6 +608,11 @@ LWLockNewTrancheId(void)
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
/* We use the ShmemLock spinlock to protect LWLockCounter */
+ /*
+ * TODO: We have retained ShmemLock global variable, should we use it here
+ * instead of main segment lock? We will need spinlock init on the global
+ * one if yes.
+ */
SpinLockAcquire(Segments[MAIN_SHMEM_SLOT].ShmemLock);
result = (*LWLockCounter)++;
SpinLockRelease(Segments[MAIN_SHMEM_SLOT].ShmemLock);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 85902788181..8c89928203b 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4656,6 +4656,12 @@ PostgresMain(const char *dbname, const char *username)
/*
* (6) check for any other interesting events that happened while we
* slept.
+ * TODO: When a backend is waiting for a command, it won't reload
+ * configuration and hence wouldn't notice change in shared_buffers. The
+ * change is only noticed after the command is received and the control
+ * comes here. We may need to improve this in case we want to resize
+ * shared buffers or perform of part of that operation in assign_hook
+ * implementation (e.g. AnonymousShmemResize()).
*/
if (ConfigReloadPending)
{
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index e8ecff5f7f0..acd94c3616c 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -345,6 +345,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+# TODO, not used anywhere, do we need it?
ShmemResize "Waiting to resize shared memory."
#
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index b25dc0199b8..6a352d3942e 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -22,6 +22,11 @@
#include "storage/condition_variable.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
+/*
+ * TODO: this header files doesn't use anything in pg_shmem.h but the files which
+ * include this file may. We should include pg_shmem.h in those files rather than
+ * here.
+ */
#include "storage/pg_shmem.h"
#include "storage/smgr.h"
#include "storage/spin.h"
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index c1a96240d79..39521208fb9 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -42,6 +42,10 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+/*
+ * TODO: should we define it in shmem.c where the previous global variables were
+ * declared? Do we need this structure outside shmem.c?
+ */
typedef struct ShmemSegment
{
PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
@@ -51,11 +55,6 @@ typedef struct ShmemSegment
* allocation */
} ShmemSegment;
-// Number of available slots for anonymous memory mappings
-#define ANON_MAPPINGS 6
-
-extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
-
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -111,6 +110,10 @@ extern void AnonymousShmemResize(int newval, void *extra);
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings slots. Each slot
* contains only certain part of the data, which size depends on NBuffers.
+ *
+ * TODO: convert this into an enum with a sentinel symbol ANON_MAPPINGS, which
+ * itself should be renamed to NUM_ANON_MAPPINGS or NUM_SHMEM_SEGMENTS or
+ * something that indicates that it's the number of shared memory segments.
*/
/* The main slot, contains everything except buffer blocks and related data. */
@@ -131,4 +134,9 @@ extern void AnonymousShmemResize(int newval, void *extra);
/* Buffer strategy status */
#define STRATEGY_SHMEM_SLOT 5
+// Number of available slots for anonymous memory mappings
+#define ANON_MAPPINGS 6
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
#endif /* PG_SHMEM_H */
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
@ 2024-11-28 18:45 ` Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2024-11-28 18:45 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: pgsql-hackers
> On Thu, Nov 28, 2024 at 12:18:54PM GMT, Robert Haas wrote:
>
> All that having been said, what does concern me a bit is our ability
> to predict what Linux will do well enough to keep what we're doing
> safe; and also whether the Linux behavior might abruptly change in the
> future. Users would be sad if we released this feature and then a
> future kernel upgrade causes PostgreSQL to completely stop working. I
> don't know how the Linux kernel developers actually feel about this
> sort of thing, but if I imagine myself as a kernel developer, I can
> totally see myself saying "well, we never promised that this would
> work in any particular way, so we're free to change it whenever we
> like." We've certainly used that argument here countless times.
Agree, at the moment I can't say for sure how reliable this behavior is
in long term. I'll try to see if there are ways to get more confidence
about that.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-25 19:33 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Re: Changing shared_buffers without restart Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2024-11-29 18:17 ` Andres Freund <andres@anarazel.de>
1 sibling, 0 replies; 167+ messages in thread
From: Andres Freund @ 2024-11-29 18:17 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Robert Haas <robertmhaas@gmail.com>; pgsql-hackers
Hi,
On 2024-11-28 17:30:32 +0100, Dmitry Dolgov wrote:
> The assumption about picking up a lowest address is just how it works right now
> on Linux, this fact is already used in the patch. The idea that we could put
> upper boundary on the size of other mappings based on total available memory
> comes from the fact that anonymous mappings, that are much larger than memory,
> will fail without overcommit.
The overcommit issue shouldn't be a big hurdle - by mmap()ing with
MAP_NORESERVE the space isn't reserved. Then madvise with MADV_POPULATE_WRITE
can be used to actually populate the used range of the mapping and MADV_REMOVE
can be used to shrink the mapping again.
> With overcommit it becomes different, but if allocations are hitting that
> limit I can imagine there are bigger problems than shared buffer resize.
I'm fairly sure it'll not work to just disregard issues around overcommit. A
overly large memory allocation, without MAP_NORESERVE, will actually reduce
the amount of memory that can be used for other allocations. That's obviously
problematic, because you'll now have a smaller shared buffers, but can't use
the memory for work_mem type allocations...
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-02-25 09:52 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-27 08:28 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 11:21 ` Re: Changing shared_buffers without restart Konstantin Knizhnik <knizhnik@garret.ru>
4 siblings, 3 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-02-25 09:52 UTC (permalink / raw)
To: pgsql-hackers; +Cc: Robert Haas <robertmhaas@gmail.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
> On Fri, Oct 18, 2024 at 09:21:19PM GMT, Dmitry Dolgov wrote:
> TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
> changing shared memory mapping layout. Any feedback is appreciated.
Hi,
Here is a new version of the patch, which contains a proposal about how to
coordinate shared memory resizing between backends. The rest is more or less
the same, a feedback about coordination is appreciated. It's a lot to read, but
the main difference is about:
1. Allowing to decouple a GUC value change from actually applying it, sort of a
"pending" change. The idea is to let a custom logic be triggered on an assign
hook, and then take responsibility for what happens later and how it's going to
be applied. This allows to use regular GUC infrastructure in cases where value
change requires some complicated processing. I was trying to make the change
not so invasive, plus it's missing GUC reporting yet.
2. Shared memory resizing patch became more complicated thanks to some
coordination between backends. The current implementation was chosen from few
more or less equal alternatives, which are evolving along following lines:
* There should be one "coordinator" process overseeing the change. Having
postmaster to fulfill this role like in this patch seems like a natural idea,
but it poses certain challenges since it doesn't have locking infrastructure.
Another option would be to elect a single backend to be a coordinator, which
will handle the postmaster as a special case. If there will ever be a
"coordinator" worker in Postgres, that would be useful here.
* The coordinator uses EmitProcSignalBarrier to reach out to all other backends
and trigger the resize process. Backends join a Barrier to synchronize and wait
untill everyone is finished.
* There is some resizing state stored in shared memory, which is there to
handle backends that were for some reason late or didn't receive the signal.
What to store there is open for discussion.
* Since we want to make sure all processes share the same understanding of what
NBuffers value is, any failure is mostly a hard stop, since to rollback the
change coordination is needed as well and sounds a bit too complicated for now.
We've tested this change manually for now, although it might be useful to try
out injection points. The testing strategy, which has caught plenty of bugs,
was simply to run pgbench workload against a running instance and change
shared_buffers on the fly. Some more subtle cases were verified by manually
injecting delays to trigger expected scenarios.
To reiterate, here is patches breakdown:
Patches 1-3 prepare the infrastructure and shared memory layout. They could be
useful even with multithreaded PostgreSQL, when there will be no need for
shared memory. I assume, in the multithreaded world there still will be need
for a contiguous chunk of memory to share between threads, and its layout would
be similar to the one with shared memory mappings. Note that patch nr 2 is
going away as soon as I'll get to implement shared memory address reservation,
but for now it's needed.
Patch 4 is a new addition to handle "pending" GUC changes.
Patch 5 actually does resizing. It's shared memory specific of course, and
utilized Linux specific mremap, meaning open portability questions.
Patch 6 is somewhat independent, but quite convenient to have. It also utilizes
Linux specific call memfd_create.
I would like to get some feedback on the synchronization part. While waiting
I'll proceed implementing shared memory address space reservation and Ashutosh
will continue with buffer eviction to support shared memory reduction.
From d88185fb3b4a3a0e102a3af52f4fb5564468db15 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 19 Feb 2025 17:43:13 +0100
Subject: [PATCH v2 1/6] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 ++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 141 +++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 5 +-
src/include/storage/buf_internals.h | 1 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
13 files changed, 272 insertions(+), 124 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index f7c8638aec5..b6301463ac7 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -313,7 +313,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -334,7 +334,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..843b1b3220f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ void *shmem; /* Pointer to the start of the mapped memory */
+ void *seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index e4d5b944e12..9d526eb43fd 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 174eed70367..4f6c707c204 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -85,7 +85,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -204,33 +204,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -360,7 +365,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 895a43fb39e..389abc82519 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -75,19 +75,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/*
@@ -96,9 +96,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -109,7 +117,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -118,9 +132,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -145,11 +159,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -179,6 +199,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -198,22 +224,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -231,6 +257,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -241,19 +273,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -268,7 +300,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -329,6 +367,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -345,9 +395,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -380,6 +430,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -388,7 +445,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -411,7 +468,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -428,8 +485,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -454,7 +511,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -473,14 +530,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -537,10 +593,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -552,15 +609,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index f1e74f184f1..40aa4014b5f 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -81,6 +81,7 @@
#include "pgstat.h"
#include "port/pg_bitutils.h"
#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/spin.h"
@@ -607,9 +608,9 @@ LWLockNewTrancheId(void)
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
/* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 1a65342177d..4595f5a9676 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -22,6 +22,7 @@
#include "storage/condition_variable.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
+#include "storage/pg_shmem.h"
#include "storage/smgr.h"
#include "storage/spin.h"
#include "utils/relcache.h"
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index e0f5f92e947..c0439f2206b 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index b99ebc9e86f..138078c29c5 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 904a336b851..5929f140236 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: 80d7f990496b1c7be61d9a00a2635b7d96b96197
--
2.45.1
From 7543fcdfc8ca1a0e1c85f397eb6dddfe1426b379 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH v2 2/6] Allow placing shared memory mapping with an offset
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is first to get
the address chosen by the kernel, then apply some offset derived from
the expected upper limit. Because we base the layout on the address
chosen by the kernel, things like address space randomization should not
be a problem, since the randomization is applied to the mmap base, which
is one per process. The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
[...free space...]
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
This approach do not impact the actual memory usage as reported by the kernel.
Here is the output of /proc/$PID/status for the master version with
shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages mapped in mm_struct
VmPeak: 422780 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 21248 kB
// Size of resident anonymous memory
RssAnon: 640 kB
// Size of resident file mappings
RssFile: 9728 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 10880 kB
Here is the same for the patch with the shared mapping placed at
an offset 10 GB:
VmPeak: 1102844 kB
VmRSS: 21376 kB
RssAnon: 640 kB
RssFile: 9856 kB
RssShmem: 10880 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
18219008 (~17 MB)
Note that currently the implementation makes assumptions about the upper limit.
Ideally it should be based on the maximum available memory.
---
src/backend/port/sysv_shmem.c | 120 +++++++++++++++++++++++++++++++++-
1 file changed, 119 insertions(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 843b1b3220f..62f01d8218a 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,63 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * By letting Linux to chose a mapping address, it will pick up the lowest
+ * possible address that do not clash with any other mappings, which will be
+ * right before locales in the example above. This information (maximum allowed
+ * size of mappings and the lowest mapping address) is enough to place every
+ * mapping as follow:
+ *
+ * - Take the lowest mapping address, which we call later the probe address.
+ * - Substract the offset of the previous mapping.
+ * - Substract the maximum allowed size for the current mapping from the
+ * address.
+ * - Place the mapping by the resulting address.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ */
+Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+ 0, /* MAIN_SHMEM_SLOT */
+};
+
+/* Remembers offset of the last mapping from the probe address */
+static Size last_offset = 0;
+
+/*
+ * Size of the mapping, which will be used to calculate anonymous mapping
+ * address. It should not be too small, otherwise there is a chance the probe
+ * mapping will be created between other mappings, leaving no room extending
+ * it. But it should not be too large either, in case if there are limitations
+ * on the mapping size. Current value is the default shared_buffers.
+ */
+#define PROBE_MAPPING_SIZE (Size) 128 * 1024 * 1024
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -673,13 +730,74 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
+ void *probe = NULL;
+
/*
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+
+ /*
+ * Try to create mapping at an address, which will allow to extend it
+ * later:
+ *
+ * - First create the temporary probe mapping of a fixed size and let
+ * kernel to place it at address of its choice. By the virtue of the
+ * probe mapping size we expect it to be located at the lowest
+ * possible address, expecting some non mapped space above.
+ *
+ * - Unmap the probe mapping, remember the address.
+ *
+ * - Create an actual anonymous mapping at that address with the
+ * offset. The offset is calculated in such a way to allow growing
+ * the mapping withing certain boundaries. For this mapping we use
+ * MAP_FIXED_NOREPLACE, which will error out with EEXIST if there is
+ * any mapping clash.
+ *
+ * - If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized
+ * without a restart.
+ */
+ probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
+
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: probe mmap(%zu) failed: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
+ else
+ {
+ Size offset = last_offset + SHMEM_EXTRA_SIZE_LIMIT[next_free_segment] + allocsize;
+ last_offset = offset;
+
+ munmap(probe, PROBE_MAPPING_SIZE);
+
+ ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ mmap_errno = errno;
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m",
+ MappingName(mapping->shmem_segment), allocsize, probe - offset);
+ }
+
+ }
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ /*
+ * Fallback to the portable way of creating a mapping.
+ */
+ allocsize = mapping->shmem_size;
+
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
}
--
2.45.1
From d7af86299878acb73019ac699ef57c120199f1ee Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Mon, 24 Feb 2025 20:08:28 +0100
Subject: [PATCH v2 3/6] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 19 ++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 5 +-
src/backend/storage/buffer/freelist.c | 4 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 99 insertions(+), 36 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 62f01d8218a..59aa67cb135 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -149,8 +149,13 @@ static int next_free_segment = 0;
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
*/
-Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
- 0, /* MAIN_SHMEM_SLOT */
+Size SHMEM_EXTRA_SIZE_LIMIT[6] = {
+ 0, /* MAIN_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 1024 * 10, /* BUFFERS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 1024 * 1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* STRATEGY_SHMEM_SEGMENT */
};
/* Remembers offset of the last mapping from the probe address */
@@ -179,6 +184,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1f8e03190..f5b9290a640 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -61,7 +61,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -73,22 +76,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -98,8 +101,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -153,33 +157,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..ac449954dab 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -59,10 +59,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 336715b6c63..4919a92f2be 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -491,9 +491,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 4f6c707c204..68778522591 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -112,7 +112,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 7c1e4316dde..bb7fe02e243 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -297,7 +297,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 138078c29c5..ba0192baf95 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -105,7 +105,29 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.45.1
From 0173967e8b0fd6c23b158c34b92651fc37ab7660 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 19 Feb 2025 17:45:40 +0100
Subject: [PATCH v2 4/6] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 2 +-
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 16 ++++----
8 files changed, 57 insertions(+), 36 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index f9bf5ba7509..ff82ba0a53d 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2188,7 +2188,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 4ad6e236d69..f24c2a0d252 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
#ifdef USE_PREFETCH
/*
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index 61ea3722ae2..cdf21847d7e 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1949,7 +1949,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1982,7 +1982,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2005,7 +2005,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2028,7 +2028,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 1149d89d7a1..13fb8c31702 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3555,7 +3555,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 12192445218..bab1c5d08f6 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 1233e07d7da..ce9f258100d 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 951451a9765..3e380f29e5a 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,12 +81,12 @@ extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
extern bool check_maintenance_io_concurrency(int *newval, void **extra,
GucSource source);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -141,13 +141,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -163,7 +163,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.45.1
From 78ea0efde8799445b90a70ca321e40b75fea52c9 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 20 Feb 2025 21:12:26 +0100
Subject: [PATCH v2 5/6] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a global
Barrier to coordinate backends that simultaneously change
shared_buffers, and pieces in shared memory to coordinate backends that
are too late to the party for some reason.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier. Afterwards it does shared memory
resize itself.
* All the backends, that participate in ProcSignal mechanism,
recalculate shared memory size based on the new NBuffers and extend it
using mremap.
* When finished, a backend waits on a global ShmemControl barrier,
untill all backends will be finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all backends are using new value, one backend will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f5a2bd04000-7f5a32e52000 /dev/zero (deleted)
7f5a39252000-7f5a4030e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d7ba000 /dev/zero (deleted)
7f5a53bba000-7f5a5ad26000 /dev/zero (deleted)
7f5a9ad26000-7f5aa9d94000 /dev/zero (deleted)
^ buffers mapping, ~240 MB
7f5d29d94000-7f5d30e00000 /dev/zero (deleted)
-- 512 MB
7f5a2bd04000-7f5a33274000 /dev/zero (deleted)
7f5a39252000-7f5a4057e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d9fa000 /dev/zero (deleted)
7f5a53bba000-7f5a5b1a6000 /dev/zero (deleted)
7f5a9ad26000-7f5ac1f14000 /dev/zero (deleted)
^ buffers mapping, ~625 MB
7f5d29d94000-7f5d30f80000 /dev/zero (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it reads something.
Note, that mremap is Linux specific, thus the implementation not very
portable.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 300 ++++++++++++++++++
src/backend/postmaster/postmaster.c | 15 +
src/backend/storage/buffer/buf_init.c | 152 ++++++++-
src/backend/storage/ipc/ipci.c | 11 +
src/backend/storage/ipc/procsignal.c | 45 +++
src/backend/storage/ipc/shmem.c | 14 +-
src/backend/tcop/postgres.c | 15 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/storage/bufmgr.h | 1 +
src/include/storage/ipc.h | 2 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 24 ++
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
15 files changed, 577 insertions(+), 12 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 59aa67cb135..35a8ff92175 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,17 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/procsignal.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -105,6 +109,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -859,6 +870,274 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * mremap, which will fail if is not sufficient space to expand the mapping.
+ * When finished, based on the new and old values initialize new buffer blocks
+ * if any.
+ *
+ * If reinitializing took place, as the last step this function broadcasts
+ * NSharedBuffers to it's new value, allowing any other backends to rely on
+ * this new value and skip buffers reinitialization.
+ */
+static bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+
+ /*
+ * Fail hard if faced any issues. In theory we could try to handle this
+ * more gracefully and proceed with shared memory as before, but some
+ * other backends might have succeeded and have different size. If we
+ * would like to go this way, to be consistent we would need to
+ * synchronize again, and it's not clear if it's worth the effort.
+ */
+ if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not resize shared memory %p to %d (%zu): %m",
+ m->shmem, NBuffers, m->shmem_size)));
+ else
+ {
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ ResizeBufferPool(NBuffersOld, true);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+ else
+ ResizeBufferPool(NBuffersOld, false);
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Do the resize and make sure to wait on
+ * the provided barrier until all simultaneously participating backends finish
+ * resizing as well, otherwise we face danger of inconsistency between
+ * backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * After attaching to the barrier we could be in any of states:
+ *
+ * - Initial SHMEM_RESIZE_REQUESTED, nothing has been done yet
+ * - SHMEM_RESIZE_START, some of the backends have started to resize
+ * - SHMEM_RESIZE_DONE, participating backends have finished resizing
+ * - SHMEM_RESIZE_REQUESTED after the reset, the shared memory was already
+ * resized
+ *
+ * The first three states take place while the actual resize is in
+ * progress, and all we need to do is join and proceed with resizing. This
+ * way all simultaneously participating backends will remap and wait until
+ * one of them initialize new buffers.
+ *
+ * The last state happens when we are too late and everything is already
+ * done. In that case proceed as well, relying on AnonymousShmemResize not
+ * reinitialize anything since the NSharedBuffers is already broadcasted.
+ */
+ BarrierAttach(barrier);
+
+ /* First phase means the resize has begun, SHMEM_RESIZE_START */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
+
+ /* XXX: Split mremap and buffer reinitialization into two barrier phases */
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ /* Allow the last backend to reset the barrier */
+ if (BarrierArriveAndDetach(barrier))
+ ResetShmemBarrier();
+
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ elog(DEBUG1, "Set pending signal");
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Coordinate all existing processes to make sure they all will have consistent
+ * view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ * There might still be some backends, that are sleeping or for some
+ * other reason not doing any work yet and have old NBuffers -- but as
+ * soon as they will get some time slice, they will acquire the new
+ * value.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE);
+
+ AnonymousShmemResize();
+
+ /*
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier, and there is no sequential logic
+ * we have to perform afterwards. In fact even if we would like to wait
+ * here, it wouldn't be possible -- we're in the postmaster, without any
+ * waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1174,3 +1453,24 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier(int phase)
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ if (BarrierPhase(barrier) == phase)
+ {
+ ereport(LOG,
+ (errmsg("ProcSignal barrier is in phase %d, waiting", phase)));
+ BarrierAttach(barrier);
+ BarrierArriveAndWait(barrier, 0);
+ BarrierDetach(barrier);
+ }
+}
+
+void
+ResetShmemBarrier(void)
+{
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+}
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index bb22b13adef..f3e508141b2 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -418,6 +418,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1680,6 +1681,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2026,6 +2030,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index f5b9290a640..b7de0ab6b0d 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -23,6 +23,41 @@ ConditionVariableMinimallyPadded *BufferIOCVArray;
WritebackContext BackendWritebackContext;
CkptSortItem *CkptBufferIds;
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
/*
* Data Structures:
@@ -72,7 +107,19 @@ BufferManagerShmemInit(void)
bool foundBufs,
foundDescs,
foundIOCV,
- foundBufCkpt;
+ foundBufCkpt,
+ foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+ }
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -153,6 +200,109 @@ BufferManagerShmemInit(void)
&backend_flush_after);
}
+/*
+ * Reinitialize shared memory structures, which size depends on NBuffers. It's
+ * similar to InitBufferPool, but applied only to the buffers in the range
+ * between NBuffersOld and NBuffers.
+ *
+ * NBuffersOld tells what was the original value of NBuffersOld. It will be
+ * used to identify new and not yet initialized buffers.
+ *
+ * initNew flag indicates that the caller wants new buffers to be initialized.
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
+ */
+void
+ResizeBufferPool(int NBuffersOld, bool initNew)
+{
+ bool foundBufs,
+ foundDescs,
+ foundIOCV,
+ foundBufCkpt;
+ int i;
+ elog(DEBUG1, "Resizing buffer pool from %d to %d", NBuffersOld, NBuffers);
+
+ /* XXX: Only increasing of shared_buffers is supported in this function */
+ if(NBuffersOld > NBuffers)
+ return;
+
+ /* Align descriptors to a cacheline boundary. */
+ BufferDescriptors = (BufferDescPadded *)
+ ShmemInitStructInSegment("Buffer Descriptors",
+ NBuffers * sizeof(BufferDescPadded),
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+
+ /* Align condition variables to cacheline boundary. */
+ BufferIOCVArray = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
+
+ /*
+ * The array used to sort to-be-checkpointed buffer ids is located in
+ * shared memory, to avoid having to allocate significant amounts of
+ * memory at runtime. As that'd be in the middle of a checkpoint, or when
+ * the checkpointer is restarted, memory allocation failures would be
+ * painful.
+ */
+ CkptBufferIds = (CkptSortItem *)
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+
+ /* Align buffer pool on IO page size boundary. */
+ BufferBlocks = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSegment("Buffer Blocks",
+ NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
+
+ /*
+ * It's enough to only resize shmem structures, if some other backend will
+ * do initialization of new buffers for us.
+ */
+ if (!initNew)
+ return;
+
+ elog(DEBUG1, "Initialize new buffers");
+
+ /*
+ * Initialize the headers for new buffers.
+ */
+ for (i = NBuffersOld; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
+
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+
+ buf->buf_id = i;
+
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
+
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
+
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ }
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+
+ /* Init other shared buffer-management stuff */
+ StrategyInitialize(!foundDescs);
+
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+}
+
/*
* BufferManagerShmemSize
*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 68778522591..a2c635f288e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -83,6 +83,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -149,6 +152,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, InjectionPointShmemSize());
size = add_size(size, SlotSyncShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 7401b6e625e..bec0e00f901 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -108,6 +109,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -168,6 +173,42 @@ ProcSignalInit(bool cancel_key_valid, int32 cancel_key)
ProcSignalSlot *slot;
uint64 barrier_generation;
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -573,6 +614,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 389abc82519..226b38ba979 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -493,17 +493,13 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
*/
if (result->size != size)
- {
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
- }
+ result->size = size;
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 13fb8c31702..04cdd0d24d8 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4267,6 +4268,20 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /*
+ * Verify the shared barrier, if it's still active: join and wait.
+ *
+ * XXX: Any potential race condition if not a single backend has
+ * incremented the barrier phase?
+ */
+ WaitOnShmemBarrier(SHMEM_RESIZE_START);
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index e199f071628..012acb98169 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -154,6 +154,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
WAL_RECEIVER_WAIT_START "Waiting for startup process to send initial data for streaming replication."
@@ -346,6 +348,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 3cde94a1759..efdaa71c8fb 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2339,14 +2339,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index bb7fe02e243..fff80214822 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -298,6 +298,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
extern Size BufferManagerShmemSize(int);
+extern void ResizeBufferPool(int, bool);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index c0439f2206b..5f5b45c88bd 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
extern void proc_exit(int code) pg_attribute_noreturn();
extern void shmem_exit(int code);
@@ -83,5 +84,6 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index cf565452382..61e89c6e8fd 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -83,3 +83,4 @@ PG_LWLOCK(49, WALSummarizer)
PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
+PG_LWLOCK(53, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index ba0192baf95..b597df0d3a3 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,23 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -105,6 +123,12 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(int phase);
+extern void ResetShmemBarrier(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 022fd8ed933..4c9973dc2d9 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index fb39c915d76..5bf6d099808 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2671,6 +2671,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.45.1
From e511bab55891a2d60152e913df6c20e20314e71b Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 23 Feb 2025 14:42:39 +0100
Subject: [PATCH v2 6/6] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f5a2bd04000-7f5a32e52000 rw-s 00000000 00:01 1845 /memfd:strategy (deleted)
7f5a39252000-7f5a4030e000 rw-s 00000000 00:01 1842 /memfd:checkpoint (deleted)
7f5a4670e000-7f5a4d7ba000 rw-s 00000000 00:01 1839 /memfd:iocv (deleted)
7f5a53bba000-7f5a5ad26000 rw-s 00000000 00:01 1836 /memfd:descriptors (deleted)
7f5a9ad26000-7f5aa9d94000 rw-s 00000000 00:01 1833 /memfd:buffers (deleted)
7f5d29d94000-7f5d30e00000 rw-s 00000000 00:01 1830 /memfd:main (deleted)
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 46 +++++++++++++++++++++++++++++------
src/include/portability/mem.h | 2 +-
2 files changed, 39 insertions(+), 9 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 35a8ff92175..8864866f26c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -105,6 +105,7 @@ typedef struct AnonymousMapping
void *shmem; /* Pointer to the start of the mapped memory */
void *seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -125,7 +126,7 @@ static int next_free_segment = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -152,9 +153,9 @@ static int next_free_segment = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* [...free space...]
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* [...free space...]
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -717,6 +718,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment), 0);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
@@ -734,8 +747,13 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ /*
+ * Do not use an anonymous file here yet. When adding it, do not forget
+ * to use ftruncate and flags MFD_HUGETLB & MFD_HUGE_2MB/MFD_HUGE_1GB
+ * in memfd_create.
+ */
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
@@ -771,7 +789,8 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* - First create the temporary probe mapping of a fixed size and let
* kernel to place it at address of its choice. By the virtue of the
* probe mapping size we expect it to be located at the lowest
- * possible address, expecting some non mapped space above.
+ * possible address, expecting some non mapped space above. The probe
+ * is does not need to be backed by an anonymous file.
*
* - Unmap the probe mapping, remember the address.
*
@@ -786,7 +805,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* without a restart.
*/
probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS, -1, 0);
if (probe == MAP_FAILED)
{
@@ -802,8 +821,14 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
munmap(probe, PROBE_MAPPING_SIZE);
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
{
@@ -822,8 +847,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
*/
allocsize = mapping->shmem_size;
+ /* Specify the segment file size using allocsize. */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
@@ -917,6 +945,8 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ ftruncate(m->segment_fd, new_size);
/*
* Fail hard if faced any issues. In theory we could try to handle this
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
--
2.45.1
Attachments:
[text/plain] v2-0001-Allow-to-use-multiple-shared-memory-mappings.patch (30.1K, ../../xuuxvlhom2tiinwwnh7r6wds74o2fkwryy6palehytuzm76l4t@3q7lszfqic3b/2-v2-0001-Allow-to-use-multiple-shared-memory-mappings.patch)
download | inline diff:
From d88185fb3b4a3a0e102a3af52f4fb5564468db15 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 19 Feb 2025 17:43:13 +0100
Subject: [PATCH v2 1/6] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 ++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 141 +++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 5 +-
src/include/storage/buf_internals.h | 1 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
13 files changed, 272 insertions(+), 124 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index f7c8638aec5..b6301463ac7 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -313,7 +313,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -334,7 +334,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..843b1b3220f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ void *shmem; /* Pointer to the start of the mapped memory */
+ void *seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index e4d5b944e12..9d526eb43fd 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 174eed70367..4f6c707c204 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -85,7 +85,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -204,33 +204,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -360,7 +365,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 895a43fb39e..389abc82519 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -75,19 +75,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/*
@@ -96,9 +96,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -109,7 +117,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -118,9 +132,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -145,11 +159,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -179,6 +199,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -198,22 +224,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -231,6 +257,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -241,19 +273,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -268,7 +300,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -329,6 +367,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -345,9 +395,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -380,6 +430,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -388,7 +445,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -411,7 +468,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -428,8 +485,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -454,7 +511,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -473,14 +530,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -537,10 +593,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -552,15 +609,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index f1e74f184f1..40aa4014b5f 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -81,6 +81,7 @@
#include "pgstat.h"
#include "port/pg_bitutils.h"
#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/spin.h"
@@ -607,9 +608,9 @@ LWLockNewTrancheId(void)
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
/* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 1a65342177d..4595f5a9676 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -22,6 +22,7 @@
#include "storage/condition_variable.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
+#include "storage/pg_shmem.h"
#include "storage/smgr.h"
#include "storage/spin.h"
#include "utils/relcache.h"
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index e0f5f92e947..c0439f2206b 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index b99ebc9e86f..138078c29c5 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 904a336b851..5929f140236 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: 80d7f990496b1c7be61d9a00a2635b7d96b96197
--
2.45.1
[text/plain] v2-0002-Allow-placing-shared-memory-mapping-with-an-offse.patch (8.7K, ../../xuuxvlhom2tiinwwnh7r6wds74o2fkwryy6palehytuzm76l4t@3q7lszfqic3b/3-v2-0002-Allow-placing-shared-memory-mapping-with-an-offse.patch)
download | inline diff:
From 7543fcdfc8ca1a0e1c85f397eb6dddfe1426b379 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH v2 2/6] Allow placing shared memory mapping with an offset
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is first to get
the address chosen by the kernel, then apply some offset derived from
the expected upper limit. Because we base the layout on the address
chosen by the kernel, things like address space randomization should not
be a problem, since the randomization is applied to the mmap base, which
is one per process. The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
[...free space...]
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
This approach do not impact the actual memory usage as reported by the kernel.
Here is the output of /proc/$PID/status for the master version with
shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages mapped in mm_struct
VmPeak: 422780 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 21248 kB
// Size of resident anonymous memory
RssAnon: 640 kB
// Size of resident file mappings
RssFile: 9728 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 10880 kB
Here is the same for the patch with the shared mapping placed at
an offset 10 GB:
VmPeak: 1102844 kB
VmRSS: 21376 kB
RssAnon: 640 kB
RssFile: 9856 kB
RssShmem: 10880 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
18219008 (~17 MB)
Note that currently the implementation makes assumptions about the upper limit.
Ideally it should be based on the maximum available memory.
---
src/backend/port/sysv_shmem.c | 120 +++++++++++++++++++++++++++++++++-
1 file changed, 119 insertions(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 843b1b3220f..62f01d8218a 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,63 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * By letting Linux to chose a mapping address, it will pick up the lowest
+ * possible address that do not clash with any other mappings, which will be
+ * right before locales in the example above. This information (maximum allowed
+ * size of mappings and the lowest mapping address) is enough to place every
+ * mapping as follow:
+ *
+ * - Take the lowest mapping address, which we call later the probe address.
+ * - Substract the offset of the previous mapping.
+ * - Substract the maximum allowed size for the current mapping from the
+ * address.
+ * - Place the mapping by the resulting address.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ */
+Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+ 0, /* MAIN_SHMEM_SLOT */
+};
+
+/* Remembers offset of the last mapping from the probe address */
+static Size last_offset = 0;
+
+/*
+ * Size of the mapping, which will be used to calculate anonymous mapping
+ * address. It should not be too small, otherwise there is a chance the probe
+ * mapping will be created between other mappings, leaving no room extending
+ * it. But it should not be too large either, in case if there are limitations
+ * on the mapping size. Current value is the default shared_buffers.
+ */
+#define PROBE_MAPPING_SIZE (Size) 128 * 1024 * 1024
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -673,13 +730,74 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
+ void *probe = NULL;
+
/*
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+
+ /*
+ * Try to create mapping at an address, which will allow to extend it
+ * later:
+ *
+ * - First create the temporary probe mapping of a fixed size and let
+ * kernel to place it at address of its choice. By the virtue of the
+ * probe mapping size we expect it to be located at the lowest
+ * possible address, expecting some non mapped space above.
+ *
+ * - Unmap the probe mapping, remember the address.
+ *
+ * - Create an actual anonymous mapping at that address with the
+ * offset. The offset is calculated in such a way to allow growing
+ * the mapping withing certain boundaries. For this mapping we use
+ * MAP_FIXED_NOREPLACE, which will error out with EEXIST if there is
+ * any mapping clash.
+ *
+ * - If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized
+ * without a restart.
+ */
+ probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
+
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: probe mmap(%zu) failed: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
+ else
+ {
+ Size offset = last_offset + SHMEM_EXTRA_SIZE_LIMIT[next_free_segment] + allocsize;
+ last_offset = offset;
+
+ munmap(probe, PROBE_MAPPING_SIZE);
+
+ ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ mmap_errno = errno;
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m",
+ MappingName(mapping->shmem_segment), allocsize, probe - offset);
+ }
+
+ }
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ /*
+ * Fallback to the portable way of creating a mapping.
+ */
+ allocsize = mapping->shmem_size;
+
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
}
--
2.45.1
[text/plain] v2-0003-Introduce-multiple-shmem-segments-for-shared-buff.patch (11.1K, ../../xuuxvlhom2tiinwwnh7r6wds74o2fkwryy6palehytuzm76l4t@3q7lszfqic3b/4-v2-0003-Introduce-multiple-shmem-segments-for-shared-buff.patch)
download | inline diff:
From d7af86299878acb73019ac699ef57c120199f1ee Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Mon, 24 Feb 2025 20:08:28 +0100
Subject: [PATCH v2 3/6] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 19 ++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 5 +-
src/backend/storage/buffer/freelist.c | 4 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 99 insertions(+), 36 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 62f01d8218a..59aa67cb135 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -149,8 +149,13 @@ static int next_free_segment = 0;
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
*/
-Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
- 0, /* MAIN_SHMEM_SLOT */
+Size SHMEM_EXTRA_SIZE_LIMIT[6] = {
+ 0, /* MAIN_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 1024 * 10, /* BUFFERS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 1024 * 1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* STRATEGY_SHMEM_SEGMENT */
};
/* Remembers offset of the last mapping from the probe address */
@@ -179,6 +184,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1f8e03190..f5b9290a640 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -61,7 +61,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -73,22 +76,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -98,8 +101,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -153,33 +157,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..ac449954dab 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -59,10 +59,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 336715b6c63..4919a92f2be 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -491,9 +491,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 4f6c707c204..68778522591 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -112,7 +112,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 7c1e4316dde..bb7fe02e243 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -297,7 +297,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 138078c29c5..ba0192baf95 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -105,7 +105,29 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.45.1
[text/plain] v2-0004-Introduce-pending-flag-for-GUC-assign-hooks.patch (11.9K, ../../xuuxvlhom2tiinwwnh7r6wds74o2fkwryy6palehytuzm76l4t@3q7lszfqic3b/5-v2-0004-Introduce-pending-flag-for-GUC-assign-hooks.patch)
download | inline diff:
From 0173967e8b0fd6c23b158c34b92651fc37ab7660 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 19 Feb 2025 17:45:40 +0100
Subject: [PATCH v2 4/6] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 2 +-
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 16 ++++----
8 files changed, 57 insertions(+), 36 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index f9bf5ba7509..ff82ba0a53d 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2188,7 +2188,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 4ad6e236d69..f24c2a0d252 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
#ifdef USE_PREFETCH
/*
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index 61ea3722ae2..cdf21847d7e 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1949,7 +1949,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1982,7 +1982,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2005,7 +2005,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2028,7 +2028,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 1149d89d7a1..13fb8c31702 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3555,7 +3555,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 12192445218..bab1c5d08f6 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 1233e07d7da..ce9f258100d 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 951451a9765..3e380f29e5a 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,12 +81,12 @@ extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
extern bool check_maintenance_io_concurrency(int *newval, void **extra,
GucSource source);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -141,13 +141,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -163,7 +163,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.45.1
[text/plain] v2-0005-Allow-to-resize-shared-memory-without-restart.patch (32.9K, ../../xuuxvlhom2tiinwwnh7r6wds74o2fkwryy6palehytuzm76l4t@3q7lszfqic3b/6-v2-0005-Allow-to-resize-shared-memory-without-restart.patch)
download | inline diff:
From 78ea0efde8799445b90a70ca321e40b75fea52c9 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 20 Feb 2025 21:12:26 +0100
Subject: [PATCH v2 5/6] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a global
Barrier to coordinate backends that simultaneously change
shared_buffers, and pieces in shared memory to coordinate backends that
are too late to the party for some reason.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier. Afterwards it does shared memory
resize itself.
* All the backends, that participate in ProcSignal mechanism,
recalculate shared memory size based on the new NBuffers and extend it
using mremap.
* When finished, a backend waits on a global ShmemControl barrier,
untill all backends will be finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all backends are using new value, one backend will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f5a2bd04000-7f5a32e52000 /dev/zero (deleted)
7f5a39252000-7f5a4030e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d7ba000 /dev/zero (deleted)
7f5a53bba000-7f5a5ad26000 /dev/zero (deleted)
7f5a9ad26000-7f5aa9d94000 /dev/zero (deleted)
^ buffers mapping, ~240 MB
7f5d29d94000-7f5d30e00000 /dev/zero (deleted)
-- 512 MB
7f5a2bd04000-7f5a33274000 /dev/zero (deleted)
7f5a39252000-7f5a4057e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d9fa000 /dev/zero (deleted)
7f5a53bba000-7f5a5b1a6000 /dev/zero (deleted)
7f5a9ad26000-7f5ac1f14000 /dev/zero (deleted)
^ buffers mapping, ~625 MB
7f5d29d94000-7f5d30f80000 /dev/zero (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it reads something.
Note, that mremap is Linux specific, thus the implementation not very
portable.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 300 ++++++++++++++++++
src/backend/postmaster/postmaster.c | 15 +
src/backend/storage/buffer/buf_init.c | 152 ++++++++-
src/backend/storage/ipc/ipci.c | 11 +
src/backend/storage/ipc/procsignal.c | 45 +++
src/backend/storage/ipc/shmem.c | 14 +-
src/backend/tcop/postgres.c | 15 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/storage/bufmgr.h | 1 +
src/include/storage/ipc.h | 2 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 24 ++
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
15 files changed, 577 insertions(+), 12 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 59aa67cb135..35a8ff92175 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,17 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/procsignal.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -105,6 +109,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -859,6 +870,274 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * mremap, which will fail if is not sufficient space to expand the mapping.
+ * When finished, based on the new and old values initialize new buffer blocks
+ * if any.
+ *
+ * If reinitializing took place, as the last step this function broadcasts
+ * NSharedBuffers to it's new value, allowing any other backends to rely on
+ * this new value and skip buffers reinitialization.
+ */
+static bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+
+ /*
+ * Fail hard if faced any issues. In theory we could try to handle this
+ * more gracefully and proceed with shared memory as before, but some
+ * other backends might have succeeded and have different size. If we
+ * would like to go this way, to be consistent we would need to
+ * synchronize again, and it's not clear if it's worth the effort.
+ */
+ if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not resize shared memory %p to %d (%zu): %m",
+ m->shmem, NBuffers, m->shmem_size)));
+ else
+ {
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ ResizeBufferPool(NBuffersOld, true);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+ else
+ ResizeBufferPool(NBuffersOld, false);
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Do the resize and make sure to wait on
+ * the provided barrier until all simultaneously participating backends finish
+ * resizing as well, otherwise we face danger of inconsistency between
+ * backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * After attaching to the barrier we could be in any of states:
+ *
+ * - Initial SHMEM_RESIZE_REQUESTED, nothing has been done yet
+ * - SHMEM_RESIZE_START, some of the backends have started to resize
+ * - SHMEM_RESIZE_DONE, participating backends have finished resizing
+ * - SHMEM_RESIZE_REQUESTED after the reset, the shared memory was already
+ * resized
+ *
+ * The first three states take place while the actual resize is in
+ * progress, and all we need to do is join and proceed with resizing. This
+ * way all simultaneously participating backends will remap and wait until
+ * one of them initialize new buffers.
+ *
+ * The last state happens when we are too late and everything is already
+ * done. In that case proceed as well, relying on AnonymousShmemResize not
+ * reinitialize anything since the NSharedBuffers is already broadcasted.
+ */
+ BarrierAttach(barrier);
+
+ /* First phase means the resize has begun, SHMEM_RESIZE_START */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
+
+ /* XXX: Split mremap and buffer reinitialization into two barrier phases */
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ /* Allow the last backend to reset the barrier */
+ if (BarrierArriveAndDetach(barrier))
+ ResetShmemBarrier();
+
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ elog(DEBUG1, "Set pending signal");
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Coordinate all existing processes to make sure they all will have consistent
+ * view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ * There might still be some backends, that are sleeping or for some
+ * other reason not doing any work yet and have old NBuffers -- but as
+ * soon as they will get some time slice, they will acquire the new
+ * value.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE);
+
+ AnonymousShmemResize();
+
+ /*
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier, and there is no sequential logic
+ * we have to perform afterwards. In fact even if we would like to wait
+ * here, it wouldn't be possible -- we're in the postmaster, without any
+ * waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1174,3 +1453,24 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier(int phase)
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ if (BarrierPhase(barrier) == phase)
+ {
+ ereport(LOG,
+ (errmsg("ProcSignal barrier is in phase %d, waiting", phase)));
+ BarrierAttach(barrier);
+ BarrierArriveAndWait(barrier, 0);
+ BarrierDetach(barrier);
+ }
+}
+
+void
+ResetShmemBarrier(void)
+{
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+}
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index bb22b13adef..f3e508141b2 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -418,6 +418,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1680,6 +1681,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2026,6 +2030,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index f5b9290a640..b7de0ab6b0d 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -23,6 +23,41 @@ ConditionVariableMinimallyPadded *BufferIOCVArray;
WritebackContext BackendWritebackContext;
CkptSortItem *CkptBufferIds;
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
/*
* Data Structures:
@@ -72,7 +107,19 @@ BufferManagerShmemInit(void)
bool foundBufs,
foundDescs,
foundIOCV,
- foundBufCkpt;
+ foundBufCkpt,
+ foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+ }
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -153,6 +200,109 @@ BufferManagerShmemInit(void)
&backend_flush_after);
}
+/*
+ * Reinitialize shared memory structures, which size depends on NBuffers. It's
+ * similar to InitBufferPool, but applied only to the buffers in the range
+ * between NBuffersOld and NBuffers.
+ *
+ * NBuffersOld tells what was the original value of NBuffersOld. It will be
+ * used to identify new and not yet initialized buffers.
+ *
+ * initNew flag indicates that the caller wants new buffers to be initialized.
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
+ */
+void
+ResizeBufferPool(int NBuffersOld, bool initNew)
+{
+ bool foundBufs,
+ foundDescs,
+ foundIOCV,
+ foundBufCkpt;
+ int i;
+ elog(DEBUG1, "Resizing buffer pool from %d to %d", NBuffersOld, NBuffers);
+
+ /* XXX: Only increasing of shared_buffers is supported in this function */
+ if(NBuffersOld > NBuffers)
+ return;
+
+ /* Align descriptors to a cacheline boundary. */
+ BufferDescriptors = (BufferDescPadded *)
+ ShmemInitStructInSegment("Buffer Descriptors",
+ NBuffers * sizeof(BufferDescPadded),
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+
+ /* Align condition variables to cacheline boundary. */
+ BufferIOCVArray = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
+
+ /*
+ * The array used to sort to-be-checkpointed buffer ids is located in
+ * shared memory, to avoid having to allocate significant amounts of
+ * memory at runtime. As that'd be in the middle of a checkpoint, or when
+ * the checkpointer is restarted, memory allocation failures would be
+ * painful.
+ */
+ CkptBufferIds = (CkptSortItem *)
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+
+ /* Align buffer pool on IO page size boundary. */
+ BufferBlocks = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSegment("Buffer Blocks",
+ NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
+
+ /*
+ * It's enough to only resize shmem structures, if some other backend will
+ * do initialization of new buffers for us.
+ */
+ if (!initNew)
+ return;
+
+ elog(DEBUG1, "Initialize new buffers");
+
+ /*
+ * Initialize the headers for new buffers.
+ */
+ for (i = NBuffersOld; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
+
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+
+ buf->buf_id = i;
+
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
+
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
+
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ }
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+
+ /* Init other shared buffer-management stuff */
+ StrategyInitialize(!foundDescs);
+
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+}
+
/*
* BufferManagerShmemSize
*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 68778522591..a2c635f288e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -83,6 +83,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -149,6 +152,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, InjectionPointShmemSize());
size = add_size(size, SlotSyncShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 7401b6e625e..bec0e00f901 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -108,6 +109,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -168,6 +173,42 @@ ProcSignalInit(bool cancel_key_valid, int32 cancel_key)
ProcSignalSlot *slot;
uint64 barrier_generation;
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -573,6 +614,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 389abc82519..226b38ba979 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -493,17 +493,13 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
*/
if (result->size != size)
- {
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
- }
+ result->size = size;
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 13fb8c31702..04cdd0d24d8 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4267,6 +4268,20 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /*
+ * Verify the shared barrier, if it's still active: join and wait.
+ *
+ * XXX: Any potential race condition if not a single backend has
+ * incremented the barrier phase?
+ */
+ WaitOnShmemBarrier(SHMEM_RESIZE_START);
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index e199f071628..012acb98169 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -154,6 +154,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
WAL_RECEIVER_WAIT_START "Waiting for startup process to send initial data for streaming replication."
@@ -346,6 +348,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 3cde94a1759..efdaa71c8fb 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2339,14 +2339,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index bb7fe02e243..fff80214822 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -298,6 +298,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
extern Size BufferManagerShmemSize(int);
+extern void ResizeBufferPool(int, bool);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index c0439f2206b..5f5b45c88bd 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
extern void proc_exit(int code) pg_attribute_noreturn();
extern void shmem_exit(int code);
@@ -83,5 +84,6 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index cf565452382..61e89c6e8fd 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -83,3 +83,4 @@ PG_LWLOCK(49, WALSummarizer)
PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
+PG_LWLOCK(53, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index ba0192baf95..b597df0d3a3 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,23 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -105,6 +123,12 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(int phase);
+extern void ResetShmemBarrier(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 022fd8ed933..4c9973dc2d9 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index fb39c915d76..5bf6d099808 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2671,6 +2671,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.45.1
[text/plain] v2-0006-Use-anonymous-files-to-back-shared-memory-segment.patch (6.7K, ../../xuuxvlhom2tiinwwnh7r6wds74o2fkwryy6palehytuzm76l4t@3q7lszfqic3b/7-v2-0006-Use-anonymous-files-to-back-shared-memory-segment.patch)
download | inline diff:
From e511bab55891a2d60152e913df6c20e20314e71b Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 23 Feb 2025 14:42:39 +0100
Subject: [PATCH v2 6/6] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f5a2bd04000-7f5a32e52000 rw-s 00000000 00:01 1845 /memfd:strategy (deleted)
7f5a39252000-7f5a4030e000 rw-s 00000000 00:01 1842 /memfd:checkpoint (deleted)
7f5a4670e000-7f5a4d7ba000 rw-s 00000000 00:01 1839 /memfd:iocv (deleted)
7f5a53bba000-7f5a5ad26000 rw-s 00000000 00:01 1836 /memfd:descriptors (deleted)
7f5a9ad26000-7f5aa9d94000 rw-s 00000000 00:01 1833 /memfd:buffers (deleted)
7f5d29d94000-7f5d30e00000 rw-s 00000000 00:01 1830 /memfd:main (deleted)
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 46 +++++++++++++++++++++++++++++------
src/include/portability/mem.h | 2 +-
2 files changed, 39 insertions(+), 9 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 35a8ff92175..8864866f26c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -105,6 +105,7 @@ typedef struct AnonymousMapping
void *shmem; /* Pointer to the start of the mapped memory */
void *seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -125,7 +126,7 @@ static int next_free_segment = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -152,9 +153,9 @@ static int next_free_segment = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* [...free space...]
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* [...free space...]
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -717,6 +718,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment), 0);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
@@ -734,8 +747,13 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ /*
+ * Do not use an anonymous file here yet. When adding it, do not forget
+ * to use ftruncate and flags MFD_HUGETLB & MFD_HUGE_2MB/MFD_HUGE_1GB
+ * in memfd_create.
+ */
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
@@ -771,7 +789,8 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* - First create the temporary probe mapping of a fixed size and let
* kernel to place it at address of its choice. By the virtue of the
* probe mapping size we expect it to be located at the lowest
- * possible address, expecting some non mapped space above.
+ * possible address, expecting some non mapped space above. The probe
+ * is does not need to be backed by an anonymous file.
*
* - Unmap the probe mapping, remember the address.
*
@@ -786,7 +805,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* without a restart.
*/
probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS, -1, 0);
if (probe == MAP_FAILED)
{
@@ -802,8 +821,14 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
munmap(probe, PROBE_MAPPING_SIZE);
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
{
@@ -822,8 +847,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
*/
allocsize = mapping->shmem_size;
+ /* Specify the segment file size using allocsize. */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
@@ -917,6 +945,8 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ ftruncate(m->segment_fd, new_size);
/*
* Fail hard if faced any issues. In theory we could try to handle this
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
--
2.45.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-02-27 08:28 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-02-27 08:28 UTC (permalink / raw)
To: pgsql-hackers; +Cc: Robert Haas <robertmhaas@gmail.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
> On Tue, Feb 25, 2025 at 10:52:05AM GMT, Dmitry Dolgov wrote:
> > On Fri, Oct 18, 2024 at 09:21:19PM GMT, Dmitry Dolgov wrote:
> > TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
> > changing shared memory mapping layout. Any feedback is appreciated.
>
> Hi,
>
> Here is a new version of the patch, which contains a proposal about how to
> coordinate shared memory resizing between backends. The rest is more or less
> the same, a feedback about coordination is appreciated. It's a lot to read, but
> the main difference is about:
Just one note, there are still couple of compilation warnings in the
code, which I haven't addressed yet. Those will go away in the next
version.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-27 08:28 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-02-28 11:52 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-02-28 11:52 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Thu, Feb 27, 2025 at 1:58 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Tue, Feb 25, 2025 at 10:52:05AM GMT, Dmitry Dolgov wrote:
> > > On Fri, Oct 18, 2024 at 09:21:19PM GMT, Dmitry Dolgov wrote:
> > > TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
> > > changing shared memory mapping layout. Any feedback is appreciated.
> >
> > Hi,
> >
> > Here is a new version of the patch, which contains a proposal about how to
> > coordinate shared memory resizing between backends. The rest is more or less
> > the same, a feedback about coordination is appreciated. It's a lot to read, but
> > the main difference is about:
>
> Just one note, there are still couple of compilation warnings in the
> code, which I haven't addressed yet. Those will go away in the next
> version.
PFA the patchset which implements shrinking shared buffers.
0001-0006 are same as the previous patchset
0007 fixes compilation warnings from previous patches - I think those
should be absorbed into their respective patches
0008 adds TODOs that need some code changes or at least need some
consideration. Some of them might point to the causes of Assertion
failures seen with this patch set.
0009 adds WIP support for shrinking shared buffers - I think this
should be absorbed into 0005
0010 WIP fix for Assertion failures seen from BgBufferSync() - I am
still investigating those.
I am using the attached script to shake the patch well. It runs
pgbench and concurrently resizes the shared_buffers. I am seeing
Assertion failures when running the script in both cases, expanding
and shrinking the buffers. I am investigating "failed
Assert("strategy_delta >= 0")," next.
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] 0004-Introduce-pending-flag-for-GUC-assign-hooks-20250228.patch (11.9K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/2-0004-Introduce-pending-flag-for-GUC-assign-hooks-20250228.patch)
download | inline diff:
From 7de04680820a8aa3de7b13e4a0f33471a51524fe Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 19 Feb 2025 17:45:40 +0100
Subject: [PATCH 04/11] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 2 +-
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 16 ++++----
8 files changed, 57 insertions(+), 36 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 75d5554c77c..cf79df49cb7 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2267,7 +2267,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 4ad6e236d69..f24c2a0d252 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
#ifdef USE_PREFETCH
/*
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index bddd6465de2..3e6d83c8775 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1949,7 +1949,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1982,7 +1982,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2005,7 +2005,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2028,7 +2028,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 1149d89d7a1..13fb8c31702 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3555,7 +3555,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 12192445218..bab1c5d08f6 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 1233e07d7da..ce9f258100d 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 87999218d68..1fb4c519c7f 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,12 +81,12 @@ extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
extern bool check_maintenance_io_concurrency(int *newval, void **extra,
GucSource source);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -141,13 +141,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -163,7 +163,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.34.1
[text/x-patch] 0003-Introduce-multiple-shmem-segments-for-share-20250228.patch (11.1K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/3-0003-Introduce-multiple-shmem-segments-for-share-20250228.patch)
download | inline diff:
From df9ed8581fbe8cf4cf1f6bb53e5ced94c7bfe19d Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Mon, 24 Feb 2025 20:08:28 +0100
Subject: [PATCH 03/11] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 19 ++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 5 +-
src/backend/storage/buffer/freelist.c | 4 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 99 insertions(+), 36 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 62f01d8218a..59aa67cb135 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -149,8 +149,13 @@ static int next_free_segment = 0;
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
*/
-Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
- 0, /* MAIN_SHMEM_SLOT */
+Size SHMEM_EXTRA_SIZE_LIMIT[6] = {
+ 0, /* MAIN_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 1024 * 10, /* BUFFERS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 1024 * 1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ (Size) 1024 * 1024 * 100, /* STRATEGY_SHMEM_SEGMENT */
};
/* Remembers offset of the last mapping from the probe address */
@@ -179,6 +184,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1f8e03190..f5b9290a640 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -61,7 +61,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -73,22 +76,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -98,8 +101,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -153,33 +157,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..ac449954dab 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -59,10 +59,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 336715b6c63..4919a92f2be 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -491,9 +491,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 4f6c707c204..68778522591 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -112,7 +112,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 7c1e4316dde..bb7fe02e243 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -297,7 +297,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 138078c29c5..ba0192baf95 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -105,7 +105,29 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.34.1
[text/x-patch] 0005-Allow-to-resize-shared-memory-without-resta-20250228.patch (32.9K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/4-0005-Allow-to-resize-shared-memory-without-resta-20250228.patch)
download | inline diff:
From 71b4321af793a8a5682d5f9b1490fb3718e74759 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 20 Feb 2025 21:12:26 +0100
Subject: [PATCH 05/11] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a global
Barrier to coordinate backends that simultaneously change
shared_buffers, and pieces in shared memory to coordinate backends that
are too late to the party for some reason.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier. Afterwards it does shared memory
resize itself.
* All the backends, that participate in ProcSignal mechanism,
recalculate shared memory size based on the new NBuffers and extend it
using mremap.
* When finished, a backend waits on a global ShmemControl barrier,
untill all backends will be finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all backends are using new value, one backend will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f5a2bd04000-7f5a32e52000 /dev/zero (deleted)
7f5a39252000-7f5a4030e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d7ba000 /dev/zero (deleted)
7f5a53bba000-7f5a5ad26000 /dev/zero (deleted)
7f5a9ad26000-7f5aa9d94000 /dev/zero (deleted)
^ buffers mapping, ~240 MB
7f5d29d94000-7f5d30e00000 /dev/zero (deleted)
-- 512 MB
7f5a2bd04000-7f5a33274000 /dev/zero (deleted)
7f5a39252000-7f5a4057e000 /dev/zero (deleted)
7f5a4670e000-7f5a4d9fa000 /dev/zero (deleted)
7f5a53bba000-7f5a5b1a6000 /dev/zero (deleted)
7f5a9ad26000-7f5ac1f14000 /dev/zero (deleted)
^ buffers mapping, ~625 MB
7f5d29d94000-7f5d30f80000 /dev/zero (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it reads something.
Note, that mremap is Linux specific, thus the implementation not very
portable.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 300 ++++++++++++++++++
src/backend/postmaster/postmaster.c | 15 +
src/backend/storage/buffer/buf_init.c | 152 ++++++++-
src/backend/storage/ipc/ipci.c | 11 +
src/backend/storage/ipc/procsignal.c | 45 +++
src/backend/storage/ipc/shmem.c | 14 +-
src/backend/tcop/postgres.c | 15 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/storage/bufmgr.h | 1 +
src/include/storage/ipc.h | 2 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 24 ++
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
15 files changed, 577 insertions(+), 12 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 59aa67cb135..35a8ff92175 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,17 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/procsignal.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -105,6 +109,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -859,6 +870,274 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * mremap, which will fail if is not sufficient space to expand the mapping.
+ * When finished, based on the new and old values initialize new buffer blocks
+ * if any.
+ *
+ * If reinitializing took place, as the last step this function broadcasts
+ * NSharedBuffers to it's new value, allowing any other backends to rely on
+ * this new value and skip buffers reinitialization.
+ */
+static bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+
+ /*
+ * Fail hard if faced any issues. In theory we could try to handle this
+ * more gracefully and proceed with shared memory as before, but some
+ * other backends might have succeeded and have different size. If we
+ * would like to go this way, to be consistent we would need to
+ * synchronize again, and it's not clear if it's worth the effort.
+ */
+ if (mremap(m->shmem, m->shmem_size, new_size, 0) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not resize shared memory %p to %d (%zu): %m",
+ m->shmem, NBuffers, m->shmem_size)));
+ else
+ {
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ ResizeBufferPool(NBuffersOld, true);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+ else
+ ResizeBufferPool(NBuffersOld, false);
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Do the resize and make sure to wait on
+ * the provided barrier until all simultaneously participating backends finish
+ * resizing as well, otherwise we face danger of inconsistency between
+ * backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * After attaching to the barrier we could be in any of states:
+ *
+ * - Initial SHMEM_RESIZE_REQUESTED, nothing has been done yet
+ * - SHMEM_RESIZE_START, some of the backends have started to resize
+ * - SHMEM_RESIZE_DONE, participating backends have finished resizing
+ * - SHMEM_RESIZE_REQUESTED after the reset, the shared memory was already
+ * resized
+ *
+ * The first three states take place while the actual resize is in
+ * progress, and all we need to do is join and proceed with resizing. This
+ * way all simultaneously participating backends will remap and wait until
+ * one of them initialize new buffers.
+ *
+ * The last state happens when we are too late and everything is already
+ * done. In that case proceed as well, relying on AnonymousShmemResize not
+ * reinitialize anything since the NSharedBuffers is already broadcasted.
+ */
+ BarrierAttach(barrier);
+
+ /* First phase means the resize has begun, SHMEM_RESIZE_START */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
+
+ /* XXX: Split mremap and buffer reinitialization into two barrier phases */
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ /* Allow the last backend to reset the barrier */
+ if (BarrierArriveAndDetach(barrier))
+ ResetShmemBarrier();
+
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ elog(DEBUG1, "Set pending signal");
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Coordinate all existing processes to make sure they all will have consistent
+ * view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ * There might still be some backends, that are sleeping or for some
+ * other reason not doing any work yet and have old NBuffers -- but as
+ * soon as they will get some time slice, they will acquire the new
+ * value.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE);
+
+ AnonymousShmemResize();
+
+ /*
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier, and there is no sequential logic
+ * we have to perform afterwards. In fact even if we would like to wait
+ * here, it wouldn't be possible -- we're in the postmaster, without any
+ * waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1174,3 +1453,24 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier(int phase)
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ if (BarrierPhase(barrier) == phase)
+ {
+ ereport(LOG,
+ (errmsg("ProcSignal barrier is in phase %d, waiting", phase)));
+ BarrierAttach(barrier);
+ BarrierArriveAndWait(barrier, 0);
+ BarrierDetach(barrier);
+ }
+}
+
+void
+ResetShmemBarrier(void)
+{
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+}
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index bb22b13adef..f3e508141b2 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -418,6 +418,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1680,6 +1681,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2026,6 +2030,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index f5b9290a640..b7de0ab6b0d 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -23,6 +23,41 @@ ConditionVariableMinimallyPadded *BufferIOCVArray;
WritebackContext BackendWritebackContext;
CkptSortItem *CkptBufferIds;
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
/*
* Data Structures:
@@ -72,7 +107,19 @@ BufferManagerShmemInit(void)
bool foundBufs,
foundDescs,
foundIOCV,
- foundBufCkpt;
+ foundBufCkpt,
+ foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+ }
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -153,6 +200,109 @@ BufferManagerShmemInit(void)
&backend_flush_after);
}
+/*
+ * Reinitialize shared memory structures, which size depends on NBuffers. It's
+ * similar to InitBufferPool, but applied only to the buffers in the range
+ * between NBuffersOld and NBuffers.
+ *
+ * NBuffersOld tells what was the original value of NBuffersOld. It will be
+ * used to identify new and not yet initialized buffers.
+ *
+ * initNew flag indicates that the caller wants new buffers to be initialized.
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
+ */
+void
+ResizeBufferPool(int NBuffersOld, bool initNew)
+{
+ bool foundBufs,
+ foundDescs,
+ foundIOCV,
+ foundBufCkpt;
+ int i;
+ elog(DEBUG1, "Resizing buffer pool from %d to %d", NBuffersOld, NBuffers);
+
+ /* XXX: Only increasing of shared_buffers is supported in this function */
+ if(NBuffersOld > NBuffers)
+ return;
+
+ /* Align descriptors to a cacheline boundary. */
+ BufferDescriptors = (BufferDescPadded *)
+ ShmemInitStructInSegment("Buffer Descriptors",
+ NBuffers * sizeof(BufferDescPadded),
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+
+ /* Align condition variables to cacheline boundary. */
+ BufferIOCVArray = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
+
+ /*
+ * The array used to sort to-be-checkpointed buffer ids is located in
+ * shared memory, to avoid having to allocate significant amounts of
+ * memory at runtime. As that'd be in the middle of a checkpoint, or when
+ * the checkpointer is restarted, memory allocation failures would be
+ * painful.
+ */
+ CkptBufferIds = (CkptSortItem *)
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+
+ /* Align buffer pool on IO page size boundary. */
+ BufferBlocks = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSegment("Buffer Blocks",
+ NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
+
+ /*
+ * It's enough to only resize shmem structures, if some other backend will
+ * do initialization of new buffers for us.
+ */
+ if (!initNew)
+ return;
+
+ elog(DEBUG1, "Initialize new buffers");
+
+ /*
+ * Initialize the headers for new buffers.
+ */
+ for (i = NBuffersOld; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
+
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+
+ buf->buf_id = i;
+
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
+
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
+
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ }
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+
+ /* Init other shared buffer-management stuff */
+ StrategyInitialize(!foundDescs);
+
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+}
+
/*
* BufferManagerShmemSize
*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 68778522591..a2c635f288e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -83,6 +83,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -149,6 +152,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, InjectionPointShmemSize());
size = add_size(size, SlotSyncShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 7401b6e625e..bec0e00f901 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -108,6 +109,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -168,6 +173,42 @@ ProcSignalInit(bool cancel_key_valid, int32 cancel_key)
ProcSignalSlot *slot;
uint64 barrier_generation;
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -573,6 +614,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 389abc82519..226b38ba979 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -493,17 +493,13 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
*/
if (result->size != size)
- {
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
- }
+ result->size = size;
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 13fb8c31702..04cdd0d24d8 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4267,6 +4268,20 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /*
+ * Verify the shared barrier, if it's still active: join and wait.
+ *
+ * XXX: Any potential race condition if not a single backend has
+ * incremented the barrier phase?
+ */
+ WaitOnShmemBarrier(SHMEM_RESIZE_START);
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index ccf73781d81..947f13cb1fa 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -154,6 +154,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
@@ -346,6 +348,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 42728189322..01faf705582 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2329,14 +2329,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index bb7fe02e243..fff80214822 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -298,6 +298,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
extern Size BufferManagerShmemSize(int);
+extern void ResizeBufferPool(int, bool);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index c0439f2206b..5f5b45c88bd 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
extern void proc_exit(int code) pg_attribute_noreturn();
extern void shmem_exit(int code);
@@ -83,5 +84,6 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index ff897515769..e5d8cd183cf 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -83,3 +83,4 @@ PG_LWLOCK(49, WALSummarizer)
PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
+PG_LWLOCK(53, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index ba0192baf95..b597df0d3a3 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,23 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -105,6 +123,12 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(int phase);
+extern void ResetShmemBarrier(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 022fd8ed933..4c9973dc2d9 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index b6c170ac249..00f6b9d7d7d 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2668,6 +2668,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.34.1
[text/x-patch] 0002-Allow-placing-shared-memory-mapping-with-an-20250228.patch (8.7K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/5-0002-Allow-placing-shared-memory-mapping-with-an-20250228.patch)
download | inline diff:
From 436475da9bbb9f8f370332931016f4a5cbd9d393 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH 02/11] Allow placing shared memory mapping with an offset
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is first to get
the address chosen by the kernel, then apply some offset derived from
the expected upper limit. Because we base the layout on the address
chosen by the kernel, things like address space randomization should not
be a problem, since the randomization is applied to the mmap base, which
is one per process. The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
[...free space...]
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
This approach do not impact the actual memory usage as reported by the kernel.
Here is the output of /proc/$PID/status for the master version with
shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages mapped in mm_struct
VmPeak: 422780 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 21248 kB
// Size of resident anonymous memory
RssAnon: 640 kB
// Size of resident file mappings
RssFile: 9728 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 10880 kB
Here is the same for the patch with the shared mapping placed at
an offset 10 GB:
VmPeak: 1102844 kB
VmRSS: 21376 kB
RssAnon: 640 kB
RssFile: 9856 kB
RssShmem: 10880 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
18219008 (~17 MB)
Note that currently the implementation makes assumptions about the upper limit.
Ideally it should be based on the maximum available memory.
---
src/backend/port/sysv_shmem.c | 120 +++++++++++++++++++++++++++++++++-
1 file changed, 119 insertions(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 843b1b3220f..62f01d8218a 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,63 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * By letting Linux to chose a mapping address, it will pick up the lowest
+ * possible address that do not clash with any other mappings, which will be
+ * right before locales in the example above. This information (maximum allowed
+ * size of mappings and the lowest mapping address) is enough to place every
+ * mapping as follow:
+ *
+ * - Take the lowest mapping address, which we call later the probe address.
+ * - Substract the offset of the previous mapping.
+ * - Substract the maximum allowed size for the current mapping from the
+ * address.
+ * - Place the mapping by the resulting address.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * [...free space...]
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ */
+Size SHMEM_EXTRA_SIZE_LIMIT[1] = {
+ 0, /* MAIN_SHMEM_SLOT */
+};
+
+/* Remembers offset of the last mapping from the probe address */
+static Size last_offset = 0;
+
+/*
+ * Size of the mapping, which will be used to calculate anonymous mapping
+ * address. It should not be too small, otherwise there is a chance the probe
+ * mapping will be created between other mappings, leaving no room extending
+ * it. But it should not be too large either, in case if there are limitations
+ * on the mapping size. Current value is the default shared_buffers.
+ */
+#define PROBE_MAPPING_SIZE (Size) 128 * 1024 * 1024
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -673,13 +730,74 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
+ void *probe = NULL;
+
/*
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+
+ /*
+ * Try to create mapping at an address, which will allow to extend it
+ * later:
+ *
+ * - First create the temporary probe mapping of a fixed size and let
+ * kernel to place it at address of its choice. By the virtue of the
+ * probe mapping size we expect it to be located at the lowest
+ * possible address, expecting some non mapped space above.
+ *
+ * - Unmap the probe mapping, remember the address.
+ *
+ * - Create an actual anonymous mapping at that address with the
+ * offset. The offset is calculated in such a way to allow growing
+ * the mapping withing certain boundaries. For this mapping we use
+ * MAP_FIXED_NOREPLACE, which will error out with EEXIST if there is
+ * any mapping clash.
+ *
+ * - If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized
+ * without a restart.
+ */
+ probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
+
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: probe mmap(%zu) failed: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
+ else
+ {
+ Size offset = last_offset + SHMEM_EXTRA_SIZE_LIMIT[next_free_segment] + allocsize;
+ last_offset = offset;
+
+ munmap(probe, PROBE_MAPPING_SIZE);
+
+ ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ mmap_errno = errno;
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m",
+ MappingName(mapping->shmem_segment), allocsize, probe - offset);
+ }
+
+ }
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ /*
+ * Fallback to the portable way of creating a mapping.
+ */
+ allocsize = mapping->shmem_size;
+
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
}
--
2.34.1
[text/x-patch] 0001-Allow-to-use-multiple-shared-memory-mapping-20250228.patch (30.1K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/6-0001-Allow-to-use-multiple-shared-memory-mapping-20250228.patch)
download | inline diff:
From 4fb360f80ddb353d013d7b8e999b88522a36a604 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 19 Feb 2025 17:43:13 +0100
Subject: [PATCH 01/11] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 ++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 141 +++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 5 +-
src/include/storage/buf_internals.h | 1 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
13 files changed, 272 insertions(+), 124 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index f7c8638aec5..b6301463ac7 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -313,7 +313,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -334,7 +334,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..843b1b3220f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ void *shmem; /* Pointer to the start of the mapped memory */
+ void *seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index e4d5b944e12..9d526eb43fd 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 174eed70367..4f6c707c204 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -85,7 +85,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -204,33 +204,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -360,7 +365,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 895a43fb39e..389abc82519 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -75,19 +75,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/*
@@ -96,9 +96,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -109,7 +117,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -118,9 +132,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -145,11 +159,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -179,6 +199,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -198,22 +224,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -231,6 +257,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -241,19 +273,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -268,7 +300,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -329,6 +367,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -345,9 +395,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -380,6 +430,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -388,7 +445,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -411,7 +468,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -428,8 +485,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -454,7 +511,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -473,14 +530,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -537,10 +593,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -552,15 +609,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index f1e74f184f1..40aa4014b5f 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -81,6 +81,7 @@
#include "pgstat.h"
#include "port/pg_bitutils.h"
#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/spin.h"
@@ -607,9 +608,9 @@ LWLockNewTrancheId(void)
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
/* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 1a65342177d..4595f5a9676 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -22,6 +22,7 @@
#include "storage/condition_variable.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
+#include "storage/pg_shmem.h"
#include "storage/smgr.h"
#include "storage/spin.h"
#include "utils/relcache.h"
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index e0f5f92e947..c0439f2206b 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index b99ebc9e86f..138078c29c5 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 904a336b851..5929f140236 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: eaf502747bacee0122668eb1ba3979f86b8d8342
--
2.34.1
[text/x-patch] 0006-Use-anonymous-files-to-back-shared-memory-s-20250228.patch (6.7K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/7-0006-Use-anonymous-files-to-back-shared-memory-s-20250228.patch)
download | inline diff:
From be911372e4de4b2b98699512007ff8055dbea2f2 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 23 Feb 2025 14:42:39 +0100
Subject: [PATCH 06/11] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f5a2bd04000-7f5a32e52000 rw-s 00000000 00:01 1845 /memfd:strategy (deleted)
7f5a39252000-7f5a4030e000 rw-s 00000000 00:01 1842 /memfd:checkpoint (deleted)
7f5a4670e000-7f5a4d7ba000 rw-s 00000000 00:01 1839 /memfd:iocv (deleted)
7f5a53bba000-7f5a5ad26000 rw-s 00000000 00:01 1836 /memfd:descriptors (deleted)
7f5a9ad26000-7f5aa9d94000 rw-s 00000000 00:01 1833 /memfd:buffers (deleted)
7f5d29d94000-7f5d30e00000 rw-s 00000000 00:01 1830 /memfd:main (deleted)
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 46 +++++++++++++++++++++++++++++------
src/include/portability/mem.h | 2 +-
2 files changed, 39 insertions(+), 9 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 35a8ff92175..8864866f26c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -105,6 +105,7 @@ typedef struct AnonymousMapping
void *shmem; /* Pointer to the start of the mapped memory */
void *seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -125,7 +126,7 @@ static int next_free_segment = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -152,9 +153,9 @@ static int next_free_segment = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* [...free space...]
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* [...free space...]
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -717,6 +718,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment), 0);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
@@ -734,8 +747,13 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ /*
+ * Do not use an anonymous file here yet. When adding it, do not forget
+ * to use ftruncate and flags MFD_HUGETLB & MFD_HUGE_2MB/MFD_HUGE_1GB
+ * in memfd_create.
+ */
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
@@ -771,7 +789,8 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* - First create the temporary probe mapping of a fixed size and let
* kernel to place it at address of its choice. By the virtue of the
* probe mapping size we expect it to be located at the lowest
- * possible address, expecting some non mapped space above.
+ * possible address, expecting some non mapped space above. The probe
+ * is does not need to be backed by an anonymous file.
*
* - Unmap the probe mapping, remember the address.
*
@@ -786,7 +805,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* without a restart.
*/
probe = mmap(NULL, PROBE_MAPPING_SIZE, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS | MAP_ANONYMOUS, -1, 0);
if (probe == MAP_FAILED)
{
@@ -802,8 +821,14 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
munmap(probe, PROBE_MAPPING_SIZE);
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, -1, 0);
+ PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
{
@@ -822,8 +847,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
*/
allocsize = mapping->shmem_size;
+ /* Specify the segment file size using allocsize. */
+ ftruncate(mapping->segment_fd, allocsize);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
@@ -917,6 +945,8 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ ftruncate(m->segment_fd, new_size);
/*
* Fail hard if faced any issues. In theory we could try to handle this
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
--
2.34.1
[text/x-patch] 0010-WIP-Reinitialize-buffer-sync-strategy-20250228.patch (6.6K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/8-0010-WIP-Reinitialize-buffer-sync-strategy-20250228.patch)
download | inline diff:
From 67ab7fe07bd0bd9f6a1f7ebbeaa7f27e32cf0cf3 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 27 Jan 2025 15:46:23 +0530
Subject: [PATCH 10/11] WIP: Reinitialize buffer sync strategy
Resizing the shared buffers renders the saved state of BgBufferSync() invalid.
Hence reinitialize it. The state is saved in static variables inside the
function and thus can not be accessed from outside the function. Hence we add an
argument to BgBufferSync() to request the function to reset the state.
TODO: Ideally we should save this state in some global structure and add a
function to reset it.
TODO: StrategyInitialize() initializes buffer lookup table but
StrategyReInitialize() doesn't. Where should we reinitialize the buffer
lookup table?
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 6 +++++
src/backend/postmaster/bgwriter.c | 2 +-
src/backend/storage/buffer/buf_init.c | 2 +-
src/backend/storage/buffer/bufmgr.c | 20 +++++++++++++++--
src/backend/storage/buffer/freelist.c | 32 +++++++++++++++++++++++++++
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 2 +-
7 files changed, 60 insertions(+), 5 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index f084a0747ff..66d6d4b4333 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1007,7 +1007,13 @@ AnonymousShmemResize(void)
* backend who can not lock the LWLock conditionally won't resize the
* buffers.
*/
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /* If we are bgwriter wipe out the previous state and start anew. */
+ BgBufferSync(NULL, true);
}
+ }
return true;
}
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index 3eff5dc6f0e..3cd421b23c0 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -231,7 +231,7 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
/*
* Do one cycle of dirty-buffer writing.
*/
- can_hibernate = BgBufferSync(&wb_context);
+ can_hibernate = BgBufferSync(&wb_context, false);
/* Report pending statistics to the cumulative stats system */
pgstat_report_bgwriter();
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index d8139a899bb..6b6f6e6ae08 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -305,7 +305,7 @@ ResizeBufferPool(int NBuffersOld, bool initNew)
GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
- StrategyInitialize(!foundDescs);
+ StrategyReInitialize();
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b6bec73e6b7..a7948e75e76 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3179,7 +3179,10 @@ BufferSync(int flags)
* BgBufferSync -- Write out some dirty buffers in the pool.
*
* This is called periodically by the background writer process.
- *
+ *
+ * If `reset` = true, the function discards any saved information and starts
+ * anew.
+ *
* Returns true if it's appropriate for the bgwriter process to go into
* low-power hibernation mode. (This happens if the strategy clock sweep
* has been "lapped" and no buffer allocations have occurred recently,
@@ -3187,7 +3190,7 @@ BufferSync(int flags)
* bgwriter_lru_maxpages to 0.)
*/
bool
-BgBufferSync(WritebackContext *wb_context)
+BgBufferSync(WritebackContext *wb_context, bool reset)
{
/* info obtained from freelist.c */
int strategy_buf_id;
@@ -3230,6 +3233,19 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ if (reset)
+ {
+ saved_info_valid = false;
+
+ /*
+ * Return from here, if we don't have a valid WritebackContext. Next time
+ * this function will be executed with a valid WritebackContext, it will
+ * start over again.
+ */
+ if (!wb_context)
+ return false;
+ }
+
/*
* Find out where the freelist clock sweep currently is, and how many
* buffer allocations have happened since our last call.
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 45a6e768332..4bb000f36e1 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -528,6 +528,38 @@ StrategyInitialize(bool init)
Assert(!init);
}
+/*
+ * StrategyReInitialize -- re-initialize the buffer cache replacement
+ * strategy.
+ *
+ * To be called when resizing buffer manager and only from the coordinator.
+ */
+void
+StrategyReInitialize(void)
+{
+ bool found;
+
+ /*
+ * Get or create the shared strategy control block. This is mostly not
+ * required since we are not moving the starting pointer anyway.
+ */
+ StrategyControl = (BufferStrategyControl *)
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, STRATEGY_SHMEM_SEGMENT);
+
+ SpinLockInit(&StrategyControl->buffer_strategy_lock);
+
+ /* Initialize the clock sweep pointer */
+ pg_atomic_init_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /* Clear statistics */
+ StrategyControl->completePasses = 0;
+ pg_atomic_init_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* No pending notification */
+ StrategyControl->bgwprocno = -1;
+}
/* ----------------------------------------------------------------
* Backend-private buffer ring management
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 416c405fe4e..77bcaf1b43c 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -437,6 +437,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyReInitialize(void);
extern bool have_free_buffer(void);
/* buf_table.c */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index fff80214822..02db81fa416 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -288,7 +288,7 @@ extern bool ConditionalLockBufferForCleanup(Buffer buffer);
extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
-extern bool BgBufferSync(struct WritebackContext *wb_context);
+extern bool BgBufferSync(struct WritebackContext *wb_context, bool reset);
extern void LimitAdditionalPins(uint32 *additional_pins);
extern void LimitAdditionalLocalPins(uint32 *additional_pins);
--
2.34.1
[text/x-patch] 0007-Fix-compilation-failures-in-previous-patche-20250228.patch (2.4K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/9-0007-Fix-compilation-failures-in-previous-patche-20250228.patch)
download | inline diff:
From b1f510627937398ceb472f09b9b28f368a1201ea Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 25 Feb 2025 17:31:05 +0530
Subject: [PATCH 07/11] Fix compilation failures in previous patches
---
src/backend/port/sysv_shmem.c | 22 +++++++++++++++++-----
1 file changed, 17 insertions(+), 5 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 8864866f26c..992ed849dc0 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -825,16 +825,20 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* Specify the segment file size using allocsize, which contains
* potentially modified size.
*/
- ftruncate(mapping->segment_fd, allocsize);
+ if (ftruncate(mapping->segment_fd, allocsize) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not set the size of file backing shared memory %p: %m",
+ mapping->shmem)));
- ptr = mmap(probe - offset, allocsize, PROT_READ | PROT_WRITE,
+ ptr = mmap((void *)((char *) probe - offset), allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS | MAP_FIXED_NOREPLACE, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
{
DebugMappings();
elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m",
- MappingName(mapping->shmem_segment), allocsize, probe - offset);
+ MappingName(mapping->shmem_segment), allocsize, ((void *) ((char *)probe - offset)));
}
}
@@ -848,7 +852,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
allocsize = mapping->shmem_size;
/* Specify the segment file size using allocsize. */
- ftruncate(mapping->segment_fd, allocsize);
+ if (ftruncate(mapping->segment_fd, allocsize) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not set the size of file backing shared memory %p: %m",
+ mapping->shmem)));
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, mapping->segment_fd, 0);
@@ -946,7 +954,11 @@ AnonymousShmemResize(void)
continue;
/* Resize the backing anon file. */
- ftruncate(m->segment_fd, new_size);
+ if (ftruncate(m->segment_fd, new_size) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not resize file backing shared memory %p: %m",
+ m->shmem)));
/*
* Fail hard if faced any issues. In theory we could try to handle this
--
2.34.1
[text/x-patch] 0009-WIP-Support-shrinking-shared-buffers-20250228.patch (8.1K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/10-0009-WIP-Support-shrinking-shared-buffers-20250228.patch)
download | inline diff:
From 17a5e2416006e885d0d0a7bada02e56c6bb40486 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 27 Feb 2025 17:39:45 +0530
Subject: [PATCH 09/11] WIP: Support shrinking shared buffers
When shrinking the shared buffers pool, each buffer in the area being shrunk
needs to be flushed if it's dirty so as not to loose the changes to that buffer
after shrinking. Also, each such buffer needs to be removed from the buffer
mapping table so that backends do not access it after shrinking. This needs to
be done before we remap the shared memory segments related to buffer pools. If
a buffer being evicted is pinned, we raise a FATAL error. TODO: Ideally we
should be just rolling back the buffer pool resizing operation and try it again.
But we need infrastructure to do so.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 79 ++++++++++++++++---
src/backend/storage/buffer/buf_init.c | 5 +-
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/pg_shmem.h | 3 +-
4 files changed, 70 insertions(+), 18 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 2b144d45cf0..f084a0747ff 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -933,14 +933,6 @@ AnonymousShmemResize(void)
*/
pending_pm_shmem_resize = false;
- /*
- * XXX: Currently only increasing of shared_buffers is supported. For
- * decreasing something similar has to be done, but buffer blocks with
- * data have to be drained first.
- */
- if(NBuffersOld > NBuffers)
- return false;
-
for(int i = 0; i < next_free_segment; i++)
{
/* Note that CalculateShmemSize indirectly depends on NBuffers */
@@ -998,8 +990,6 @@ AnonymousShmemResize(void)
* reinitialize the new portion of buffer pool. Every other
* process will wait on the shared barrier for that to finish,
* since it's a part of the SHMEM_RESIZE_DONE phase.
- *
- * XXX: This is the right place for buffer eviction as well.
*/
ResizeBufferPool(NBuffersOld, true);
@@ -1022,6 +1012,52 @@ AnonymousShmemResize(void)
return true;
}
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+static bool
+EvictExtraBuffers()
+{
+ bool result = true;
+
+ /*
+ * If the buffer being evicated is locked, this function will need to wait.
+ * This function should not be called from a Postmaster since it can not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * Let only one backend perform eviction. We could split the work across all
+ * the backends but that doesn't seem necessary. The first backend to acquire sets its own PID as the evictor PID so that other backends do not perform eviction. Any backend which can not take this lock already knows that some backend is evicting the buffers without looking at evictor_pid. All the backends which do not perform eviction still wait for this phase to finish and thus release lock before the next phase begins. Thus the same LWLock can be used to select a leader for each phase. TODO: This comment would better be placed at a place common to all phases.
+ */
+ if (LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ if (ShmemCtrl->evictor_pid == 0)
+ {
+ ShmemCtrl->evictor_pid = MyProcPid;
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer b = NBuffers + 1; b <= NBuffersOld; b++)
+ {
+ if (!EvictUnpinnedBuffer(b))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", b);
+ result = false;
+ }
+ }
+ }
+ LWLockRelease(ShmemResizeLock);
+ }
+
+ return result;
+}
+
/*
* We are asked to resize shared memory. Do the resize and make sure to wait on
* the provided barrier until all simultaneously participating backends finish
@@ -1065,15 +1101,31 @@ ProcessBarrierShmemResize(Barrier *barrier)
/* First phase means the resize has begun, SHMEM_RESIZE_START */
BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
+ /*
+ * Evict extra buffers when shrinking shared buffers. We need to do this
+ * while the memory for extra buffers is still mapped i.e. before remapping
+ * the shared memory segments to a smaller memory area.
+ */
+ if (NBuffersOld > NBuffers)
+ {
+ /*
+ * TODO: If the buffer eviction fails for any reason, we should gracefully rollback the shared buffer resizing and try again. But the infrastructure to do so is not available right now. Hence just raise a FATAL so that the system restarts.
+ */
+ if (!EvictExtraBuffers())
+ elog(FATAL, "buffer eviction failed");
+
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_EVICT);
+ }
+
/* XXX: Split mremap and buffer reinitialization into two barrier phases */
AnonymousShmemResize();
/* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
- /* Allow the last backend to reset the barrier */
+ /* Allow the last backend to reset the control area. */
if (BarrierArriveAndDetach(barrier))
- ResetShmemBarrier();
+ ResetShmemCtrl();
return true;
}
@@ -1518,7 +1570,8 @@ WaitOnShmemBarrier(int phase)
}
void
-ResetShmemBarrier(void)
+ResetShmemCtrl(void)
{
BarrierInit(&ShmemCtrl->Barrier, 0);
+ ShmemCtrl->evictor_pid = 0;
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 248fbf1633b..d8139a899bb 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -119,6 +119,7 @@ BufferManagerShmemInit(void)
/* Initialize with the currently known value */
pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
BarrierInit(&ShmemCtrl->Barrier, 0);
+ ShmemCtrl->evictor_pid = 0;
}
/* Align descriptors to a cacheline boundary. */
@@ -228,10 +229,6 @@ ResizeBufferPool(int NBuffersOld, bool initNew)
int i;
elog(DEBUG1, "Resizing buffer pool from %d to %d", NBuffersOld, NBuffers);
- /* XXX: Only increasing of shared_buffers is supported in this function */
- if(NBuffersOld > NBuffers)
- return;
-
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
ShmemInitStructInSegment("Buffer Descriptors",
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 4203c987edc..a4a1e855c48 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,7 @@ REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it c
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_EVICT "Waiting for other backends to finish buffer evication phase."
SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 3f103d708a5..3793f369313 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -68,6 +68,7 @@ extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
typedef struct
{
pg_atomic_uint32 NSharedBuffers;
+ pid_t evictor_pid;
Barrier Barrier;
} ShmemControl;
@@ -131,7 +132,7 @@ bool ProcessBarrierShmemResize(Barrier *barrier);
void assign_shared_buffers(int newval, void *extra, bool *pending);
void AdjustShmemSize(void);
extern void WaitOnShmemBarrier(int phase);
-extern void ResetShmemBarrier(void);
+extern void ResetShmemCtrl(void);
/*
* To be able to dynamically resize largest parts of the data stored in shared
--
2.34.1
[text/x-patch] 0008-Add-TODOs-and-questions-about-previous-comm-20250228.patch (11.4K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/11-0008-Add-TODOs-and-questions-about-previous-comm-20250228.patch)
download | inline diff:
From 677a089ac8a8e0f9e700d60bbe59fc91f5b5276b Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 6 Jan 2025 14:40:51 +0530
Subject: [PATCH 08/11] Add TODOs and questions about previous commits
The commit just marks the places which need more work or whether the
code raises some questions. This is not an exhaustive list of TODOs.
More TODOs may come up as I work with these patches further.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 8 +++++-
src/backend/storage/buffer/buf_init.c | 12 +++++++++
src/backend/storage/buffer/bufmgr.c | 5 ++++
src/backend/storage/buffer/freelist.c | 3 +++
src/backend/storage/ipc/ipc.c | 4 +++
src/backend/storage/ipc/ipci.c | 10 +++++++
src/backend/storage/ipc/shmem.c | 26 ++++++++++++++++---
src/backend/storage/lmgr/lwlock.c | 5 ++++
src/backend/tcop/postgres.c | 6 +++++
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/buf_internals.h | 5 ++++
src/include/storage/pg_shmem.h | 4 +++
12 files changed, 85 insertions(+), 4 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 992ed849dc0..2b144d45cf0 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1011,7 +1011,13 @@ AnonymousShmemResize(void)
LWLockRelease(ShmemResizeLock);
}
- }
+
+ /*
+ * TODO: Shouldn't we call ResizeBufferPool() here as well? Or those
+ * backend who can not lock the LWLock conditionally won't resize the
+ * buffers.
+ */
+ }
return true;
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index b7de0ab6b0d..248fbf1633b 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -211,6 +211,12 @@ BufferManagerShmemInit(void)
* initNew flag indicates that the caller wants new buffers to be initialized.
* No locks are taking in this function, it is the caller responsibility to
* make sure only one backend can work with new buffers.
+ *
+ * TODO: Avoid code duplication with BufferManagerShmemInit() and also assess
+ * which functionality in the latter is required in this function.
+ * similar to BufferManagerShmemInit, but applied only to the buffers in the
+ * range between NBuffersOld and NBuffers.
+ *
*/
void
ResizeBufferPool(int NBuffersOld, bool initNew)
@@ -293,6 +299,12 @@ ResizeBufferPool(int NBuffersOld, bool initNew)
}
/* Correct last entry of linked list */
+ /*
+ * TODO: I think this needs to be done only when expanding the buffers.
+ *
+ * TODO: We should also fix the freelist to not point to a shrunk
+ * buffer and to append the new buffers to the existing free list.
+ */
GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 80b0d0c5ded..b6bec73e6b7 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2974,6 +2974,11 @@ BufferSync(int flags)
UnlockBufHdr(bufHdr, buf_state);
/* Check for barrier events in case NBuffers is large. */
+ /*
+ * TODO: If we allow buffer resizing while this loop is being executed,
+ * the loop will need to consider the new value of NBuffers. Hence avoid
+ * resizing if this loop is being executed.
+ */
if (ProcSignalBarrierPending)
ProcessProcSignalBarrier();
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 4919a92f2be..45a6e768332 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -484,6 +484,9 @@ StrategyInitialize(bool init)
* a new entry before deleting the old. In principle this could be
* happening in each partition concurrently, so we could need as many as
* NBuffers + NUM_BUFFER_PARTITIONS entries.
+ *
+ * TODO: If we are resizing, we need to preserve the earlier entries, don't
+ * we?
*/
InitBufTable(NBuffers + NUM_BUFFER_PARTITIONS);
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 9d526eb43fd..38bd6f41130 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -70,6 +70,10 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
+/*
+ * TODO: Why do we need to increase this by 20? I didn't notice any new calls to
+ * on_shmem_exit or on_proc_exit or before_shmem_exit.
+ */
#define MAX_ON_EXITS 40
struct ONEXIT
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index a2c635f288e..029ad28fe0b 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -204,6 +204,12 @@ AttachSharedMemoryStructs(void)
/*
* CreateSharedMemoryAndSemaphores
* Creates and initializes shared memory and semaphores.
+ *
+ * TODO: IMO this function should be rewritten to calculate the size of each
+ * shared memory slot or mapping. Instead of passing slot number to
+ * CalculateShmemSize, we should instead let each shared memory module use their
+ * own slot number and update the required sizes in the corresponding mapping.
+ * Then allocate shared memory in each of the mappings.
*/
void
CreateSharedMemoryAndSemaphores(void)
@@ -225,6 +231,10 @@ CreateSharedMemoryAndSemaphores(void)
* Create the shmem segment.
*
* XXX: Do multiple shims are needed, one per segment?
+ *
+ * TODO: while each slot will return a different shim, only the last one
+ * is passed to dsm_postmaster_startup(). Is that right? Shouldn't we
+ * pass all of them or none.
*/
seghdr = PGSharedMemoryCreate(size, &shim);
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 226b38ba979..b569d125507 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -86,6 +86,11 @@ ShmemSegment Segments[ANON_MAPPINGS];
* Primary index hashtable for shmem, for simplicity we use a single for all
* shared memory segments. There can be performance consequences of that, and
* an alternative option would be to have one index per shared memory segments.
+ *
+ * TODO: shouldn't this be part of the ShmemSegment structure? Some shared
+ * memory segments that hold only one structure do not need their pointers to be
+ * stored in the shared hash table, instead they could be part of the segments
+ * itself.
*/
static HTAB *ShmemIndex = NULL;
@@ -493,9 +498,18 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. Verify the structure's size:
- * - If it's the same, we've found the expected structure.
- * - If it's different, we're resizing the expected structure.
+ * already. Verify the structure's size: - If it's the same, we've found
+ * the expected structure. - If it's different, we're resizing the
+ * expected structure.
+ *
+ * TODO: This works because every structure that needs to be resized
+ * resides in a shmem slot by itself. But it won't work if a slot
+ * contains more structures, that need to be resized, placed in adjacent
+ * memory. Also we are not updating the Shmem stats like freeoffset. I
+ * think we will keep all resizable structures in a slot for themselves,
+ * and not have a hash table in such slots since resizing the hash table
+ * itself might cause memory to be allocated next to the resizable
+ * structure making it difficult to resize it.
*/
if (result->size != size)
result->size = size;
@@ -587,6 +601,12 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
hash_seq_init(&hstat, ShmemIndex);
+ /*
+ * TODO: For the sake of completeness we should rotate through all the slots
+ * (after saving slotwise ShmemIndex, if any). Do we want to also output
+ * shmem slot name, but that would expose the slotified structure of shared
+ * memory.
+ */
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
/* XXX: take all shared memory segments into account. */
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 40aa4014b5f..8b6fe9b24f7 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -608,6 +608,11 @@ LWLockNewTrancheId(void)
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
/* We use the ShmemLock spinlock to protect LWLockCounter */
+ /*
+ * TODO: We have retained ShmemLock global variable, should we use it here
+ * instead of main segment lock? We will need spinlock init on the global
+ * one if yes.
+ */
SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 04cdd0d24d8..e88448f2846 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4671,6 +4671,12 @@ PostgresMain(const char *dbname, const char *username)
/*
* (6) check for any other interesting events that happened while we
* slept.
+ * TODO: When a backend is waiting for a command, it won't reload
+ * configuration and hence wouldn't notice change in shared_buffers. The
+ * change is only noticed after the command is received and the control
+ * comes here. We may need to improve this in case we want to resize
+ * shared buffers or perform of part of that operation in assign_hook
+ * implementation (e.g. AnonymousShmemResize()).
*/
if (ConfigReloadPending)
{
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 947f13cb1fa..4203c987edc 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -348,6 +348,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+# TODO, not used anywhere, do we need it?
ShmemResize "Waiting to resize shared memory."
#
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 4595f5a9676..416c405fe4e 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -22,6 +22,11 @@
#include "storage/condition_variable.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
+/*
+ * TODO: this header files doesn't use anything in pg_shmem.h but the files which
+ * include this file may. We should include pg_shmem.h in those files rather than
+ * here.
+ */
#include "storage/pg_shmem.h"
#include "storage/smgr.h"
#include "storage/spin.h"
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index b597df0d3a3..3f103d708a5 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -43,6 +43,10 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+/*
+ * TODO: should we define it in shmem.c where the previous global variables were
+ * declared? Do we need this structure outside shmem.c?
+ */
typedef struct ShmemSegment
{
PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
--
2.34.1
[application/x-shellscript] pgbench-concurrent-resize-buffers.sh (2.1K, ../../CAExHW5utfpJ+WTipMLCPYTixn-34HbNCxn-_SvcyQd-XkafU5g@mail.gmail.com/12-pgbench-concurrent-resize-buffers.sh)
download
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-02-28 12:01 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-03-20 08:55 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2 siblings, 2 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-02-28 12:01 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Tue, Feb 25, 2025 at 3:22 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Fri, Oct 18, 2024 at 09:21:19PM GMT, Dmitry Dolgov wrote:
> > TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
> > changing shared memory mapping layout. Any feedback is appreciated.
>
> Hi,
>
> Here is a new version of the patch, which contains a proposal about how to
> coordinate shared memory resizing between backends. The rest is more or less
> the same, a feedback about coordination is appreciated. It's a lot to read, but
> the main difference is about:
Thanks Dmitry for the summary.
>
> 1. Allowing to decouple a GUC value change from actually applying it, sort of a
> "pending" change. The idea is to let a custom logic be triggered on an assign
> hook, and then take responsibility for what happens later and how it's going to
> be applied. This allows to use regular GUC infrastructure in cases where value
> change requires some complicated processing. I was trying to make the change
> not so invasive, plus it's missing GUC reporting yet.
>
> 2. Shared memory resizing patch became more complicated thanks to some
> coordination between backends. The current implementation was chosen from few
> more or less equal alternatives, which are evolving along following lines:
>
> * There should be one "coordinator" process overseeing the change. Having
> postmaster to fulfill this role like in this patch seems like a natural idea,
> but it poses certain challenges since it doesn't have locking infrastructure.
> Another option would be to elect a single backend to be a coordinator, which
> will handle the postmaster as a special case. If there will ever be a
> "coordinator" worker in Postgres, that would be useful here.
>
> * The coordinator uses EmitProcSignalBarrier to reach out to all other backends
> and trigger the resize process. Backends join a Barrier to synchronize and wait
> untill everyone is finished.
>
> * There is some resizing state stored in shared memory, which is there to
> handle backends that were for some reason late or didn't receive the signal.
> What to store there is open for discussion.
>
> * Since we want to make sure all processes share the same understanding of what
> NBuffers value is, any failure is mostly a hard stop, since to rollback the
> change coordination is needed as well and sounds a bit too complicated for now.
>
I think we should add a way to monitor the progress of resizing; at
least whether resizing is complete and whether the new GUC value is in
effect.
> We've tested this change manually for now, although it might be useful to try
> out injection points. The testing strategy, which has caught plenty of bugs,
> was simply to run pgbench workload against a running instance and change
> shared_buffers on the fly. Some more subtle cases were verified by manually
> injecting delays to trigger expected scenarios.
I have shared a script with my changes but it's far from being full
testing. We will need to use injection points to test specific
scenarios.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-03-20 08:55 ` Ni Ku <jakkuniku@gmail.com>
2025-03-20 10:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Ni Ku @ 2025-03-20 08:55 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Dmitry / Ashutosh,
Thanks for the patch set. I've been doing some testing with it and in
particular want to see if this solution would work with hugepage bufferpool.
I ran some simple tests (outside of PG) on linux kernel v6.1, which has
this commit that added some hugepage support to mremap (
https://patchwork.kernel.org/project/linux-mm/patch/20211013195825.3058275-1-almasrymina@google.com/
).
From reading the kernel code and testing, for a hugepage-backed mapping it
seems mremap supports only shrinking but not growing. Further, for
shrinking, what I observed is that after mremap is called the hugepage
memory
is not released back to the OS, rather it's released when the fd is closed
(or when the memory is unmapped for a mapping created with MAP_ANONYMOUS).
I'm not sure if this behavior is expected, but being able to release memory
back to the OS immediately after mremap would be important for use cases
such as supporting "serverless" PG instances on the cloud.
I'm no expert in the linux kernel so I could be missing something. It'd be
great if you or somebody can comment on these observations and whether this
mremap-based solution would work with hugepage bufferpool.
I also attached the test program in case someone can spot I did something
wrong.
Regards,
Jack Ng
On Tue, Mar 18, 2025 at 11:02 AM Ashutosh Bapat <
ashutosh.bapat.oss@gmail.com> wrote:
> On Tue, Feb 25, 2025 at 3:22 PM Dmitry Dolgov <9erthalion6@gmail.com>
> wrote:
> >
> > > On Fri, Oct 18, 2024 at 09:21:19PM GMT, Dmitry Dolgov wrote:
> > > TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
> > > changing shared memory mapping layout. Any feedback is appreciated.
> >
> > Hi,
> >
> > Here is a new version of the patch, which contains a proposal about how
> to
> > coordinate shared memory resizing between backends. The rest is more or
> less
> > the same, a feedback about coordination is appreciated. It's a lot to
> read, but
> > the main difference is about:
>
> Thanks Dmitry for the summary.
>
> >
> > 1. Allowing to decouple a GUC value change from actually applying it,
> sort of a
> > "pending" change. The idea is to let a custom logic be triggered on an
> assign
> > hook, and then take responsibility for what happens later and how it's
> going to
> > be applied. This allows to use regular GUC infrastructure in cases where
> value
> > change requires some complicated processing. I was trying to make the
> change
> > not so invasive, plus it's missing GUC reporting yet.
> >
> > 2. Shared memory resizing patch became more complicated thanks to some
> > coordination between backends. The current implementation was chosen
> from few
> > more or less equal alternatives, which are evolving along following
> lines:
> >
> > * There should be one "coordinator" process overseeing the change. Having
> > postmaster to fulfill this role like in this patch seems like a natural
> idea,
> > but it poses certain challenges since it doesn't have locking
> infrastructure.
> > Another option would be to elect a single backend to be a coordinator,
> which
> > will handle the postmaster as a special case. If there will ever be a
> > "coordinator" worker in Postgres, that would be useful here.
> >
> > * The coordinator uses EmitProcSignalBarrier to reach out to all other
> backends
> > and trigger the resize process. Backends join a Barrier to synchronize
> and wait
> > untill everyone is finished.
> >
> > * There is some resizing state stored in shared memory, which is there to
> > handle backends that were for some reason late or didn't receive the
> signal.
> > What to store there is open for discussion.
> >
> > * Since we want to make sure all processes share the same understanding
> of what
> > NBuffers value is, any failure is mostly a hard stop, since to rollback
> the
> > change coordination is needed as well and sounds a bit too complicated
> for now.
> >
>
> I think we should add a way to monitor the progress of resizing; at
> least whether resizing is complete and whether the new GUC value is in
> effect.
>
> > We've tested this change manually for now, although it might be useful
> to try
> > out injection points. The testing strategy, which has caught plenty of
> bugs,
> > was simply to run pgbench workload against a running instance and change
> > shared_buffers on the fly. Some more subtle cases were verified by
> manually
> > injecting delays to trigger expected scenarios.
>
> I have shared a script with my changes but it's far from being full
> testing. We will need to use injection points to test specific
> scenarios.
>
> --
> Best Wishes,
> Ashutosh Bapat
>
>
>
>
>
Attachments:
[application/octet-stream] hugepage_remap.c (1.1K, ../../CAPuPUJz4EB5NU1ah3NH9HjBaq-dCfJMAgmou7YY5sofQ-xBbQQ@mail.gmail.com/3-hugepage_remap.c)
download
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-03-20 08:55 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
@ 2025-03-20 10:21 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-03-21 08:48 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-03-20 10:21 UTC (permalink / raw)
To: Ni Ku <jakkuniku@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Thu, Mar 20, 2025 at 04:55:47PM GMT, Ni Ku wrote:
>
> I ran some simple tests (outside of PG) on linux kernel v6.1, which has
> this commit that added some hugepage support to mremap (
> https://patchwork.kernel.org/project/linux-mm/patch/20211013195825.3058275-1-almasrymina@google.com/
> ).
>
> From reading the kernel code and testing, for a hugepage-backed mapping it
> seems mremap supports only shrinking but not growing. Further, for
> shrinking, what I observed is that after mremap is called the hugepage
> memory
> is not released back to the OS, rather it's released when the fd is closed
> (or when the memory is unmapped for a mapping created with MAP_ANONYMOUS).
> I'm not sure if this behavior is expected, but being able to release memory
> back to the OS immediately after mremap would be important for use cases
> such as supporting "serverless" PG instances on the cloud.
>
> I'm no expert in the linux kernel so I could be missing something. It'd be
> great if you or somebody can comment on these observations and whether this
> mremap-based solution would work with hugepage bufferpool.
Hm, I think you're right. I didn't realize there is such limitation, but
just verified on the latest kernel build and hit the same condition on
increasing hugetlb mapping you've mentioned above. That's annoying of
course, but I've got another approach I was originally experimenting
with -- instead of mremap do munmap and mmap with the new size and rely
on the anonymous fd to keep the memory content in between. I'm currently
reworking mmap'ing part of the patch, let me check if this new approach
is something we could universally rely on.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-03-20 08:55 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-03-20 10:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-03-21 08:48 ` Ni Ku <jakkuniku@gmail.com>
2025-03-21 09:31 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ni Ku @ 2025-03-21 08:48 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Thanks for your insights and confirmation, Dmitry.
Right, I think the anonymous fd approach would work to keep the memory
contents intact in between munmap and mmap with the new size, so bufferpool
expansion would work.
But it seems shrinking would still be problematic, since that approach
requires the anonymous fd to remain open (for memory content protection),
and so munmap would not release the memory back to the OS right away (gets
released when the fd is closed). From testing this is true for hugepage
memory at least.
Is there a way around this? Or maybe I misunderstood what you have in mind
;)
Regards,
Jack Ng
On Thu, Mar 20, 2025 at 6:21 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > On Thu, Mar 20, 2025 at 04:55:47PM GMT, Ni Ku wrote:
> >
> > I ran some simple tests (outside of PG) on linux kernel v6.1, which has
> > this commit that added some hugepage support to mremap (
> >
> https://patchwork.kernel.org/project/linux-mm/patch/20211013195825.3058275-1-almasrymina@google.com/
> > ).
> >
> > From reading the kernel code and testing, for a hugepage-backed mapping
> it
> > seems mremap supports only shrinking but not growing. Further, for
> > shrinking, what I observed is that after mremap is called the hugepage
> > memory
> > is not released back to the OS, rather it's released when the fd is
> closed
> > (or when the memory is unmapped for a mapping created with
> MAP_ANONYMOUS).
> > I'm not sure if this behavior is expected, but being able to release
> memory
> > back to the OS immediately after mremap would be important for use cases
> > such as supporting "serverless" PG instances on the cloud.
> >
> > I'm no expert in the linux kernel so I could be missing something. It'd
> be
> > great if you or somebody can comment on these observations and whether
> this
> > mremap-based solution would work with hugepage bufferpool.
>
> Hm, I think you're right. I didn't realize there is such limitation, but
> just verified on the latest kernel build and hit the same condition on
> increasing hugetlb mapping you've mentioned above. That's annoying of
> course, but I've got another approach I was originally experimenting
> with -- instead of mremap do munmap and mmap with the new size and rely
> on the anonymous fd to keep the memory content in between. I'm currently
> reworking mmap'ing part of the patch, let me check if this new approach
> is something we could universally rely on.
>
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-03-20 08:55 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-03-20 10:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-03-21 08:48 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
@ 2025-03-21 09:31 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-03-21 10:30 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-03-21 09:31 UTC (permalink / raw)
To: Ni Ku <jakkuniku@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Fri, Mar 21, 2025 at 04:48:30PM GMT, Ni Ku wrote:
> Thanks for your insights and confirmation, Dmitry.
> Right, I think the anonymous fd approach would work to keep the memory
> contents intact in between munmap and mmap with the new size, so bufferpool
> expansion would work.
> But it seems shrinking would still be problematic, since that approach
> requires the anonymous fd to remain open (for memory content protection),
> and so munmap would not release the memory back to the OS right away (gets
> released when the fd is closed). From testing this is true for hugepage
> memory at least.
> Is there a way around this? Or maybe I misunderstood what you have in mind
> ;)
The anonymous file will be truncated to it's new shrinked size before
mapping it second time (I think this part is missing in your test
example), to my understanding after a quick look at do_vmi_align_munmap,
this should be enough to make the memory reclaimable.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-03-20 08:55 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-03-20 10:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-03-21 08:48 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-03-21 09:31 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-03-21 10:30 ` Ni Ku <jakkuniku@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Ni Ku @ 2025-03-21 10:30 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
You're right Dmitry, truncating the anonymous file before mapping it again
does the trick! I see 'HugePages_Free' increases to the expected size right
after the ftruncate call for shrinking.
This alternative approach looks very promising. Thanks.
Regards,
Jack Ng
On Fri, Mar 21, 2025 at 5:31 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > On Fri, Mar 21, 2025 at 04:48:30PM GMT, Ni Ku wrote:
> > Thanks for your insights and confirmation, Dmitry.
> > Right, I think the anonymous fd approach would work to keep the memory
> > contents intact in between munmap and mmap with the new size, so
> bufferpool
> > expansion would work.
> > But it seems shrinking would still be problematic, since that approach
> > requires the anonymous fd to remain open (for memory content protection),
> > and so munmap would not release the memory back to the OS right away
> (gets
> > released when the fd is closed). From testing this is true for hugepage
> > memory at least.
> > Is there a way around this? Or maybe I misunderstood what you have in
> mind
> > ;)
>
> The anonymous file will be truncated to it's new shrinked size before
> mapping it second time (I think this part is missing in your test
> example), to my understanding after a quick look at do_vmi_align_munmap,
> this should be enough to make the memory reclaimable.
>
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-04-07 06:20 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-04-07 06:20 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Fri, Feb 28, 2025 at 5:31 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> I think we should add a way to monitor the progress of resizing; at
> least whether resizing is complete and whether the new GUC value is in
> effect.
>
I further tested this approach by tracing the barrier synchronization
using the attached patch with adds a bunch of elogs().
I ran pgbench load and simultaneously
executed following commands on a psql connection
#alter system set shared_buffers to '200MB';
ALTER SYSTEM
#select pg_reload_conf();
pg_reload_conf
----------------
t
(1 row)
#show shared_buffers;
shared_buffers
----------------
200MB
(1 row)
#select count(*) from pg_stat_activity;
count
-------
6
(1 row)
#select pg_backend_pid(); - the backend where all these commands were executed
pg_backend_pid
----------------
878405
(1 row)
I see the following in the postgresql error logs.
2025-03-12 11:04:53.812 IST [878167] LOG: received SIGHUP, reloading
configuration files
2025-03-12 11:04:53.813 IST [878405] LOG: Handle a barrier for shmem
resizing from 16384 to -1, 0
2025-03-12 11:04:53.813 IST [878341] LOG: Handle a barrier for shmem
resizing from 16384 to -1, 0
2025-03-12 11:04:53.813 IST [878341] LOG: Handle a barrier for shmem
resizing from 16384 to -1, 0
2025-03-12 11:04:53.813 IST [878341] LOG: Handle a barrier for shmem
resizing from 16384 to -1, 0
2025-03-12 11:04:53.813 IST [878341] LOG: Handle a barrier for shmem
resizing from 16384 to -1, 0
-- not all backends have reloaded configuration.
2025-03-12 11:04:53.813 IST [878173] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.813 IST [878173] LOG: attached when barrier was at phase 0
2025-03-12 11:04:53.813 IST [878173] LOG: reached barrier phase 1
2025-03-12 11:04:53.813 IST [878171] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.813 IST [878172] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.813 IST [878171] LOG: attached when barrier was at phase 1
2025-03-12 11:04:53.813 IST [878172] LOG: attached when barrier was at phase 1
2025-03-12 11:04:53.813 IST [878340] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.813 IST [878340] STATEMENT: UPDATE
pgbench_branches SET bbalance = bbalance + 1367 WHERE bid = 8;
2025-03-12 11:04:53.813 IST [878340] LOG: attached when barrier was at phase 1
2025-03-12 11:04:53.813 IST [878340] STATEMENT: UPDATE
pgbench_branches SET bbalance = bbalance + 1367 WHERE bid = 8;
2025-03-12 11:04:53.813 IST [878338] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.813 IST [878338] STATEMENT: UPDATE
pgbench_accounts SET abalance = abalance + -209 WHERE aid = 453662;
2025-03-12 11:04:53.813 IST [878339] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.813 IST [878339] STATEMENT: UPDATE
pgbench_accounts SET abalance = abalance + -3449 WHERE aid = 159726;
2025-03-12 11:04:53.813 IST [878338] LOG: attached when barrier was at phase 1
2025-03-12 11:04:53.813 IST [878338] STATEMENT: UPDATE
pgbench_accounts SET abalance = abalance + -209 WHERE aid = 453662;
2025-03-12 11:04:53.813 IST [878339] LOG: attached when barrier was at phase 1
2025-03-12 11:04:53.813 IST [878339] STATEMENT: UPDATE
pgbench_accounts SET abalance = abalance + -3449 WHERE aid = 159726;
2025-03-12 11:04:53.813 IST [878341] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.813 IST [878341] STATEMENT: BEGIN;
2025-03-12 11:04:53.814 IST [878341] LOG: attached when barrier was at phase 1
2025-03-12 11:04:53.814 IST [878341] STATEMENT: BEGIN;
2025-03-12 11:04:53.814 IST [878337] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.814 IST [878337] STATEMENT: UPDATE pgbench_tellers
SET tbalance = tbalance + -1996 WHERE tid = 392;
2025-03-12 11:04:53.814 IST [878337] LOG: attached when barrier was at phase 1
2025-03-12 11:04:53.814 IST [878337] STATEMENT: UPDATE pgbench_tellers
SET tbalance = tbalance + -1996 WHERE tid = 392;
2025-03-12 11:04:53.814 IST [878168] LOG: Handle a barrier for shmem
resizing from 16384 to -1, 0
2025-03-12 11:04:53.814 IST [878172] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878171] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878340] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878340] STATEMENT: UPDATE
pgbench_branches SET bbalance = bbalance + 1367 WHERE bid = 8;
2025-03-12 11:04:53.814 IST [878338] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878338] STATEMENT: UPDATE
pgbench_accounts SET abalance = abalance + -209 WHERE aid = 453662;
2025-03-12 11:04:53.814 IST [878341] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878341] STATEMENT: BEGIN;
2025-03-12 11:04:53.814 IST [878337] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878337] STATEMENT: UPDATE pgbench_tellers
SET tbalance = tbalance + -1996 WHERE tid = 392;
2025-03-12 11:04:53.814 IST [878173] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878339] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878339] STATEMENT: UPDATE
pgbench_accounts SET abalance = abalance + -3449 WHERE aid = 159726;
2025-03-12 11:04:53.814 IST [878172] LOG: reached barrier phase 3
2025-03-12 11:04:53.814 IST [878340] LOG: reached barrier phase 3
2025-03-12 11:04:53.814 IST [878340] STATEMENT: UPDATE
pgbench_branches SET bbalance = bbalance + 1367 WHERE bid = 8;
2025-03-12 11:04:53.814 IST [878341] LOG: reached barrier phase 3
2025-03-12 11:04:53.814 IST [878341] STATEMENT: BEGIN;
2025-03-12 11:04:53.814 IST [878339] LOG: reached barrier phase 3
2025-03-12 11:04:53.814 IST [878339] STATEMENT: UPDATE
pgbench_accounts SET abalance = abalance + -3449 WHERE aid = 159726;
2025-03-12 11:04:53.814 IST [878171] LOG: reached barrier phase 3
2025-03-12 11:04:53.814 IST [878338] LOG: reached barrier phase 3
2025-03-12 11:04:53.814 IST [878338] STATEMENT: UPDATE
pgbench_accounts SET abalance = abalance + -209 WHERE aid = 453662;
2025-03-12 11:04:53.814 IST [878337] LOG: reached barrier phase 3
2025-03-12 11:04:53.814 IST [878337] STATEMENT: UPDATE pgbench_tellers
SET tbalance = tbalance + -1996 WHERE tid = 392;
2025-03-12 11:04:53.814 IST [878337] LOG: buffer resizing operation
finished at phase 4
2025-03-12 11:04:53.814 IST [878337] STATEMENT: UPDATE pgbench_tellers
SET tbalance = tbalance + -1996 WHERE tid = 392;
2025-03-12 11:04:53.814 IST [878168] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.814 IST [878168] LOG: attached when barrier was at phase 0
2025-03-12 11:04:53.814 IST [878168] LOG: reached barrier phase 1
2025-03-12 11:04:53.814 IST [878168] LOG: reached barrier phase 2
2025-03-12 11:04:53.814 IST [878168] LOG: buffer resizing operation
finished at phase 3
2025-03-12 11:04:53.815 IST [878169] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:53.815 IST [878169] LOG: attached when barrier was at phase 0
2025-03-12 11:04:53.815 IST [878169] LOG: reached barrier phase 1
2025-03-12 11:04:53.815 IST [878169] LOG: reached barrier phase 2
2025-03-12 11:04:53.815 IST [878169] LOG: buffer resizing operation
finished at phase 3
2025-03-12 11:04:55.965 IST [878405] LOG: Handle a barrier for shmem
resizing from 16384 to -1, 0
2025-03-12 11:04:55.965 IST [878405] LOG: Handle a barrier for shmem
resizing from 16384 to -1, 0
2025-03-12 11:04:55.965 IST [878405] LOG: Handle a barrier for shmem
resizing from 16384 to 25600, 1
2025-03-12 11:04:55.965 IST [878405] STATEMENT: show shared_buffers;
2025-03-12 11:04:55.965 IST [878405] LOG: attached when barrier was at phase 0
2025-03-12 11:04:55.965 IST [878405] STATEMENT: show shared_buffers;
2025-03-12 11:04:55.965 IST [878405] LOG: reached barrier phase 1
2025-03-12 11:04:55.965 IST [878405] STATEMENT: show shared_buffers;
2025-03-12 11:04:55.965 IST [878405] LOG: reached barrier phase 2
2025-03-12 11:04:55.965 IST [878405] STATEMENT: show shared_buffers;
2025-03-12 11:04:55.965 IST [878405] LOG: buffer resizing operation
finished at phase 3
2025-03-12 11:04:55.965 IST [878405] STATEMENT: show shared_buffers;
To tell the story in short. pid 173 (for the sake of brevity I am just
mentioning the last three digits of PID) attached to the barrier first
and immediately reached phase 1. 171, 172, 340, 338, 339, 341, 337 -
all attached barrier in phase 1. All of these backends completed the
phases in synchronous fashion. But 168, 169 and 405 were yet to attach
to the barrier since they hadn't loaded their configurations yet. Each
of these backends then finished all phases independent of others.
For your reference
#select pid, application_name, backend_type from pg_stat_activity
where pid in (878169, 878168);
pid | application_name | backend_type
--------+------------------+-------------------
878168 | | checkpointer
878169 | | background writer
(2 rows)
This is because the BarrierArriveAndWait() only waits for all the
attached backends. It doesn't wait for backends which are yet to
attach. I think what we want is *all* the backends should execute all
the phases synchronously and wait for others to finish. If we don't do
that, there's a possibility that some of them would see inconsistent
buffer states or even worse may not have necessary memory mapped and
resized - thus causing segfaults. Am I correct?
I think what needs to be done is that every backend should wait for other
backends to attach themselves to the barrier before moving to the
first phase. One way I can think of is we use two signal barriers -
one to ensure that all the backends have attached themselves and
second for the actual resizing. But then the postmaster needs to wait for
all the processes to process the first signal barrier. A postmaster can
not wait on anything. Maybe there's a way to poll, but I didn't find
it. Does that mean that we have to make some other backend a coordinator?
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-04-07 08:43 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-09 05:42 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-07 08:43 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Mon, Apr 07, 2025 at 11:50:46AM GMT, Ashutosh Bapat wrote:
> This is because the BarrierArriveAndWait() only waits for all the
> attached backends. It doesn't wait for backends which are yet to
> attach. I think what we want is *all* the backends should execute all
> the phases synchronously and wait for others to finish. If we don't do
> that, there's a possibility that some of them would see inconsistent
> buffer states or even worse may not have necessary memory mapped and
> resized - thus causing segfaults. Am I correct?
>
> I think what needs to be done is that every backend should wait for other
> backends to attach themselves to the barrier before moving to the
> first phase. One way I can think of is we use two signal barriers -
> one to ensure that all the backends have attached themselves and
> second for the actual resizing. But then the postmaster needs to wait for
> all the processes to process the first signal barrier. A postmaster can
> not wait on anything. Maybe there's a way to poll, but I didn't find
> it. Does that mean that we have to make some other backend a coordinator?
Yes, you're right, plain dynamic Barrier does not ensure all available
processes will be synchronized. I was aware about the scenario you
describe, it's mentioned in commentaries for the resize function. I was
under the impression this should be enough, but after some more thinking
I'm not so sure anymore. Let me try to structure it as a list of
possible corner cases that we need to worry about:
* New backend spawned while we're busy resizing shared memory. Those
should wait until the resizing is complete and get the new size as well.
* Old backend receives a resize message, but exits before attempting to
resize. Those should be excluded from coordination.
* A backend is blocked and not responding before or after the
ProcSignalBarrier message was sent. I'm thinking about a failure
situation, when one rogue backend is doing something without checking
for interrupts. We need to wait for those to become responsive, and
potentially abort shared memory resize after some timeout.
* Backends join the barrier in disjoint groups with some time in
between, which is longer than what it takes to resize shared memory.
That means that relying only on the shared dynamic barrier is not
enough -- it will only synchronize resize procedure withing those
groups.
Out of those I think the third poses some problems, e.g. if we shrinking
the shared memory, but one backend is accessing buffer pool without
checking for interrupts. In the v3 implementation this won't be handled
correctly, other backends will ignore such rogue process. Independently
from that we could reason about the logic much easier if it's guaranteed
that all the process to resize shared memory will wait for each other to
start simultaneously.
Looks like to achieve that we need a slightly different combination of a
global Barrier and ProcSignalBarrier mechanism. We can't use
ProcSignalBarrier as it is, because processes need to wait for each
other, and at the same time finish processing to bump the generation. We
also can't use a simple dynamic Barrier due to possibility of disjoint
groups of processes. A static Barrier is also not easier, because we
would need somehow to know exact number of processes, which might change
over time.
I think a relatively elegant solution is to extend ProcSignalBarrier
mechanism to track not only pss_barrierGeneration, as a sign that
everything was processed, but also something like
pss_barrierReceivedGeneration, indicating that the message was received
everywhere but not processed yet. That would be enough to allow
processes to wait until the resize message was received everywhere, then
use a global Barrier to wait until all processes are finished. It's
somehow similar to your proposal to use two signals, but has less
implementation overhead.
This would also allow different solutions regarding error handling. E.g.
we could do an unbounded waiting for all processes we expect to resize,
assuming that the user will be able to intervene and fix an issue if
there is any. Or we can do a timed waiting, and abort the resize after
some timeout of not all processes are ready yet. In the new v4 version
of the patch the first option is implemented.
On top of that there are following changes:
* Shared memory address space is now reserved for future usage, making
shared memory segments clash (e.g. due to memory allocation)
impossible. There is a new GUC to control how much space to reserve,
which is called max_available_memory -- on the assumption that most of
the time it would make sense to set its value to the total amount of
memory on the machine. I'm open for suggestions regarding the name.
* There is one more patch to address hugepages remap. As mentioned in
this thread above, Linux kernel has certain limitations when it comes
to mremap for segments allocated with huge pages. To work around it's
possible to replace mremap with a sequence of unmap and map again,
relying on the anon file behind the segment to keep the memory
content. I haven't found any downsides of this approach so far, but it
makes the anonymous file patch 0007 mandatory.
From 15b87a1cb89d3f31b656e27d07ead5aa935f1643 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH v4 1/8] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 ++++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 141 +++++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 13 ++-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
12 files changed, 278 insertions(+), 125 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index f7c8638aec5..b6301463ac7 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -313,7 +313,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -334,7 +334,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..56af0231d24 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 567739b5be9..5b55bec8d9d 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2fa045e6b0f..8b38e985327 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -363,7 +368,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 895a43fb39e..389abc82519 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -75,19 +75,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/*
@@ -96,9 +96,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -109,7 +117,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -118,9 +132,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -145,11 +159,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -179,6 +199,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -198,22 +224,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -231,6 +257,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -241,19 +273,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -268,7 +300,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -329,6 +367,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -345,9 +395,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -380,6 +430,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -388,7 +445,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -411,7 +468,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -428,8 +485,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -454,7 +511,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -473,14 +530,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -537,10 +593,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -552,15 +609,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 3df29658f18..8241c061507 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -618,10 +620,15 @@ LWLockNewTrancheId(void)
int *LWLockCounter;
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
- /* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ /*
+ * We use the ShmemLock spinlock to protect LWLockCounter.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
+ */
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..6ebda479ced 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index b99ebc9e86f..138078c29c5 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 904a336b851..5929f140236 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: 5e1915439085014140314979c4dd5e23bd677cac
--
2.45.1
From eae77d430e6e6cc3ec95b2cf613e4b3ae095e75e Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH v4 2/8] Address space reservation for shared memory
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is:
* To reserve some address space via mmap'ing a large chunk of memory
with PROT_NONE and MAP_NORESERVE. This way we prepare a playground for
preparing shared memory layout without risking anything interfering
with that.
* To slice the reserved space up into sections, one to use for each
shared segment.
* Allocate shared memory segments out of corresponding slices and
leaving unclaimed space in between them. This is implemented via
mmap'ing memory at a specified address from the reserved space with
MAP_FIXED.
The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
7f444196c000-7f470a800000 # reserved space
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
Things like address space randomization should not be a problem in this
context, since the randomization is applied to the mmap base, which is
one per process.
This approach also do not impact the actual memory usage as reported by
the kernel. Here is the output of /proc/$PID/status for the master
version with shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages
// mapped in mm_struct. It corresponds to the mapped reserved space
// and is the only number that grows with it.
VmPeak: 2043192 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 22908 kB
// Size of resident anonymous memory
RssAnon: 768 kB
// Size of resident file mappings
RssFile: 10364 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 11776 kB
Here is the same for the patch when reserving 20GB of space:
VmPeak: 21250648 kB
VmRSS: 22948 kB
RssAnon: 768 kB
RssFile: 10404 kB
RssShmem: 11776 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
17637376 (~16.8 MB)
To control the amount of space reserved a new GUC max_available_memory
is introduced. Ideally it should be based on the maximum available
memory, hense the name.
---
src/backend/port/sysv_shmem.c | 284 ++++++++++++++++++++++++----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 5 +-
src/backend/utils/init/globals.c | 1 +
src/backend/utils/misc/guc_tables.c | 14 ++
src/include/storage/pg_shmem.h | 4 +-
6 files changed, 271 insertions(+), 39 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 56af0231d24..a0f03ff868f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,66 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * To achieve that we first reserve some shared memory address space by
+ * mmap'ing a segment of MaxAvailableMemory size with PROT_NONE and
+ * MAP_NORESERVE (these flags allow to make sure this space will not be used by
+ * anything else, yet do not count against memory limits). Having the reserved
+ * space, we allocate out of it actual chunks of shared memory as usual,
+ * updating a pointer to the current available reserved space for the next
+ * allocation with the gap between segments in mind.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f442e010000-7f443a800000 # reserved empty space
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f444196c000-7f470a800000 # reserved empty space
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * [...]
+ *
+ * The reserved space pointer is calculated to slice up the total reserved
+ * space into fixed fractions of address space for each segment, as specified
+ * in the SHMEM_RESIZE_RATIO array.
+ */
+static double SHMEM_RESIZE_RATIO[1] = {
+ 1.0, /* MAIN_SHMEM_SLOT */
+};
+
+/*
+ * Offset from the beginning of the reserved space, which indicates currently
+ * available range. New shared memory segments have to be allocated at this
+ * offset related to the reserved space.
+ */
+static Size reserved_offset = 0;
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -626,39 +686,198 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
*
* This function will modify mapping size to the actual size of the allocation,
* if it ends up allocating a segment that is larger than requested.
+ *
+ * Note that we do not switch from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
*/
static void
-CreateAnonymousSegment(AnonymousMapping *mapping)
+CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
{
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS;
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* ReserveAnonymousMemory should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the request size to a suitable large value */
GetHugePageSize(&hugepagesize, &mmap_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
+ }
+#endif
+
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p",
+ MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+
+ /*
+ * Try to create mapping at an address out of the reserved range, which
+ * will allow to extend it later. Use reserved_offset to allocate the
+ * segment, then update currently available reserved range.
+ *
+ * If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized without
+ * a restart.
+ */
+ ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_FIXED, -1, 0);
+ mmap_errno = errno;
+
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m, "
+ "fallback to the non-resizable allocation",
+ MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS, -1, 0);
+ mmap_errno = errno;
+ }
+ else
+ {
+ Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+
+ reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ errno = mmap_errno;
+ DebugMappings();
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
+ (mmap_errno == ENOMEM) ?
+ errhint("This error usually means that PostgreSQL's request "
+ "for a shared memory segment exceeded available memory, "
+ "swap space, or huge pages. To reduce the request size "
+ "(currently %zu bytes), reduce PostgreSQL's shared "
+ "memory usage, perhaps by reducing \"shared_buffers\" or "
+ "\"max_connections\".",
+ allocsize) : 0));
+ }
+
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
+}
+
+/*
+ * ReserveAnonymousMemory
+ *
+ * Reserve shared memory address space, from which shared memory segments are
+ * going to be sliced out. The goal of this exercise is to support segments
+ * resizing, for which we need a reserved space free of potential clashes with
+ * other mmap'd areas that are not under our control. Reservation is done via
+ * mmap, and will not allocate any memory until it will be actually used, and
+ * MAP_NORESERVE allows to make it not counting againt kernel reservation
+ * limits (e.g. in cgroups or for huge pages). Do not get confused because of
+ * MAP_NORESERVE -- we need to reserve some space, but not the actual memory,
+ * and that is that this flag is about.
+ *
+ * Note, that with MAP_NORESERVE a reservation with hugetlb will succeed even
+ * if there is actually not enough huge pages. Hence this function is
+ * responsible for deciding whether to use huge pages or not. To achieve that
+ * we need to probe first and try to allocate needed memory for all segments --
+ * if this succeeds, we unmap the probe segment and use hugetlb; if it fails,
+ * we proceed with the regular memory.
+ */
+void *
+ReserveAnonymousMemory(Size reserve_size)
+{
+ Size allocsize = reserve_size;
+ void *ptr = MAP_FAILED;
+ int mmap_errno = 0;
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ *
+ * We could actually have a mix and match of segments with and without
+ * huge pages. But in that case we need to have multiple reservation
+ * spaces to use corresponding memory (hugetlb adress space reserved
+ * for hugetlb segments, regular memory for others), and it doesn't
+ * seem to worth the complexity for now.
+ */
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ int numSemas;
+ Size segment_size = CalculateShmemSize(&numSemas, segment);
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
- DebugMappings();
- elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- MappingName(mapping->shmem_segment), allocsize);
+ /* No huge pages, we will go with the regular page size */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB "
+ "failed, huge pages disabled: %m", total_size);
+ }
+ else
+ {
+ /*
+ * All fine, unmap the temporary segment and proceed with reserving
+ * using huge pages.
+ */
+ if (munmap(ptr, total_size) < 0)
+ elog(LOG, "reservice space: munmap(%p, %zu) failed: %m",
+ ptr, total_size);
+
+ /* Round up the requested size to a suitable large value. */
+ if (allocsize % hugepagesize != 0)
+ allocsize += hugepagesize - (allocsize % hugepagesize);
+
+ elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB",
+ allocsize);
+ ptr = mmap(NULL, allocsize, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | MAP_NORESERVE | mmap_flags,
+ -1, 0);
+ mmap_errno = errno;
+
+ /* This should not happen, but handle errors anyway */
+ if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
+ {
+ elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB "
+ "failed, huge pages disabled: %m", allocsize);
+ }
}
}
#endif
@@ -666,10 +885,12 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
/*
* Report whether huge pages are in use. This needs to be tracked before
* the second mmap() call if attempting to use huge pages failed
- * previously.
+ * previously. At this point ptr is either pointing to the probe segment,
+ * if we couldn't mmap it, or the reservation space.
*/
SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
@@ -677,10 +898,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
+ allocsize = reserve_size;
+
+ elog(DEBUG1, "reserving space: mmap(%zu)", allocsize);
+ ptr = mmap(NULL, allocsize, PROT_NONE,
+ MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0);
}
if (ptr == MAP_FAILED)
@@ -688,20 +910,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
errno = mmap_errno;
DebugMappings();
ereport(FATAL,
- (errmsg("segment[%s]: could not map anonymous shared memory: %m",
- MappingName(mapping->shmem_segment)),
+ (errmsg("reserving space: could not map anonymous shared "
+ "memory: %m"),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
- "for a shared memory segment exceeded available memory, "
- "swap space, or huge pages. To reduce the request size "
- "(currently %zu bytes), reduce PostgreSQL's shared "
- "memory usage, perhaps by reducing \"shared_buffers\" or "
- "\"max_connections\".",
+ "for a reserved shared memory address space exceeded "
+ "available memory, swap space, or huge pages. To "
+ "reduce the request reservation size (currently %zu "
+ "bytes), reduce PostgreSQL's \"maximum_shared_buffers\".",
allocsize) : 0));
}
- mapping->shmem = ptr;
- mapping->shmem_size = allocsize;
+ return ptr;
}
/*
@@ -740,7 +960,7 @@ AnonymousShmemDetach(int status, Datum arg)
*/
PGShmemHeader *
PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim)
+ PGShmemHeader **shim, Pointer base)
{
IpcMemoryKey NextShmemSegID;
void *memAddress;
@@ -760,14 +980,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -782,7 +994,7 @@ PGSharedMemoryCreate(Size size,
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
/* On success, mapping data will be modified. */
- CreateAnonymousSegment(mapping);
+ CreateAnonymousSegment(mapping, base);
next_free_segment++;
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..ce719f1b412 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -205,7 +205,7 @@ EnableLockPagesPrivilege(int elevel)
*/
PGShmemHeader *
PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim)
+ PGShmemHeader **shim, Pointer base)
{
void *memAddress;
PGShmemHeader *hdr;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8b38e985327..076888c0172 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -203,9 +203,12 @@ CreateSharedMemoryAndSemaphores(void)
PGShmemHeader *seghdr;
Size size;
int numSemas;
+ void *base;
Assert(!IsUnderPostmaster);
+ base = ReserveAnonymousMemory((Size) MaxAvailableMemory * BLCKSZ);
+
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
/* Compute the size of the shared-memory block */
@@ -217,7 +220,7 @@ CreateSharedMemoryAndSemaphores(void)
*
* XXX: Do multiple shims are needed, one per segment?
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(size, &shim, base);
/*
* Make sure that huge pages are never reported as "unknown" while the
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index 2152aad97d9..1d42a5856c0 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffers = 16384;
+int MaxAvailableMemory = 131072;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 4eaeca89f2c..dede37f7905 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2364,6 +2364,20 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"max_available_memory", PGC_SIGHUP, RESOURCES_MEM,
+ gettext_noop("Sets the upper limit for the shared_buffers value."),
+ gettext_noop("Shared memory could be resized at runtime, this "
+ "parameters sets the upper limit for it, beyond which "
+ "resizing would not be supported. Normally this value "
+ "would be the same as the total available memory."),
+ GUC_UNIT_BLOCKS
+ },
+ &MaxAvailableMemory,
+ 131072, 16, INT_MAX / 2,
+ NULL, NULL, NULL
+ },
+
{
{"vacuum_buffer_usage_limit", PGC_USERSET, RESOURCES_MEM,
gettext_noop("Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum."),
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 138078c29c5..4a83e255652 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -60,6 +60,7 @@ extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
+extern PGDLLIMPORT int MaxAvailableMemory;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -100,10 +101,11 @@ extern void PGSharedMemoryNoReAttach(void);
#endif
extern PGShmemHeader *PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim);
+ PGShmemHeader **shim, Pointer base);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+void *ReserveAnonymousMemory(Size reserve_size);
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
--
2.45.1
From dca1257476fc4c0718fec35b11ba0f4a4e57151b Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 15 Mar 2025 16:38:59 +0100
Subject: [PATCH v4 3/8] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 24 +++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 105 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a0f03ff868f..f46d9d5d9cd 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -147,10 +147,18 @@ static int next_free_segment = 0;
*
* The reserved space pointer is calculated to slice up the total reserved
* space into fixed fractions of address space for each segment, as specified
- * in the SHMEM_RESIZE_RATIO array.
+ * in the SHMEM_RESIZE_RATIO array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take
+ * up to 60% of the whole space when resizing, based on the fact that it most
+ * likely will be the main consumer of this memory. Those numbers are pulled
+ * out of thin air for now, makes sense to evaluate them more precise.
*/
-static double SHMEM_RESIZE_RATIO[1] = {
- 1.0, /* MAIN_SHMEM_SLOT */
+static double SHMEM_RESIZE_RATIO[6] = {
+ 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.6, /* BUFFERS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -182,6 +190,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1dc488a42..bd68b69ee98 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +77,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -156,33 +160,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..a9952b36eba 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -22,6 +22,7 @@
#include "postgres.h"
#include "storage/buf_internals.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -59,10 +60,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 336715b6c63..81543cb5ced 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -491,9 +492,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 076888c0172..9d00b80b4f8 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index f2192ceb271..1977001e533 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -308,7 +308,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 4a83e255652..c5009a1cd73 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -107,7 +107,29 @@ extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
void *ReserveAnonymousMemory(Size reserve_size);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.45.1
From 24704e57aea0ee94fbfb37ca3a4ea4fcf050a738 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:40:32 +0200
Subject: [PATCH v4 4/8] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 61 insertions(+), 40 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index ec40c0b7c42..9aa426992a2 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2321,7 +2321,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index a9f2a3a3062..e715a6f01c2 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,13 +1161,13 @@ assign_maintenance_io_concurrency(int newval, void *extra)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra)
+assign_io_max_combine_limit(int newval, void *extra, bool *pending)
{
io_max_combine_limit = newval;
io_combine_limit = Min(io_max_combine_limit, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra)
+assign_io_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit_guc = newval;
io_combine_limit = Min(io_max_combine_limit, io_combine_limit_guc);
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index e5171467de1..2a6a587ef76 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1952,7 +1952,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1985,7 +1985,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2008,7 +2008,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2031,7 +2031,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 6ae9f38f0c8..b1fba850f02 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3593,7 +3593,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 667df448732..bb681f5bc60 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f619100467d..8802ad8a3cb 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 799fa7ace68..c8300cffa8e 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,14 +81,14 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
-extern void assign_io_max_combine_limit(int newval, void *extra);
-extern void assign_io_combine_limit(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
+extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -143,13 +143,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -165,7 +165,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.45.1
From 619b10ec409185995a4a3ffd56972f1efa493c45 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 4 Apr 2025 21:46:14 +0200
Subject: [PATCH v4 5/8] Introduce pss_barrierReceivedGeneration
Currently WaitForProcSignalBarrier allows to make sure the message sent
via EmitProcSignalBarrier was processed by all ProcSignal mechanism
participants.
Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
which will be updated when a process has received the message, but not
processed it yet. This makes it possible to support a new mode of
waiting, when ProcSignal participants want to synchronize message
processing. To do that, a participant can wait via
WaitForProcSignalBarrierReceived when processing a message, effectively
making sure that all processes are going to start processing
ProcSignalBarrier simultaneously.
---
src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------
src/include/storage/procsignal.h | 1 +
2 files changed, 54 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index b7c39a4c5f0..8e313ad9bf8 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -58,7 +58,10 @@
* of it. For such use cases, we set a bit in pss_barrierCheckMask and then
* increment the current "barrier generation"; when the new barrier generation
* (or greater) appears in the pss_barrierGeneration flag of every process,
- * we know that the message has been received everywhere.
+ * we know that the message has been received and processed everywhere. In case
+ * if we only need to know only that the message was received everywhere (e.g.
+ * receiving processes need to handle the message in a coordinated fashion)
+ * use pss_barrierReceivedGeneration in the same way.
*/
typedef struct
{
@@ -70,6 +73,7 @@ typedef struct
/* Barrier-related fields (not protected by pss_mutex) */
pg_atomic_uint64 pss_barrierGeneration;
+ pg_atomic_uint64 pss_barrierReceivedGeneration;
pg_atomic_uint32 pss_barrierCheckMask;
ConditionVariable pss_barrierCV;
} ProcSignalSlot;
@@ -151,6 +155,8 @@ ProcSignalShmemInit(void)
slot->pss_cancel_key_len = 0;
MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration,
+ PG_UINT64_MAX);
pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
ConditionVariableInit(&slot->pss_barrierCV);
}
@@ -198,6 +204,8 @@ ProcSignalInit(char *cancel_key, int cancel_key_len)
barrier_generation =
pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration);
pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration,
+ barrier_generation);
if (cancel_key_len > 0)
memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len);
@@ -262,6 +270,7 @@ CleanupProcSignalState(int status, Datum arg)
* no barrier waits block on it.
*/
pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX);
SpinLockRelease(&slot->pss_mutex);
@@ -415,12 +424,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type)
return generation;
}
-/*
- * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
- * requested by a specific call to EmitProcSignalBarrier() have taken effect.
- */
-void
-WaitForProcSignalBarrier(uint64 generation)
+static void
+WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly)
{
Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration));
@@ -435,12 +440,17 @@ WaitForProcSignalBarrier(uint64 generation)
uint64 oldval;
/*
- * It's important that we check only pss_barrierGeneration here and
- * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared
- * before the barrier is actually absorbed, but pss_barrierGeneration
+ * It's important that we check only pss_barrierGeneration &
+ * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in
+ * pss_barrierCheckMask get cleared before the barrier is actually
+ * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration
* is updated only afterward.
*/
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
while (oldval < generation)
{
if (ConditionVariableTimedSleep(&slot->pss_barrierCV,
@@ -449,7 +459,11 @@ WaitForProcSignalBarrier(uint64 generation)
ereport(LOG,
(errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier",
(int) pg_atomic_read_u32(&slot->pss_pid))));
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
}
ConditionVariableCancelSleep();
}
@@ -463,12 +477,33 @@ WaitForProcSignalBarrier(uint64 generation)
* The caller is probably calling this function because it wants to read
* the shared state or perform further writes to shared state once all
* backends are known to have absorbed the barrier. However, the read of
- * pss_barrierGeneration was performed unlocked; insert a memory barrier
- * to separate it from whatever follows.
+ * pss_barrierGeneration & pss_barrierReceivedGeneration was performed
+ * unlocked; insert a memory barrier to separate it from whatever follows.
*/
pg_memory_barrier();
}
+/*
+ * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
+ * requested by a specific call to EmitProcSignalBarrier() have taken effect.
+ */
+void
+WaitForProcSignalBarrier(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, false);
+}
+
+/*
+ * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all
+ * backends have observed the message sent by a specific call to
+ * EmitProcSignalBarrier().
+ */
+void
+WaitForProcSignalBarrierReceived(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, true);
+}
+
/*
* Handle receipt of an interrupt indicating a global barrier event.
*
@@ -522,6 +557,10 @@ ProcessProcSignalBarrier(void)
if (local_gen == shared_gen)
return;
+ /* The message is observed, record that */
+ pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
+ shared_gen);
+
/*
* Get and clear the flags that are set for this backend. Note that
* pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 016dfd9b3f6..defd8b66a19 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, char *cancel_key, int cancel_key_l
extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type);
extern void WaitForProcSignalBarrier(uint64 generation);
+extern void WaitForProcSignalBarrierReceived(uint64 generation);
extern void ProcessProcSignalBarrier(void);
extern void procsignal_sigusr1_handler(SIGNAL_ARGS);
--
2.45.1
From 886a3ea87408e628bea08a9c77116343616ad032 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:47:16 +0200
Subject: [PATCH v4 6/8] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers and extend it using mremap. One elected process signals the
postmaster to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f90cde00000-7f90d4fa6000 /dev/zero (deleted)
7f90d4fa6000-7f914de00000
7f914de00000-7f915cfa8000 /dev/zero (deleted)
^ buffers mapping, ~241 MB
7f915cfa8000-7f944de00000
7f944de00000-7f94550a8000 /dev/zero (deleted)
7f94550a8000-7f94cde00000
7f94cde00000-7f94d4fe8000 /dev/zero (deleted)
7f94d4fe8000-7f954de00000
7f954de00000-7f9554ff6000 /dev/zero (deleted)
7f9554ff6000-7f958de00000
7f958de00000-7f959508a000 /dev/zero (deleted)
7f959508a000-7f95cde00000
-- 512 MB
7f90cde00000-7f90d5126000 /dev/zero (deleted)
7f90d5126000-7f914de00000
7f914de00000-7f9175128000 /dev/zero (deleted)
^ buffers mapping, ~627 MB
7f9175128000-7f944de00000
7f944de00000-7f9455528000 /dev/zero (deleted)
7f9455528000-7f94cde00000
7f94cde00000-7f94d5228000 /dev/zero (deleted)
7f94d5228000-7f954de00000
7f954de00000-7f9555266000 /dev/zero (deleted)
7f9555266000-7f958de00000
7f958de00000-7f95954aa000 /dev/zero (deleted)
7f95954aa000-7f95cde00000
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it receives something.
Note, that mremap is Linux specific, thus the implementation not very
portable.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 413 ++++++++++++++++++
src/backend/postmaster/postmaster.c | 18 +
src/backend/storage/buffer/buf_init.c | 75 ++--
src/backend/storage/ipc/ipci.c | 18 +-
src/backend/storage/ipc/procsignal.c | 46 ++
src/backend/storage/ipc/shmem.c | 23 +-
src/backend/tcop/postgres.c | 10 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/miscadmin.h | 1 +
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 26 ++
src/include/storage/pmsignal.h | 1 +
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
17 files changed, 603 insertions(+), 43 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index f46d9d5d9cd..a3437973784 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -105,6 +111,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -176,6 +189,49 @@ static Size reserved_offset = 0;
*/
static bool huge_pages_on = false;
+/*
+ * Flag telling that we have prepared the memory layout to be resizable. If
+ * false after all shared memory segments creation, it means we failed to setup
+ * needed layout and falled back to the regular non-resizable approach.
+ */
+static bool shmem_resizable = false;
+
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -769,6 +825,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
{
Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+ shmem_resizable = true;
reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
}
@@ -964,6 +1021,315 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * mremap, which will fail if is not sufficient space to expand the mapping.
+ * When finished, based on the new and old values initialize new buffer blocks
+ * if any.
+ *
+ * If reinitializing took place, as the last step this function does buffers
+ * reinitialization as well and broadcasts the new value of NSharedBuffers. All
+ * of that needs to be done only by one backend, the first one that managed to
+ * grab the ShmemResizeLock.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ void *ptr = MAP_FAILED;
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ /* Clean up some reserved space to resize into */
+ if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not unmap %zu from reserved shared memory %p: %m",
+ new_size - m->shmem_size, m->shmem)));
+
+ /* Claim the unused space */
+ elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
+ if (ptr == MAP_FAILED)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m",
+ MappingName(m->shmem_segment), m->shmem, NBuffers,
+ new_size)));
+
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * Note that it's enough when only one backend will do that,
+ * even the ShmemInitStruct part. The reason is that resized
+ * shared memory will maintain the same addresses, meaning that
+ * all the pointers are still valid, and we only need to update
+ * structures size in the ShmemIndex once -- any other backend
+ * will pick up this shared structure from the index.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ BufferManagerShmemInit(NBuffersOld);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Wait for all ProcSignal participants
+ * to join the barrier, then do the resize and wait on the barrier until all
+ * participating finish resizing as well -- otherwise we face danger of
+ * inconsistency between backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * First thing to do after attaching to the barrier is to wait for others.
+ * We can't simply use BarrierArriveAndWait, because backends might arrive
+ * here in disjoint groups, e.g. first two backends, pause, then second two
+ * backends. If the resize is quick enough that can lead to a situation
+ * when the first group is already finished before the second has appeared,
+ * and the barrier will only synchonize withing those groups.
+ */
+ if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
+ WaitForProcSignalBarrierReceived(
+ pg_atomic_read_u64(&ShmemCtrl->Generation));
+
+ /*
+ * Now start the procedure, and elect one backend to ping postmaster to do
+ * the same.
+ *
+ * XXX: If we need to be able to abort resizing, this has to be done later,
+ * after the SHMEM_RESIZE_DONE.
+ */
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ {
+ Assert(IsUnderPostmaster);
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ }
+
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ BarrierDetach(barrier);
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ elog(DEBUG1, "Set pending signal");
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Start resizing procedure, making sure all existing processes will have
+ * consistent view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * We use dynamic barrier to help dealing with backends that were spawned
+ * during the resize.
+ */
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ *
+ * Those phases are ensured by joining the shared barrier associated with
+ * the procedure. Since resizing takes time, we need to take into account
+ * that during that time:
+ *
+ * - New backends can be spawned. They will check status of the barrier
+ * early during the bootstrap, and wait until everything is over to work
+ * with the new NBuffers value.
+ *
+ * - Old backends can exit before attempting to resize. Synchronization
+ * used between backends relies on ProcSignalBarrier and waits for all
+ * participants received the message at the beginning to gather all
+ * existing backends.
+ *
+ * - Some backends might be blocked and not responsing either before or
+ * after receiving the message. In the first case such backend still
+ * have ProcSignalSlot and should be waited for, in the second case
+ * shared barrier will make sure we still waiting for those backends. In
+ * any case there is an unbounded wait.
+ *
+ * - Backends might join barrier in disjoint groups with some time in
+ * between. That means that relying only on the shared dynamic barrier is
+ * not enough -- it will only synchronize resize procedure withing those
+ * groups. That's why we wait first for all participants of ProcSignal
+ * mechanism who received the message.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ pg_atomic_init_u64(&ShmemCtrl->Generation,
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
+
+ /* To order everything after setting Generation value */
+ pg_memory_barrier();
+
+ /*
+ * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
+ * the rest of the pack has started the procedure and it can resize shared
+ * memory as well.
+ *
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier. In fact even if we would like to
+ * wait here, it wouldn't be possible -- we're in the postmaster, without
+ * any waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1271,3 +1637,50 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier()
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ /* Nothing to do if resizing is not started */
+ if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
+ return;
+
+ BarrierAttach(barrier);
+
+ /* Otherwise wait through all available phases */
+ while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
+ {
+ ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
+ BarrierPhase(barrier))));
+
+ BarrierArriveAndWait(barrier, 0);
+ }
+
+ BarrierDetach(barrier);
+}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /*
+ * The barrier is missing here, it will be initialized right before
+ * starting the resizing process as a convenient way to reset it.
+ */
+
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+
+ /* shmem_resizable should be initialized by now */
+ ShmemCtrl->Resizable = shmem_resizable;
+ }
+}
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 3fe45de5da0..196f233fe0e 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -425,6 +425,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1693,6 +1694,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2038,6 +2042,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
@@ -3851,6 +3866,9 @@ process_pm_pmsignal(void)
request_state_update = true;
}
+ if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
+ AnonymousShmemResize();
+
/*
* Try to advance postmaster's state machine, if a child requests it.
*/
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index bd68b69ee98..ac844b114bd 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -24,7 +25,6 @@ ConditionVariableMinimallyPadded *BufferIOCVArray;
WritebackContext BackendWritebackContext;
CkptSortItem *CkptBufferIds;
-
/*
* Data Structures:
* buffers live in a freelist and a lookup data structure.
@@ -62,18 +62,28 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * postmaster, or in a standalone backend) or during shared-memory resize. Size
+ * of data structures initialized here depends on NBuffers, and to be able to
+ * change NBuffers without a restart we store each structure into a separate
+ * shared memory segment, which could be resized on demand.
+ *
+ * FirstBufferToInit tells where to start initializing buffers. For
+ * initialization it always will be zero, but when resizing shared-memory it
+ * indicates the number of already initialized buffers.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(void)
+BufferManagerShmemInit(int FirstBufferToInit)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
+ elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
+ FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -110,43 +120,44 @@ BufferManagerShmemInit(void)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory, or in all cases when resizing shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
- ClearBufferTag(&buf->tag);
+#ifndef EXEC_BACKEND
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = FirstBufferToInit; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ ClearBufferTag(&buf->tag);
- buf->buf_id = i;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- pgaio_wref_clear(&buf->io_wref);
+ buf->buf_id = i;
- /*
- * Initially link all the buffers together as unused. Subsequent
- * management of this list is done by freelist.c.
- */
- buf->freeNext = i + 1;
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
+#endif
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 9d00b80b4f8..abeb91e24fd 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -84,6 +84,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -151,6 +154,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, AioShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -298,7 +309,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit();
+ BufferManagerShmemInit(0);
/*
* Set up lock manager
@@ -310,6 +321,11 @@ CreateOrAttachShmemStructs(void)
*/
PredicateLockShmemInit();
+ /*
+ * Set up shared memory resize manager
+ */
+ ShmemControlInit();
+
/*
* Set up process table
*/
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 8e313ad9bf8..35c42f260a8 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -112,6 +113,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -175,6 +180,43 @@ ProcSignalInit(char *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -614,6 +656,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 389abc82519..0fd421f004e 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -493,17 +493,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
+ *
+ * XXX: There is an implicit assumption this can only happen in
+ * "resizable" segments, where only one shared structure is allowed.
+ * This has to be implemented more cleanly.
*/
if (result->size != size)
{
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
+ Size delta = size - result->size;
+
+ result->size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
}
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index b1fba850f02..58f1a05fd2a 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4311,6 +4312,15 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /* Verify the shared barrier, if it's still active: join and wait. */
+ WaitOnShmemBarrier();
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 8bce14c38fd..e0ba8384fdd 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
@@ -351,6 +353,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index dede37f7905..1e70853ccdb 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2354,14 +2354,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 0d8528b2875..405d0a7e65d 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int MaxAvailableMemory;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 1977001e533..52633dd7537 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -307,7 +307,7 @@ extern void LimitAdditionalLocalPins(uint32 *additional_pins);
extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
-extern void BufferManagerShmemInit(void);
+extern void BufferManagerShmemInit(int);
extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6ebda479ced..bb7ae4d33b3 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index a9681738146..558da6fdd55 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
+PG_LWLOCK(54, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index c5009a1cd73..2e47b222cbb 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,25 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+ pg_atomic_uint64 Generation;
+ bool Resizable;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -107,6 +127,12 @@ extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
void *ReserveAnonymousMemory(Size reserve_size);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 67fa9ac06e1..27bc6a81191 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,6 +42,7 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index defd8b66a19..522b8de1e02 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 1a30437ad96..6755b302858 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2738,6 +2738,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.45.1
From 0e3c671082743f2826a7e8a96a19a071f5c8aeb3 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 15 Mar 2025 16:39:45 +0100
Subject: [PATCH v4 7/8] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f90cde00000-7f90d5126000 rw-s 00000000 00:01 5463 /memfd:main (deleted)
7f90d5126000-7f914de00000 ---p 00000000 00:00 0
7f914de00000-7f9175128000 rw-s 00000000 00:01 5466 /memfd:buffers (deleted)
7f9175128000-7f944de00000 ---p 00000000 00:00 0
7f944de00000-7f9455528000 rw-s 00000000 00:01 5469 /memfd:descriptors (deleted)
7f9455528000-7f94cde00000 ---p 00000000 00:00 0
7f94cde00000-7f94d5228000 rw-s 00000000 00:01 5472 /memfd:iocv (deleted)
7f94d5228000-7f954de00000 ---p 00000000 00:00 0
7f954de00000-7f9555266000 rw-s 00000000 00:01 5475 /memfd:checkpoint (deleted)
7f9555266000-7f958de00000 ---p 00000000 00:00 0
7f958de00000-7f95954aa000 rw-s 00000000 00:01 5478 /memfd:strategy (deleted)
7f95954aa000-7f95cde00000 ---p 00000000 00:00 0
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 73 +++++++++++++++++++++++++++++-----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/portability/mem.h | 2 +-
src/include/storage/pg_shmem.h | 3 +-
5 files changed, 68 insertions(+), 14 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a3437973784..87000a24eea 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -107,6 +107,7 @@ typedef struct AnonymousMapping
Pointer shmem; /* Pointer to the start of the mapped memory */
Pointer seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -127,7 +128,7 @@ static int next_free_segment = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -150,9 +151,9 @@ static int next_free_segment = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* 7f442e010000-7f443a800000 # reserved empty space
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* 7f444196c000-7f470a800000 # reserved empty space
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -643,13 +644,14 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* *hugepagesize and *mmap_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -708,6 +710,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -718,7 +721,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -727,6 +739,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -734,6 +748,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -771,7 +787,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
- int mmap_flags = PG_MMAP_FLAGS;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
#ifndef MAP_HUGETLB
/* ReserveAnonymousMemory should have dealt with this case */
@@ -785,7 +801,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
/* Round up the request size to a suitable large value */
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
@@ -794,6 +810,29 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
}
#endif
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment),
+ memfd_flags);
+
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(mapping->shmem_segment))));
+
elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p",
MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
@@ -807,7 +846,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
* a restart.
*/
ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | MAP_FIXED, -1, 0);
+ mmap_flags | MAP_FIXED, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
@@ -817,8 +856,15 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
"fallback to the non-resizable allocation",
MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+ /* Specify the segment file size using allocsize. */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(mapping->shmem_segment))));
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
else
@@ -889,7 +935,7 @@ ReserveAnonymousMemory(Size reserve_size)
Size hugepagesize, total_size = 0;
int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
/*
* Figure out how much memory is needed for all segments, keeping in
@@ -1070,6 +1116,13 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ if(ftruncate(m->segment_fd, new_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
/* Clean up some reserved space to resize into */
if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
ereport(FATAL,
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index ce719f1b412..ba972106de1 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index abeb91e24fd..dc2b4becf4a 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -396,7 +396,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 2e47b222cbb..b9573520d9a 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -124,7 +124,8 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size,
PGShmemHeader **shim, Pointer base);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
void *ReserveAnonymousMemory(Size reserve_size);
bool ProcessBarrierShmemResize(Barrier *barrier);
--
2.45.1
From 08476af71724fcb3035fc907dc98a6ff351fe58e Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 5 Apr 2025 19:51:33 +0200
Subject: [PATCH v4 8/8] Support resize for hugetlb
Linux kernel has a set of limitations on remapping hugetlb segments: it
can't increase size of such segment [1], and shrinking it will not
release the memory back. In fact support for hugetlb mremap was
implemented no so long time ago [2].
As a workaround, avoid mremap for resizing shared memory. Instead unmap
the whole segment and map it back at the same address with the new size,
relying on the fact that fd for the anon file behind the segment is
still open and will keep the memory content.
[1]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/mremap.c?id=f4d2ef482...
[2]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/mm/mremap.c?id=550a7d6...
---
src/backend/port/sysv_shmem.c | 60 +++++++++++++++++++++++++----------
1 file changed, 44 insertions(+), 16 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 87000a24eea..f0b53ce1d7c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1109,6 +1109,7 @@ AnonymousShmemResize(void)
/* Note that CalculateShmemSize indirectly depends on NBuffers */
Size new_size = CalculateShmemSize(&numSemas, i);
AnonymousMapping *m = &Mappings[i];
+ int mmap_flags = PG_MMAP_FLAGS;
if (m->shmem == NULL)
continue;
@@ -1116,6 +1117,44 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+#ifndef MAP_HUGETLB
+ /* ReserveAnonymousMemory should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ Size hugepagesize;
+
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the new size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ if (new_size % hugepagesize != 0)
+ new_size += hugepagesize - (new_size % hugepagesize);
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
+ }
+#endif
+
+ /*
+ * Linux limitations do not allow us to mremap hugetlb in the way we
+ * want. E.g. no size increase is allowed, and for shrinking the memory
+ * will not be released back. To work around this unmap the segment and
+ * create a new one at the same address. Thanks for the backing anon
+ * file the content will still be kept in memory.
+ */
+ elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ if (munmap(m->shmem, m->shmem_size) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not unmap shared memory segment %s [%p]: %m",
+ MappingName(m->shmem_segment), m->shmem)));
+
/* Resize the backing anon file. */
if(ftruncate(m->segment_fd, new_size) == -1)
ereport(FATAL,
@@ -1123,25 +1162,14 @@ AnonymousShmemResize(void)
errmsg("could not truncase anonymous file for \"%s\": %m",
MappingName(m->shmem_segment))));
- /* Clean up some reserved space to resize into */
- if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
- ereport(FATAL,
- (errcode(ERRCODE_SYSTEM_ERROR),
- errmsg("could not unmap %zu from reserved shared memory %p: %m",
- new_size - m->shmem_size, m->shmem)));
-
- /* Claim the unused space */
- elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
- MappingName(m->shmem_segment), m->shmem_size,
- new_size, m->shmem);
-
- ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
+ /* Reclaim the space */
+ ptr = mmap(m->shmem, new_size, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_FIXED, m->segment_fd, 0);
if (ptr == MAP_FAILED)
ereport(FATAL,
(errcode(ERRCODE_SYSTEM_ERROR),
- errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m",
- MappingName(m->shmem_segment), m->shmem, NBuffers,
- new_size)));
+ errmsg("could not map shared memory segment %s [%p] with size %zu: %m",
+ MappingName(m->shmem_segment), m->shmem, new_size)));
reinit = true;
m->shmem_size = new_size;
--
2.45.1
Attachments:
[text/plain] v4-0001-Allow-to-use-multiple-shared-memory-mappings.patch (29.9K, ../../eqs6v4rsboazl67xz3wxc6xjkgrpfybitpl45y3lmb2br67wbj@o7czebb3rlgd/2-v4-0001-Allow-to-use-multiple-shared-memory-mappings.patch)
download | inline diff:
From 15b87a1cb89d3f31b656e27d07ead5aa935f1643 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH v4 1/8] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 ++++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 141 +++++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 13 ++-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
12 files changed, 278 insertions(+), 125 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index f7c8638aec5..b6301463ac7 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -313,7 +313,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -334,7 +334,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..56af0231d24 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 567739b5be9..5b55bec8d9d 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2fa045e6b0f..8b38e985327 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -363,7 +368,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 895a43fb39e..389abc82519 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -75,19 +75,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/*
@@ -96,9 +96,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -109,7 +117,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -118,9 +132,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -145,11 +159,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -179,6 +199,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -198,22 +224,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -231,6 +257,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -241,19 +273,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -268,7 +300,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -329,6 +367,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -345,9 +395,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -380,6 +430,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -388,7 +445,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -411,7 +468,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -428,8 +485,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -454,7 +511,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -473,14 +530,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -537,10 +593,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -552,15 +609,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 3df29658f18..8241c061507 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -618,10 +620,15 @@ LWLockNewTrancheId(void)
int *LWLockCounter;
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
- /* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ /*
+ * We use the ShmemLock spinlock to protect LWLockCounter.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
+ */
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..6ebda479ced 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index b99ebc9e86f..138078c29c5 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 904a336b851..5929f140236 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: 5e1915439085014140314979c4dd5e23bd677cac
--
2.45.1
[text/plain] v4-0002-Address-space-reservation-for-shared-memory.patch (21.6K, ../../eqs6v4rsboazl67xz3wxc6xjkgrpfybitpl45y3lmb2br67wbj@o7czebb3rlgd/3-v4-0002-Address-space-reservation-for-shared-memory.patch)
download | inline diff:
From eae77d430e6e6cc3ec95b2cf613e4b3ae095e75e Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH v4 2/8] Address space reservation for shared memory
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is:
* To reserve some address space via mmap'ing a large chunk of memory
with PROT_NONE and MAP_NORESERVE. This way we prepare a playground for
preparing shared memory layout without risking anything interfering
with that.
* To slice the reserved space up into sections, one to use for each
shared segment.
* Allocate shared memory segments out of corresponding slices and
leaving unclaimed space in between them. This is implemented via
mmap'ing memory at a specified address from the reserved space with
MAP_FIXED.
The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
7f444196c000-7f470a800000 # reserved space
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
Things like address space randomization should not be a problem in this
context, since the randomization is applied to the mmap base, which is
one per process.
This approach also do not impact the actual memory usage as reported by
the kernel. Here is the output of /proc/$PID/status for the master
version with shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages
// mapped in mm_struct. It corresponds to the mapped reserved space
// and is the only number that grows with it.
VmPeak: 2043192 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 22908 kB
// Size of resident anonymous memory
RssAnon: 768 kB
// Size of resident file mappings
RssFile: 10364 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 11776 kB
Here is the same for the patch when reserving 20GB of space:
VmPeak: 21250648 kB
VmRSS: 22948 kB
RssAnon: 768 kB
RssFile: 10404 kB
RssShmem: 11776 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
17637376 (~16.8 MB)
To control the amount of space reserved a new GUC max_available_memory
is introduced. Ideally it should be based on the maximum available
memory, hense the name.
---
src/backend/port/sysv_shmem.c | 284 ++++++++++++++++++++++++----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 5 +-
src/backend/utils/init/globals.c | 1 +
src/backend/utils/misc/guc_tables.c | 14 ++
src/include/storage/pg_shmem.h | 4 +-
6 files changed, 271 insertions(+), 39 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 56af0231d24..a0f03ff868f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,66 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * To achieve that we first reserve some shared memory address space by
+ * mmap'ing a segment of MaxAvailableMemory size with PROT_NONE and
+ * MAP_NORESERVE (these flags allow to make sure this space will not be used by
+ * anything else, yet do not count against memory limits). Having the reserved
+ * space, we allocate out of it actual chunks of shared memory as usual,
+ * updating a pointer to the current available reserved space for the next
+ * allocation with the gap between segments in mind.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f442e010000-7f443a800000 # reserved empty space
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f444196c000-7f470a800000 # reserved empty space
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * [...]
+ *
+ * The reserved space pointer is calculated to slice up the total reserved
+ * space into fixed fractions of address space for each segment, as specified
+ * in the SHMEM_RESIZE_RATIO array.
+ */
+static double SHMEM_RESIZE_RATIO[1] = {
+ 1.0, /* MAIN_SHMEM_SLOT */
+};
+
+/*
+ * Offset from the beginning of the reserved space, which indicates currently
+ * available range. New shared memory segments have to be allocated at this
+ * offset related to the reserved space.
+ */
+static Size reserved_offset = 0;
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -626,39 +686,198 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
*
* This function will modify mapping size to the actual size of the allocation,
* if it ends up allocating a segment that is larger than requested.
+ *
+ * Note that we do not switch from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
*/
static void
-CreateAnonymousSegment(AnonymousMapping *mapping)
+CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
{
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS;
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* ReserveAnonymousMemory should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the request size to a suitable large value */
GetHugePageSize(&hugepagesize, &mmap_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
+ }
+#endif
+
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p",
+ MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+
+ /*
+ * Try to create mapping at an address out of the reserved range, which
+ * will allow to extend it later. Use reserved_offset to allocate the
+ * segment, then update currently available reserved range.
+ *
+ * If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized without
+ * a restart.
+ */
+ ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_FIXED, -1, 0);
+ mmap_errno = errno;
+
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m, "
+ "fallback to the non-resizable allocation",
+ MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS, -1, 0);
+ mmap_errno = errno;
+ }
+ else
+ {
+ Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+
+ reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ errno = mmap_errno;
+ DebugMappings();
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
+ (mmap_errno == ENOMEM) ?
+ errhint("This error usually means that PostgreSQL's request "
+ "for a shared memory segment exceeded available memory, "
+ "swap space, or huge pages. To reduce the request size "
+ "(currently %zu bytes), reduce PostgreSQL's shared "
+ "memory usage, perhaps by reducing \"shared_buffers\" or "
+ "\"max_connections\".",
+ allocsize) : 0));
+ }
+
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
+}
+
+/*
+ * ReserveAnonymousMemory
+ *
+ * Reserve shared memory address space, from which shared memory segments are
+ * going to be sliced out. The goal of this exercise is to support segments
+ * resizing, for which we need a reserved space free of potential clashes with
+ * other mmap'd areas that are not under our control. Reservation is done via
+ * mmap, and will not allocate any memory until it will be actually used, and
+ * MAP_NORESERVE allows to make it not counting againt kernel reservation
+ * limits (e.g. in cgroups or for huge pages). Do not get confused because of
+ * MAP_NORESERVE -- we need to reserve some space, but not the actual memory,
+ * and that is that this flag is about.
+ *
+ * Note, that with MAP_NORESERVE a reservation with hugetlb will succeed even
+ * if there is actually not enough huge pages. Hence this function is
+ * responsible for deciding whether to use huge pages or not. To achieve that
+ * we need to probe first and try to allocate needed memory for all segments --
+ * if this succeeds, we unmap the probe segment and use hugetlb; if it fails,
+ * we proceed with the regular memory.
+ */
+void *
+ReserveAnonymousMemory(Size reserve_size)
+{
+ Size allocsize = reserve_size;
+ void *ptr = MAP_FAILED;
+ int mmap_errno = 0;
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ *
+ * We could actually have a mix and match of segments with and without
+ * huge pages. But in that case we need to have multiple reservation
+ * spaces to use corresponding memory (hugetlb adress space reserved
+ * for hugetlb segments, regular memory for others), and it doesn't
+ * seem to worth the complexity for now.
+ */
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ int numSemas;
+ Size segment_size = CalculateShmemSize(&numSemas, segment);
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
- DebugMappings();
- elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- MappingName(mapping->shmem_segment), allocsize);
+ /* No huge pages, we will go with the regular page size */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB "
+ "failed, huge pages disabled: %m", total_size);
+ }
+ else
+ {
+ /*
+ * All fine, unmap the temporary segment and proceed with reserving
+ * using huge pages.
+ */
+ if (munmap(ptr, total_size) < 0)
+ elog(LOG, "reservice space: munmap(%p, %zu) failed: %m",
+ ptr, total_size);
+
+ /* Round up the requested size to a suitable large value. */
+ if (allocsize % hugepagesize != 0)
+ allocsize += hugepagesize - (allocsize % hugepagesize);
+
+ elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB",
+ allocsize);
+ ptr = mmap(NULL, allocsize, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | MAP_NORESERVE | mmap_flags,
+ -1, 0);
+ mmap_errno = errno;
+
+ /* This should not happen, but handle errors anyway */
+ if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
+ {
+ elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB "
+ "failed, huge pages disabled: %m", allocsize);
+ }
}
}
#endif
@@ -666,10 +885,12 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
/*
* Report whether huge pages are in use. This needs to be tracked before
* the second mmap() call if attempting to use huge pages failed
- * previously.
+ * previously. At this point ptr is either pointing to the probe segment,
+ * if we couldn't mmap it, or the reservation space.
*/
SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
@@ -677,10 +898,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
+ allocsize = reserve_size;
+
+ elog(DEBUG1, "reserving space: mmap(%zu)", allocsize);
+ ptr = mmap(NULL, allocsize, PROT_NONE,
+ MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0);
}
if (ptr == MAP_FAILED)
@@ -688,20 +910,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
errno = mmap_errno;
DebugMappings();
ereport(FATAL,
- (errmsg("segment[%s]: could not map anonymous shared memory: %m",
- MappingName(mapping->shmem_segment)),
+ (errmsg("reserving space: could not map anonymous shared "
+ "memory: %m"),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
- "for a shared memory segment exceeded available memory, "
- "swap space, or huge pages. To reduce the request size "
- "(currently %zu bytes), reduce PostgreSQL's shared "
- "memory usage, perhaps by reducing \"shared_buffers\" or "
- "\"max_connections\".",
+ "for a reserved shared memory address space exceeded "
+ "available memory, swap space, or huge pages. To "
+ "reduce the request reservation size (currently %zu "
+ "bytes), reduce PostgreSQL's \"maximum_shared_buffers\".",
allocsize) : 0));
}
- mapping->shmem = ptr;
- mapping->shmem_size = allocsize;
+ return ptr;
}
/*
@@ -740,7 +960,7 @@ AnonymousShmemDetach(int status, Datum arg)
*/
PGShmemHeader *
PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim)
+ PGShmemHeader **shim, Pointer base)
{
IpcMemoryKey NextShmemSegID;
void *memAddress;
@@ -760,14 +980,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -782,7 +994,7 @@ PGSharedMemoryCreate(Size size,
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
/* On success, mapping data will be modified. */
- CreateAnonymousSegment(mapping);
+ CreateAnonymousSegment(mapping, base);
next_free_segment++;
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..ce719f1b412 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -205,7 +205,7 @@ EnableLockPagesPrivilege(int elevel)
*/
PGShmemHeader *
PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim)
+ PGShmemHeader **shim, Pointer base)
{
void *memAddress;
PGShmemHeader *hdr;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8b38e985327..076888c0172 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -203,9 +203,12 @@ CreateSharedMemoryAndSemaphores(void)
PGShmemHeader *seghdr;
Size size;
int numSemas;
+ void *base;
Assert(!IsUnderPostmaster);
+ base = ReserveAnonymousMemory((Size) MaxAvailableMemory * BLCKSZ);
+
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
/* Compute the size of the shared-memory block */
@@ -217,7 +220,7 @@ CreateSharedMemoryAndSemaphores(void)
*
* XXX: Do multiple shims are needed, one per segment?
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(size, &shim, base);
/*
* Make sure that huge pages are never reported as "unknown" while the
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index 2152aad97d9..1d42a5856c0 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffers = 16384;
+int MaxAvailableMemory = 131072;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 4eaeca89f2c..dede37f7905 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2364,6 +2364,20 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"max_available_memory", PGC_SIGHUP, RESOURCES_MEM,
+ gettext_noop("Sets the upper limit for the shared_buffers value."),
+ gettext_noop("Shared memory could be resized at runtime, this "
+ "parameters sets the upper limit for it, beyond which "
+ "resizing would not be supported. Normally this value "
+ "would be the same as the total available memory."),
+ GUC_UNIT_BLOCKS
+ },
+ &MaxAvailableMemory,
+ 131072, 16, INT_MAX / 2,
+ NULL, NULL, NULL
+ },
+
{
{"vacuum_buffer_usage_limit", PGC_USERSET, RESOURCES_MEM,
gettext_noop("Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum."),
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 138078c29c5..4a83e255652 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -60,6 +60,7 @@ extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
+extern PGDLLIMPORT int MaxAvailableMemory;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -100,10 +101,11 @@ extern void PGSharedMemoryNoReAttach(void);
#endif
extern PGShmemHeader *PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim);
+ PGShmemHeader **shim, Pointer base);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+void *ReserveAnonymousMemory(Size reserve_size);
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
--
2.45.1
[text/plain] v4-0003-Introduce-multiple-shmem-segments-for-shared-buff.patch (11.7K, ../../eqs6v4rsboazl67xz3wxc6xjkgrpfybitpl45y3lmb2br67wbj@o7czebb3rlgd/4-v4-0003-Introduce-multiple-shmem-segments-for-shared-buff.patch)
download | inline diff:
From dca1257476fc4c0718fec35b11ba0f4a4e57151b Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 15 Mar 2025 16:38:59 +0100
Subject: [PATCH v4 3/8] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 24 +++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 105 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a0f03ff868f..f46d9d5d9cd 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -147,10 +147,18 @@ static int next_free_segment = 0;
*
* The reserved space pointer is calculated to slice up the total reserved
* space into fixed fractions of address space for each segment, as specified
- * in the SHMEM_RESIZE_RATIO array.
+ * in the SHMEM_RESIZE_RATIO array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take
+ * up to 60% of the whole space when resizing, based on the fact that it most
+ * likely will be the main consumer of this memory. Those numbers are pulled
+ * out of thin air for now, makes sense to evaluate them more precise.
*/
-static double SHMEM_RESIZE_RATIO[1] = {
- 1.0, /* MAIN_SHMEM_SLOT */
+static double SHMEM_RESIZE_RATIO[6] = {
+ 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.6, /* BUFFERS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -182,6 +190,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1dc488a42..bd68b69ee98 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +77,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -156,33 +160,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..a9952b36eba 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -22,6 +22,7 @@
#include "postgres.h"
#include "storage/buf_internals.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -59,10 +60,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 336715b6c63..81543cb5ced 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -491,9 +492,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 076888c0172..9d00b80b4f8 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index f2192ceb271..1977001e533 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -308,7 +308,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 4a83e255652..c5009a1cd73 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -107,7 +107,29 @@ extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
void *ReserveAnonymousMemory(Size reserve_size);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.45.1
[text/plain] v4-0004-Introduce-pending-flag-for-GUC-assign-hooks.patch (12.9K, ../../eqs6v4rsboazl67xz3wxc6xjkgrpfybitpl45y3lmb2br67wbj@o7czebb3rlgd/5-v4-0004-Introduce-pending-flag-for-GUC-assign-hooks.patch)
download | inline diff:
From 24704e57aea0ee94fbfb37ca3a4ea4fcf050a738 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:40:32 +0200
Subject: [PATCH v4 4/8] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 61 insertions(+), 40 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index ec40c0b7c42..9aa426992a2 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2321,7 +2321,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index a9f2a3a3062..e715a6f01c2 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,13 +1161,13 @@ assign_maintenance_io_concurrency(int newval, void *extra)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra)
+assign_io_max_combine_limit(int newval, void *extra, bool *pending)
{
io_max_combine_limit = newval;
io_combine_limit = Min(io_max_combine_limit, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra)
+assign_io_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit_guc = newval;
io_combine_limit = Min(io_max_combine_limit, io_combine_limit_guc);
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index e5171467de1..2a6a587ef76 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1952,7 +1952,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1985,7 +1985,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2008,7 +2008,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2031,7 +2031,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 6ae9f38f0c8..b1fba850f02 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3593,7 +3593,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 667df448732..bb681f5bc60 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f619100467d..8802ad8a3cb 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 799fa7ace68..c8300cffa8e 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,14 +81,14 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
-extern void assign_io_max_combine_limit(int newval, void *extra);
-extern void assign_io_combine_limit(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
+extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -143,13 +143,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -165,7 +165,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.45.1
[text/plain] v4-0005-Introduce-pss_barrierReceivedGeneration.patch (7.3K, ../../eqs6v4rsboazl67xz3wxc6xjkgrpfybitpl45y3lmb2br67wbj@o7czebb3rlgd/6-v4-0005-Introduce-pss_barrierReceivedGeneration.patch)
download | inline diff:
From 619b10ec409185995a4a3ffd56972f1efa493c45 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 4 Apr 2025 21:46:14 +0200
Subject: [PATCH v4 5/8] Introduce pss_barrierReceivedGeneration
Currently WaitForProcSignalBarrier allows to make sure the message sent
via EmitProcSignalBarrier was processed by all ProcSignal mechanism
participants.
Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
which will be updated when a process has received the message, but not
processed it yet. This makes it possible to support a new mode of
waiting, when ProcSignal participants want to synchronize message
processing. To do that, a participant can wait via
WaitForProcSignalBarrierReceived when processing a message, effectively
making sure that all processes are going to start processing
ProcSignalBarrier simultaneously.
---
src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------
src/include/storage/procsignal.h | 1 +
2 files changed, 54 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index b7c39a4c5f0..8e313ad9bf8 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -58,7 +58,10 @@
* of it. For such use cases, we set a bit in pss_barrierCheckMask and then
* increment the current "barrier generation"; when the new barrier generation
* (or greater) appears in the pss_barrierGeneration flag of every process,
- * we know that the message has been received everywhere.
+ * we know that the message has been received and processed everywhere. In case
+ * if we only need to know only that the message was received everywhere (e.g.
+ * receiving processes need to handle the message in a coordinated fashion)
+ * use pss_barrierReceivedGeneration in the same way.
*/
typedef struct
{
@@ -70,6 +73,7 @@ typedef struct
/* Barrier-related fields (not protected by pss_mutex) */
pg_atomic_uint64 pss_barrierGeneration;
+ pg_atomic_uint64 pss_barrierReceivedGeneration;
pg_atomic_uint32 pss_barrierCheckMask;
ConditionVariable pss_barrierCV;
} ProcSignalSlot;
@@ -151,6 +155,8 @@ ProcSignalShmemInit(void)
slot->pss_cancel_key_len = 0;
MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration,
+ PG_UINT64_MAX);
pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
ConditionVariableInit(&slot->pss_barrierCV);
}
@@ -198,6 +204,8 @@ ProcSignalInit(char *cancel_key, int cancel_key_len)
barrier_generation =
pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration);
pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration,
+ barrier_generation);
if (cancel_key_len > 0)
memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len);
@@ -262,6 +270,7 @@ CleanupProcSignalState(int status, Datum arg)
* no barrier waits block on it.
*/
pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX);
SpinLockRelease(&slot->pss_mutex);
@@ -415,12 +424,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type)
return generation;
}
-/*
- * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
- * requested by a specific call to EmitProcSignalBarrier() have taken effect.
- */
-void
-WaitForProcSignalBarrier(uint64 generation)
+static void
+WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly)
{
Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration));
@@ -435,12 +440,17 @@ WaitForProcSignalBarrier(uint64 generation)
uint64 oldval;
/*
- * It's important that we check only pss_barrierGeneration here and
- * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared
- * before the barrier is actually absorbed, but pss_barrierGeneration
+ * It's important that we check only pss_barrierGeneration &
+ * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in
+ * pss_barrierCheckMask get cleared before the barrier is actually
+ * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration
* is updated only afterward.
*/
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
while (oldval < generation)
{
if (ConditionVariableTimedSleep(&slot->pss_barrierCV,
@@ -449,7 +459,11 @@ WaitForProcSignalBarrier(uint64 generation)
ereport(LOG,
(errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier",
(int) pg_atomic_read_u32(&slot->pss_pid))));
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
}
ConditionVariableCancelSleep();
}
@@ -463,12 +477,33 @@ WaitForProcSignalBarrier(uint64 generation)
* The caller is probably calling this function because it wants to read
* the shared state or perform further writes to shared state once all
* backends are known to have absorbed the barrier. However, the read of
- * pss_barrierGeneration was performed unlocked; insert a memory barrier
- * to separate it from whatever follows.
+ * pss_barrierGeneration & pss_barrierReceivedGeneration was performed
+ * unlocked; insert a memory barrier to separate it from whatever follows.
*/
pg_memory_barrier();
}
+/*
+ * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
+ * requested by a specific call to EmitProcSignalBarrier() have taken effect.
+ */
+void
+WaitForProcSignalBarrier(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, false);
+}
+
+/*
+ * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all
+ * backends have observed the message sent by a specific call to
+ * EmitProcSignalBarrier().
+ */
+void
+WaitForProcSignalBarrierReceived(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, true);
+}
+
/*
* Handle receipt of an interrupt indicating a global barrier event.
*
@@ -522,6 +557,10 @@ ProcessProcSignalBarrier(void)
if (local_gen == shared_gen)
return;
+ /* The message is observed, record that */
+ pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
+ shared_gen);
+
/*
* Get and clear the flags that are set for this backend. Note that
* pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 016dfd9b3f6..defd8b66a19 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, char *cancel_key, int cancel_key_l
extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type);
extern void WaitForProcSignalBarrier(uint64 generation);
+extern void WaitForProcSignalBarrierReceived(uint64 generation);
extern void ProcessProcSignalBarrier(void);
extern void procsignal_sigusr1_handler(SIGNAL_ARGS);
--
2.45.1
[text/plain] v4-0006-Allow-to-resize-shared-memory-without-restart.patch (39.8K, ../../eqs6v4rsboazl67xz3wxc6xjkgrpfybitpl45y3lmb2br67wbj@o7czebb3rlgd/7-v4-0006-Allow-to-resize-shared-memory-without-restart.patch)
download | inline diff:
From 886a3ea87408e628bea08a9c77116343616ad032 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:47:16 +0200
Subject: [PATCH v4 6/8] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers and extend it using mremap. One elected process signals the
postmaster to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f90cde00000-7f90d4fa6000 /dev/zero (deleted)
7f90d4fa6000-7f914de00000
7f914de00000-7f915cfa8000 /dev/zero (deleted)
^ buffers mapping, ~241 MB
7f915cfa8000-7f944de00000
7f944de00000-7f94550a8000 /dev/zero (deleted)
7f94550a8000-7f94cde00000
7f94cde00000-7f94d4fe8000 /dev/zero (deleted)
7f94d4fe8000-7f954de00000
7f954de00000-7f9554ff6000 /dev/zero (deleted)
7f9554ff6000-7f958de00000
7f958de00000-7f959508a000 /dev/zero (deleted)
7f959508a000-7f95cde00000
-- 512 MB
7f90cde00000-7f90d5126000 /dev/zero (deleted)
7f90d5126000-7f914de00000
7f914de00000-7f9175128000 /dev/zero (deleted)
^ buffers mapping, ~627 MB
7f9175128000-7f944de00000
7f944de00000-7f9455528000 /dev/zero (deleted)
7f9455528000-7f94cde00000
7f94cde00000-7f94d5228000 /dev/zero (deleted)
7f94d5228000-7f954de00000
7f954de00000-7f9555266000 /dev/zero (deleted)
7f9555266000-7f958de00000
7f958de00000-7f95954aa000 /dev/zero (deleted)
7f95954aa000-7f95cde00000
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it receives something.
Note, that mremap is Linux specific, thus the implementation not very
portable.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 413 ++++++++++++++++++
src/backend/postmaster/postmaster.c | 18 +
src/backend/storage/buffer/buf_init.c | 75 ++--
src/backend/storage/ipc/ipci.c | 18 +-
src/backend/storage/ipc/procsignal.c | 46 ++
src/backend/storage/ipc/shmem.c | 23 +-
src/backend/tcop/postgres.c | 10 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/miscadmin.h | 1 +
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 26 ++
src/include/storage/pmsignal.h | 1 +
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
17 files changed, 603 insertions(+), 43 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index f46d9d5d9cd..a3437973784 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -105,6 +111,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -176,6 +189,49 @@ static Size reserved_offset = 0;
*/
static bool huge_pages_on = false;
+/*
+ * Flag telling that we have prepared the memory layout to be resizable. If
+ * false after all shared memory segments creation, it means we failed to setup
+ * needed layout and falled back to the regular non-resizable approach.
+ */
+static bool shmem_resizable = false;
+
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -769,6 +825,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
{
Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+ shmem_resizable = true;
reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
}
@@ -964,6 +1021,315 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * mremap, which will fail if is not sufficient space to expand the mapping.
+ * When finished, based on the new and old values initialize new buffer blocks
+ * if any.
+ *
+ * If reinitializing took place, as the last step this function does buffers
+ * reinitialization as well and broadcasts the new value of NSharedBuffers. All
+ * of that needs to be done only by one backend, the first one that managed to
+ * grab the ShmemResizeLock.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ void *ptr = MAP_FAILED;
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ /* Clean up some reserved space to resize into */
+ if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not unmap %zu from reserved shared memory %p: %m",
+ new_size - m->shmem_size, m->shmem)));
+
+ /* Claim the unused space */
+ elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
+ if (ptr == MAP_FAILED)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m",
+ MappingName(m->shmem_segment), m->shmem, NBuffers,
+ new_size)));
+
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * Note that it's enough when only one backend will do that,
+ * even the ShmemInitStruct part. The reason is that resized
+ * shared memory will maintain the same addresses, meaning that
+ * all the pointers are still valid, and we only need to update
+ * structures size in the ShmemIndex once -- any other backend
+ * will pick up this shared structure from the index.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ BufferManagerShmemInit(NBuffersOld);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Wait for all ProcSignal participants
+ * to join the barrier, then do the resize and wait on the barrier until all
+ * participating finish resizing as well -- otherwise we face danger of
+ * inconsistency between backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * First thing to do after attaching to the barrier is to wait for others.
+ * We can't simply use BarrierArriveAndWait, because backends might arrive
+ * here in disjoint groups, e.g. first two backends, pause, then second two
+ * backends. If the resize is quick enough that can lead to a situation
+ * when the first group is already finished before the second has appeared,
+ * and the barrier will only synchonize withing those groups.
+ */
+ if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
+ WaitForProcSignalBarrierReceived(
+ pg_atomic_read_u64(&ShmemCtrl->Generation));
+
+ /*
+ * Now start the procedure, and elect one backend to ping postmaster to do
+ * the same.
+ *
+ * XXX: If we need to be able to abort resizing, this has to be done later,
+ * after the SHMEM_RESIZE_DONE.
+ */
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ {
+ Assert(IsUnderPostmaster);
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ }
+
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ BarrierDetach(barrier);
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ elog(DEBUG1, "Set pending signal");
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Start resizing procedure, making sure all existing processes will have
+ * consistent view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * We use dynamic barrier to help dealing with backends that were spawned
+ * during the resize.
+ */
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ *
+ * Those phases are ensured by joining the shared barrier associated with
+ * the procedure. Since resizing takes time, we need to take into account
+ * that during that time:
+ *
+ * - New backends can be spawned. They will check status of the barrier
+ * early during the bootstrap, and wait until everything is over to work
+ * with the new NBuffers value.
+ *
+ * - Old backends can exit before attempting to resize. Synchronization
+ * used between backends relies on ProcSignalBarrier and waits for all
+ * participants received the message at the beginning to gather all
+ * existing backends.
+ *
+ * - Some backends might be blocked and not responsing either before or
+ * after receiving the message. In the first case such backend still
+ * have ProcSignalSlot and should be waited for, in the second case
+ * shared barrier will make sure we still waiting for those backends. In
+ * any case there is an unbounded wait.
+ *
+ * - Backends might join barrier in disjoint groups with some time in
+ * between. That means that relying only on the shared dynamic barrier is
+ * not enough -- it will only synchronize resize procedure withing those
+ * groups. That's why we wait first for all participants of ProcSignal
+ * mechanism who received the message.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ pg_atomic_init_u64(&ShmemCtrl->Generation,
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
+
+ /* To order everything after setting Generation value */
+ pg_memory_barrier();
+
+ /*
+ * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
+ * the rest of the pack has started the procedure and it can resize shared
+ * memory as well.
+ *
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier. In fact even if we would like to
+ * wait here, it wouldn't be possible -- we're in the postmaster, without
+ * any waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1271,3 +1637,50 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier()
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ /* Nothing to do if resizing is not started */
+ if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
+ return;
+
+ BarrierAttach(barrier);
+
+ /* Otherwise wait through all available phases */
+ while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
+ {
+ ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
+ BarrierPhase(barrier))));
+
+ BarrierArriveAndWait(barrier, 0);
+ }
+
+ BarrierDetach(barrier);
+}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /*
+ * The barrier is missing here, it will be initialized right before
+ * starting the resizing process as a convenient way to reset it.
+ */
+
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+
+ /* shmem_resizable should be initialized by now */
+ ShmemCtrl->Resizable = shmem_resizable;
+ }
+}
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 3fe45de5da0..196f233fe0e 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -425,6 +425,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1693,6 +1694,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2038,6 +2042,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
@@ -3851,6 +3866,9 @@ process_pm_pmsignal(void)
request_state_update = true;
}
+ if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
+ AnonymousShmemResize();
+
/*
* Try to advance postmaster's state machine, if a child requests it.
*/
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index bd68b69ee98..ac844b114bd 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -24,7 +25,6 @@ ConditionVariableMinimallyPadded *BufferIOCVArray;
WritebackContext BackendWritebackContext;
CkptSortItem *CkptBufferIds;
-
/*
* Data Structures:
* buffers live in a freelist and a lookup data structure.
@@ -62,18 +62,28 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * postmaster, or in a standalone backend) or during shared-memory resize. Size
+ * of data structures initialized here depends on NBuffers, and to be able to
+ * change NBuffers without a restart we store each structure into a separate
+ * shared memory segment, which could be resized on demand.
+ *
+ * FirstBufferToInit tells where to start initializing buffers. For
+ * initialization it always will be zero, but when resizing shared-memory it
+ * indicates the number of already initialized buffers.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(void)
+BufferManagerShmemInit(int FirstBufferToInit)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
+ elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
+ FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -110,43 +120,44 @@ BufferManagerShmemInit(void)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory, or in all cases when resizing shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
- ClearBufferTag(&buf->tag);
+#ifndef EXEC_BACKEND
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = FirstBufferToInit; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ ClearBufferTag(&buf->tag);
- buf->buf_id = i;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- pgaio_wref_clear(&buf->io_wref);
+ buf->buf_id = i;
- /*
- * Initially link all the buffers together as unused. Subsequent
- * management of this list is done by freelist.c.
- */
- buf->freeNext = i + 1;
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
+#endif
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 9d00b80b4f8..abeb91e24fd 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -84,6 +84,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -151,6 +154,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, AioShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -298,7 +309,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit();
+ BufferManagerShmemInit(0);
/*
* Set up lock manager
@@ -310,6 +321,11 @@ CreateOrAttachShmemStructs(void)
*/
PredicateLockShmemInit();
+ /*
+ * Set up shared memory resize manager
+ */
+ ShmemControlInit();
+
/*
* Set up process table
*/
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 8e313ad9bf8..35c42f260a8 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -112,6 +113,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -175,6 +180,43 @@ ProcSignalInit(char *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -614,6 +656,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 389abc82519..0fd421f004e 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -493,17 +493,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
+ *
+ * XXX: There is an implicit assumption this can only happen in
+ * "resizable" segments, where only one shared structure is allowed.
+ * This has to be implemented more cleanly.
*/
if (result->size != size)
{
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
+ Size delta = size - result->size;
+
+ result->size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
}
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index b1fba850f02..58f1a05fd2a 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4311,6 +4312,15 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /* Verify the shared barrier, if it's still active: join and wait. */
+ WaitOnShmemBarrier();
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 8bce14c38fd..e0ba8384fdd 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
@@ -351,6 +353,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index dede37f7905..1e70853ccdb 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2354,14 +2354,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 0d8528b2875..405d0a7e65d 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int MaxAvailableMemory;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 1977001e533..52633dd7537 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -307,7 +307,7 @@ extern void LimitAdditionalLocalPins(uint32 *additional_pins);
extern bool EvictUnpinnedBuffer(Buffer buf);
/* in buf_init.c */
-extern void BufferManagerShmemInit(void);
+extern void BufferManagerShmemInit(int);
extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6ebda479ced..bb7ae4d33b3 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index a9681738146..558da6fdd55 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
+PG_LWLOCK(54, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index c5009a1cd73..2e47b222cbb 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,25 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+ pg_atomic_uint64 Generation;
+ bool Resizable;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -107,6 +127,12 @@ extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
void *ReserveAnonymousMemory(Size reserve_size);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 67fa9ac06e1..27bc6a81191 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,6 +42,7 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index defd8b66a19..522b8de1e02 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 1a30437ad96..6755b302858 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2738,6 +2738,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.45.1
[text/plain] v4-0007-Use-anonymous-files-to-back-shared-memory-segment.patch (10.7K, ../../eqs6v4rsboazl67xz3wxc6xjkgrpfybitpl45y3lmb2br67wbj@o7czebb3rlgd/8-v4-0007-Use-anonymous-files-to-back-shared-memory-segment.patch)
download | inline diff:
From 0e3c671082743f2826a7e8a96a19a071f5c8aeb3 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 15 Mar 2025 16:39:45 +0100
Subject: [PATCH v4 7/8] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f90cde00000-7f90d5126000 rw-s 00000000 00:01 5463 /memfd:main (deleted)
7f90d5126000-7f914de00000 ---p 00000000 00:00 0
7f914de00000-7f9175128000 rw-s 00000000 00:01 5466 /memfd:buffers (deleted)
7f9175128000-7f944de00000 ---p 00000000 00:00 0
7f944de00000-7f9455528000 rw-s 00000000 00:01 5469 /memfd:descriptors (deleted)
7f9455528000-7f94cde00000 ---p 00000000 00:00 0
7f94cde00000-7f94d5228000 rw-s 00000000 00:01 5472 /memfd:iocv (deleted)
7f94d5228000-7f954de00000 ---p 00000000 00:00 0
7f954de00000-7f9555266000 rw-s 00000000 00:01 5475 /memfd:checkpoint (deleted)
7f9555266000-7f958de00000 ---p 00000000 00:00 0
7f958de00000-7f95954aa000 rw-s 00000000 00:01 5478 /memfd:strategy (deleted)
7f95954aa000-7f95cde00000 ---p 00000000 00:00 0
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 73 +++++++++++++++++++++++++++++-----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/portability/mem.h | 2 +-
src/include/storage/pg_shmem.h | 3 +-
5 files changed, 68 insertions(+), 14 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a3437973784..87000a24eea 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -107,6 +107,7 @@ typedef struct AnonymousMapping
Pointer shmem; /* Pointer to the start of the mapped memory */
Pointer seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -127,7 +128,7 @@ static int next_free_segment = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -150,9 +151,9 @@ static int next_free_segment = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* 7f442e010000-7f443a800000 # reserved empty space
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* 7f444196c000-7f470a800000 # reserved empty space
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -643,13 +644,14 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* *hugepagesize and *mmap_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -708,6 +710,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -718,7 +721,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -727,6 +739,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -734,6 +748,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -771,7 +787,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
- int mmap_flags = PG_MMAP_FLAGS;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
#ifndef MAP_HUGETLB
/* ReserveAnonymousMemory should have dealt with this case */
@@ -785,7 +801,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
/* Round up the request size to a suitable large value */
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
@@ -794,6 +810,29 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
}
#endif
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment),
+ memfd_flags);
+
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(mapping->shmem_segment))));
+
elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p",
MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
@@ -807,7 +846,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
* a restart.
*/
ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | MAP_FIXED, -1, 0);
+ mmap_flags | MAP_FIXED, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
@@ -817,8 +856,15 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
"fallback to the non-resizable allocation",
MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+ /* Specify the segment file size using allocsize. */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(mapping->shmem_segment))));
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
else
@@ -889,7 +935,7 @@ ReserveAnonymousMemory(Size reserve_size)
Size hugepagesize, total_size = 0;
int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
/*
* Figure out how much memory is needed for all segments, keeping in
@@ -1070,6 +1116,13 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ if(ftruncate(m->segment_fd, new_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
/* Clean up some reserved space to resize into */
if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
ereport(FATAL,
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index ce719f1b412..ba972106de1 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index abeb91e24fd..dc2b4becf4a 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -396,7 +396,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 2e47b222cbb..b9573520d9a 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -124,7 +124,8 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size,
PGShmemHeader **shim, Pointer base);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
void *ReserveAnonymousMemory(Size reserve_size);
bool ProcessBarrierShmemResize(Barrier *barrier);
--
2.45.1
[text/plain] v4-0008-Support-resize-for-hugetlb.patch (4.3K, ../../eqs6v4rsboazl67xz3wxc6xjkgrpfybitpl45y3lmb2br67wbj@o7czebb3rlgd/9-v4-0008-Support-resize-for-hugetlb.patch)
download | inline diff:
From 08476af71724fcb3035fc907dc98a6ff351fe58e Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 5 Apr 2025 19:51:33 +0200
Subject: [PATCH v4 8/8] Support resize for hugetlb
Linux kernel has a set of limitations on remapping hugetlb segments: it
can't increase size of such segment [1], and shrinking it will not
release the memory back. In fact support for hugetlb mremap was
implemented no so long time ago [2].
As a workaround, avoid mremap for resizing shared memory. Instead unmap
the whole segment and map it back at the same address with the new size,
relying on the fact that fd for the anon file behind the segment is
still open and will keep the memory content.
[1]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/mremap.c?id=f4d2ef48250ad057e4f00087967b5ff366da9f39#n1593
[2]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/mm/mremap.c?id=550a7d60bd5e35a56942dba6d8a26752beb26c9f
---
src/backend/port/sysv_shmem.c | 60 +++++++++++++++++++++++++----------
1 file changed, 44 insertions(+), 16 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 87000a24eea..f0b53ce1d7c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1109,6 +1109,7 @@ AnonymousShmemResize(void)
/* Note that CalculateShmemSize indirectly depends on NBuffers */
Size new_size = CalculateShmemSize(&numSemas, i);
AnonymousMapping *m = &Mappings[i];
+ int mmap_flags = PG_MMAP_FLAGS;
if (m->shmem == NULL)
continue;
@@ -1116,6 +1117,44 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+#ifndef MAP_HUGETLB
+ /* ReserveAnonymousMemory should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ Size hugepagesize;
+
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the new size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ if (new_size % hugepagesize != 0)
+ new_size += hugepagesize - (new_size % hugepagesize);
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
+ }
+#endif
+
+ /*
+ * Linux limitations do not allow us to mremap hugetlb in the way we
+ * want. E.g. no size increase is allowed, and for shrinking the memory
+ * will not be released back. To work around this unmap the segment and
+ * create a new one at the same address. Thanks for the backing anon
+ * file the content will still be kept in memory.
+ */
+ elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ if (munmap(m->shmem, m->shmem_size) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not unmap shared memory segment %s [%p]: %m",
+ MappingName(m->shmem_segment), m->shmem)));
+
/* Resize the backing anon file. */
if(ftruncate(m->segment_fd, new_size) == -1)
ereport(FATAL,
@@ -1123,25 +1162,14 @@ AnonymousShmemResize(void)
errmsg("could not truncase anonymous file for \"%s\": %m",
MappingName(m->shmem_segment))));
- /* Clean up some reserved space to resize into */
- if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
- ereport(FATAL,
- (errcode(ERRCODE_SYSTEM_ERROR),
- errmsg("could not unmap %zu from reserved shared memory %p: %m",
- new_size - m->shmem_size, m->shmem)));
-
- /* Claim the unused space */
- elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
- MappingName(m->shmem_segment), m->shmem_size,
- new_size, m->shmem);
-
- ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
+ /* Reclaim the space */
+ ptr = mmap(m->shmem, new_size, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_FIXED, m->segment_fd, 0);
if (ptr == MAP_FAILED)
ereport(FATAL,
(errcode(ERRCODE_SYSTEM_ERROR),
- errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m",
- MappingName(m->shmem_segment), m->shmem, NBuffers,
- new_size)));
+ errmsg("could not map shared memory segment %s [%p] with size %zu: %m",
+ MappingName(m->shmem_segment), m->shmem, new_size)));
reinit = true;
m->shmem_size = new_size;
--
2.45.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-09 05:42 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-09 07:45 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-04-09 05:42 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Mon, Apr 7, 2025 at 2:13 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> In the new v4 version
> of the patch the first option is implemented.
>
The patches don't apply cleanly using git am but patch -p1 applies
them cleanly. However I see following compilation errors
RuntimeError: command "ninja" failed with error [1/1954] Generating
src/include/utils/errcodes with a custom command
[2/1954] Generating src/include/storage/lwlocknames_h with a custom command
[3/1954] Generating src/include/utils/wait_event_names with a custom command
[4/1954] Compiling C object src/port/libpgport.a.p/pg_popcount_aarch64.c.o
[5/1954] Compiling C object src/port/libpgport.a.p/pg_numa.c.o
FAILED: src/port/libpgport.a.p/pg_numa.c.o
cc -Isrc/port/libpgport.a.p -Isrc/include
-I../../coderoot/pg/src/include -fdiagnostics-color=always
-D_FILE_OFFSET_BITS=64 -Wall -Winvalid-pch -Werror -g
-fno-strict-aliasing -fwrapv -fexcess-precision=standard -D_GNU_SOURCE
-Wmissing-prototypes -Wpointer-arith -Werror=vla -Wendif-labels
-Wmissing-format-attribute -Wimplicit-fallthrough=3
-Wcast-function-type -Wshadow=compatible-local -Wformat-security
-Wdeclaration-after-statement -Wno-format-truncation
-Wno-stringop-truncation -fPIC -DFRONTEND -MD -MQ
src/port/libpgport.a.p/pg_numa.c.o -MF
src/port/libpgport.a.p/pg_numa.c.o.d -o
src/port/libpgport.a.p/pg_numa.c.o -c
../../coderoot/pg/src/port/pg_numa.c
In file included from ../../coderoot/pg/src/include/storage/spin.h:54,
from
../../coderoot/pg/src/include/storage/condition_variable.h:26,
from ../../coderoot/pg/src/include/storage/barrier.h:22,
from ../../coderoot/pg/src/include/storage/pg_shmem.h:27,
from ../../coderoot/pg/src/port/pg_numa.c:26:
../../coderoot/pg/src/include/storage/s_lock.h:93:2: error: #error
"s_lock.h may not be included from frontend code"
93 | #error "s_lock.h may not be included from frontend code"
| ^~~~~
In file included from ../../coderoot/pg/src/port/pg_numa.c:26:
../../coderoot/pg/src/include/storage/pg_shmem.h:66:9: error: unknown
type name ‘pg_atomic_uint32’
66 | pg_atomic_uint32 NSharedBuffers;
| ^~~~~~~~~~~~~~~~
../../coderoot/pg/src/include/storage/pg_shmem.h:68:9: error: unknown
type name ‘pg_atomic_uint64’
68 | pg_atomic_uint64 Generation;
| ^~~~~~~~~~~~~~~~
../../coderoot/pg/src/port/pg_numa.c: In function ‘pg_numa_get_pagesize’:
../../coderoot/pg/src/port/pg_numa.c:117:17: error: too few arguments
to function ‘GetHugePageSize’
117 | GetHugePageSize(&os_page_size, NULL);
| ^~~~~~~~~~~~~~~
In file included from ../../coderoot/pg/src/port/pg_numa.c:26:
../../coderoot/pg/src/include/storage/pg_shmem.h:127:13: note: declared here
127 | extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
| ^~~~~~~~~~~~~~~
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-09 05:42 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-04-09 07:45 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-09 07:50 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-09 07:45 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Wed, Apr 09, 2025 at 11:12:18AM GMT, Ashutosh Bapat wrote:
> On Mon, Apr 7, 2025 at 2:13 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> >
> > In the new v4 version
> > of the patch the first option is implemented.
> >
>
> The patches don't apply cleanly using git am but patch -p1 applies
> them cleanly. However I see following compilation errors
> RuntimeError: command "ninja" failed with error
Becase it's relatively meaningless to apply a patch to the tip of the
master around the release freeze time :) Commit 65c298f61fc has
introduced new usage of GetHugePageSize, which was modified in my patch.
I'm going to address it with the next rebased version, in the meantime
you can always use the specified base commit to apply the changeset:
base-commit: 5e1915439085014140314979c4dd5e23bd677cac
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-09 05:42 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-09 07:45 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-09 07:50 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-09 08:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-04-09 07:50 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Wed, Apr 9, 2025 at 1:15 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Wed, Apr 09, 2025 at 11:12:18AM GMT, Ashutosh Bapat wrote:
> > On Mon, Apr 7, 2025 at 2:13 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > >
> > > In the new v4 version
> > > of the patch the first option is implemented.
> > >
> >
> > The patches don't apply cleanly using git am but patch -p1 applies
> > them cleanly. However I see following compilation errors
> > RuntimeError: command "ninja" failed with error
>
> Becase it's relatively meaningless to apply a patch to the tip of the
> master around the release freeze time :) Commit 65c298f61fc has
> introduced new usage of GetHugePageSize, which was modified in my patch.
> I'm going to address it with the next rebased version, in the meantime
> you can always use the specified base commit to apply the changeset:
>
> base-commit: 5e1915439085014140314979c4dd5e23bd677cac
There is a higher chance that people will try these patches now than
it was two days before and more chance if they find the patches
applicable easily.
../../coderoot/pg/src/include/storage/s_lock.h:93:2: error: #error
"s_lock.h may not be included from frontend code"
How about this? Why is that happening?
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-09 05:42 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-09 07:45 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-09 07:50 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-04-09 08:19 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-09 08:19 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Wed, Apr 09, 2025 at 01:20:16PM GMT, Ashutosh Bapat wrote:
> ../../coderoot/pg/src/include/storage/s_lock.h:93:2: error: #error
> "s_lock.h may not be included from frontend code"
>
> How about this? Why is that happening?
The same -- as you can see it comes from compiling pg_numa.c, which as
it seems used in frontend and imports pg_shmem.h . I wanted to reshuffle
includes in the patch anyway, that would be a good excuse to finally do
this.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-11 14:34 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-04-11 14:34 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Mon, Apr 7, 2025 at 2:13 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> Yes, you're right, plain dynamic Barrier does not ensure all available
> processes will be synchronized. I was aware about the scenario you
> describe, it's mentioned in commentaries for the resize function. I was
> under the impression this should be enough, but after some more thinking
> I'm not so sure anymore. Let me try to structure it as a list of
> possible corner cases that we need to worry about:
>
> * New backend spawned while we're busy resizing shared memory. Those
> should wait until the resizing is complete and get the new size as well.
>
> * Old backend receives a resize message, but exits before attempting to
> resize. Those should be excluded from coordination.
Should we detach barrier in on_exit()?
>
> * A backend is blocked and not responding before or after the
> ProcSignalBarrier message was sent. I'm thinking about a failure
> situation, when one rogue backend is doing something without checking
> for interrupts. We need to wait for those to become responsive, and
> potentially abort shared memory resize after some timeout.
Right.
>
> I think a relatively elegant solution is to extend ProcSignalBarrier
> mechanism to track not only pss_barrierGeneration, as a sign that
> everything was processed, but also something like
> pss_barrierReceivedGeneration, indicating that the message was received
> everywhere but not processed yet. That would be enough to allow
> processes to wait until the resize message was received everywhere, then
> use a global Barrier to wait until all processes are finished. It's
> somehow similar to your proposal to use two signals, but has less
> implementation overhead.
The way it's implemented in v4 still has the disjoint group problem.
Assume backends p1, p2, p3. All three of them are executing
ProcessProcSignalBarrier(). All three of them updated
pss_barrierReceivedGeneration
/* The message is observed, record that */
pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
shared_gen);
p1, p2 moved faster and reached following code from ProcessBarrierShmemResize()
if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
WaitForProcSignalBarrierReceived(pg_atomic_read_u64(&ShmemCtrl->Generation));
Since all the processes have received the barrier message, p1, p2 move
ahead and go through all the next phases and finish resizing even
before p3 gets a chance to call ProcessBarrierShmemResize() and attach
itself to Barrier. This could happen because it processed some other
ProcSignalBarrier message. p1 and p2 won't wait for p3 since it has
not attached itself to the barrier. Once p1, p2 finish, p3 will attach
itself to the barrier and resize buffers again - reinitializing the
shared memory, which might has been already modified by p1 or p2. Boom
- there's memory corruption.
Either every process has to make sure that all the other extant
backends have attached themselves to the barrier OR somebody has to
ensure that and signal all the backends to proceed. The implementation
doesn't do either.
>
> * Shared memory address space is now reserved for future usage, making
> shared memory segments clash (e.g. due to memory allocation)
> impossible. There is a new GUC to control how much space to reserve,
> which is called max_available_memory -- on the assumption that most of
> the time it would make sense to set its value to the total amount of
> memory on the machine. I'm open for suggestions regarding the name.
With 0006 applied
+ /* Clean up some reserved space to resize into */
+ if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
ze, m->shmem)));
... snip ...
+ ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
We unmap the portion of reserved address space where the existing
segment would expand into. As long as we are just expanding this will
work. I am wondering how would this work for shrinking buffers? What
scheme do you have in mind?
>
> * There is one more patch to address hugepages remap. As mentioned in
> this thread above, Linux kernel has certain limitations when it comes
> to mremap for segments allocated with huge pages. To work around it's
> possible to replace mremap with a sequence of unmap and map again,
> relying on the anon file behind the segment to keep the memory
> content. I haven't found any downsides of this approach so far, but it
> makes the anonymous file patch 0007 mandatory.
In 0008
if (munmap(m->shmem, m->shmem_size) < 0)
... snip ...
/* Resize the backing anon file. */
if(ftruncate(m->segment_fd, new_size) == -1)
...
/* Reclaim the space */
ptr = mmap(m->shmem, new_size, PROT_READ | PROT_WRITE,
mmap_flags | MAP_FIXED, m->segment_fd, 0);
How are we preventing something get mapped into the space after
m->shmem + newsize? We will need to add an unallocated but reserved
addressed space map after m->shmem+newsize right?
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-04-11 15:01 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-11 15:01 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Fri, Apr 11, 2025 at 08:04:39PM GMT, Ashutosh Bapat wrote:
> On Mon, Apr 7, 2025 at 2:13 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> >
> > Yes, you're right, plain dynamic Barrier does not ensure all available
> > processes will be synchronized. I was aware about the scenario you
> > describe, it's mentioned in commentaries for the resize function. I was
> > under the impression this should be enough, but after some more thinking
> > I'm not so sure anymore. Let me try to structure it as a list of
> > possible corner cases that we need to worry about:
> >
> > * New backend spawned while we're busy resizing shared memory. Those
> > should wait until the resizing is complete and get the new size as well.
> >
> > * Old backend receives a resize message, but exits before attempting to
> > resize. Those should be excluded from coordination.
>
> Should we detach barrier in on_exit()?
Yeah, good point.
> > I think a relatively elegant solution is to extend ProcSignalBarrier
> > mechanism to track not only pss_barrierGeneration, as a sign that
> > everything was processed, but also something like
> > pss_barrierReceivedGeneration, indicating that the message was received
> > everywhere but not processed yet. That would be enough to allow
> > processes to wait until the resize message was received everywhere, then
> > use a global Barrier to wait until all processes are finished. It's
> > somehow similar to your proposal to use two signals, but has less
> > implementation overhead.
>
> The way it's implemented in v4 still has the disjoint group problem.
> Assume backends p1, p2, p3. All three of them are executing
> ProcessProcSignalBarrier(). All three of them updated
> pss_barrierReceivedGeneration
>
> /* The message is observed, record that */
> pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
> shared_gen);
>
> p1, p2 moved faster and reached following code from ProcessBarrierShmemResize()
> if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
> WaitForProcSignalBarrierReceived(pg_atomic_read_u64(&ShmemCtrl->Generation));
>
> Since all the processes have received the barrier message, p1, p2 move
> ahead and go through all the next phases and finish resizing even
> before p3 gets a chance to call ProcessBarrierShmemResize() and attach
> itself to Barrier. This could happen because it processed some other
> ProcSignalBarrier message. p1 and p2 won't wait for p3 since it has
> not attached itself to the barrier. Once p1, p2 finish, p3 will attach
> itself to the barrier and resize buffers again - reinitializing the
> shared memory, which might has been already modified by p1 or p2. Boom
> - there's memory corruption.
It won't reinitialize anything, since this logic is controlled by the
ShmemCtrl->NSharedBuffers, if it's already updated nothing will be
changed.
About the race condition you mention, there is indeed a window between
receiving the ProcSignalBarrier and attaching to the global Barrier in
resize, but I don't think any process will be able to touch buffer pool
while inside this window. Even if it happens that the remapping itself
was blazing fast that this window was enough to make one process late
(e.g. if it was busy handling some other signal as you mention), as I've
showed above it shouldn't be a problem.
I can experiment with this case though, maybe there is a way to
completely close this window to not thing about even potential
scenarios.
> > * Shared memory address space is now reserved for future usage, making
> > shared memory segments clash (e.g. due to memory allocation)
> > impossible. There is a new GUC to control how much space to reserve,
> > which is called max_available_memory -- on the assumption that most of
> > the time it would make sense to set its value to the total amount of
> > memory on the machine. I'm open for suggestions regarding the name.
>
> With 0006 applied
> + /* Clean up some reserved space to resize into */
> + if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
> ze, m->shmem)));
> ... snip ...
> + ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
>
> We unmap the portion of reserved address space where the existing
> segment would expand into. As long as we are just expanding this will
> work. I am wondering how would this work for shrinking buffers? What
> scheme do you have in mind?
I didn't like this part originally, and after changes to support hugetlb
I think it's worth it just to replace mremap with munmap/mmap. That way
there will be no such question, e.g. if a segment is getting shrinked
the unmapped area will again become a part of the reserved space.
> > * There is one more patch to address hugepages remap. As mentioned in
> > this thread above, Linux kernel has certain limitations when it comes
> > to mremap for segments allocated with huge pages. To work around it's
> > possible to replace mremap with a sequence of unmap and map again,
> > relying on the anon file behind the segment to keep the memory
> > content. I haven't found any downsides of this approach so far, but it
> > makes the anonymous file patch 0007 mandatory.
>
> In 0008
> if (munmap(m->shmem, m->shmem_size) < 0)
> ... snip ...
> /* Resize the backing anon file. */
> if(ftruncate(m->segment_fd, new_size) == -1)
> ...
> /* Reclaim the space */
> ptr = mmap(m->shmem, new_size, PROT_READ | PROT_WRITE,
> mmap_flags | MAP_FIXED, m->segment_fd, 0);
>
> How are we preventing something get mapped into the space after
> m->shmem + newsize? We will need to add an unallocated but reserved
> addressed space map after m->shmem+newsize right?
Nope, the segment is allocated from the reserved space already, with
some chunk of it left after the segment's end for resizing purposes. We
only take some part of the designated space, the rest is still reserved.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-14 05:10 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-04-14 05:10 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Fri, Apr 11, 2025 at 8:31 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > > I think a relatively elegant solution is to extend ProcSignalBarrier
> > > mechanism to track not only pss_barrierGeneration, as a sign that
> > > everything was processed, but also something like
> > > pss_barrierReceivedGeneration, indicating that the message was received
> > > everywhere but not processed yet. That would be enough to allow
> > > processes to wait until the resize message was received everywhere, then
> > > use a global Barrier to wait until all processes are finished. It's
> > > somehow similar to your proposal to use two signals, but has less
> > > implementation overhead.
> >
> > The way it's implemented in v4 still has the disjoint group problem.
> > Assume backends p1, p2, p3. All three of them are executing
> > ProcessProcSignalBarrier(). All three of them updated
> > pss_barrierReceivedGeneration
> >
> > /* The message is observed, record that */
> > pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
> > shared_gen);
> >
> > p1, p2 moved faster and reached following code from ProcessBarrierShmemResize()
> > if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
> > WaitForProcSignalBarrierReceived(pg_atomic_read_u64(&ShmemCtrl->Generation));
> >
> > Since all the processes have received the barrier message, p1, p2 move
> > ahead and go through all the next phases and finish resizing even
> > before p3 gets a chance to call ProcessBarrierShmemResize() and attach
> > itself to Barrier. This could happen because it processed some other
> > ProcSignalBarrier message. p1 and p2 won't wait for p3 since it has
> > not attached itself to the barrier. Once p1, p2 finish, p3 will attach
> > itself to the barrier and resize buffers again - reinitializing the
> > shared memory, which might has been already modified by p1 or p2. Boom
> > - there's memory corruption.
>
> It won't reinitialize anything, since this logic is controlled by the
> ShmemCtrl->NSharedBuffers, if it's already updated nothing will be
> changed.
Ah, I see it now
if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
{
Thanks for the clarification.
However, when we put back the patches to shrink buffers, we will evict
the extra buffers, and shrink - if all the processes haven't
participated in the barrier by then, some of them may try to access
those buffers - re-installing them and then bad things can happen.
>
> About the race condition you mention, there is indeed a window between
> receiving the ProcSignalBarrier and attaching to the global Barrier in
> resize, but I don't think any process will be able to touch buffer pool
> while inside this window. Even if it happens that the remapping itself
> was blazing fast that this window was enough to make one process late
> (e.g. if it was busy handling some other signal as you mention), as I've
> showed above it shouldn't be a problem.
>
> I can experiment with this case though, maybe there is a way to
> completely close this window to not thing about even potential
> scenarios.
The window may be small today but we have to make this future proof.
Multiple ProcSignalBarrier messages may be processed in a single call
to ProcessProcSignalBarrier() and if each of those takes as long as
buffer resizing, the window will get bigger and bigger. So we have to
close this window.
>
> > > * Shared memory address space is now reserved for future usage, making
> > > shared memory segments clash (e.g. due to memory allocation)
> > > impossible. There is a new GUC to control how much space to reserve,
> > > which is called max_available_memory -- on the assumption that most of
> > > the time it would make sense to set its value to the total amount of
> > > memory on the machine. I'm open for suggestions regarding the name.
> >
> > With 0006 applied
> > + /* Clean up some reserved space to resize into */
> > + if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
> > ze, m->shmem)));
> > ... snip ...
> > + ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
> >
> > We unmap the portion of reserved address space where the existing
> > segment would expand into. As long as we are just expanding this will
> > work. I am wondering how would this work for shrinking buffers? What
> > scheme do you have in mind?
>
> I didn't like this part originally, and after changes to support hugetlb
> I think it's worth it just to replace mremap with munmap/mmap. That way
> there will be no such question, e.g. if a segment is getting shrinked
> the unmapped area will again become a part of the reserved space.
>
I might have not noticed it, but are we putting two mappings one
reserved and one allocated in the same address space, so that when the
allocated mapping shrinks or expands, the reserved mapping continues
to prohibit any other mapping from appearing there? I looked at some
of the previous emails, but didn't find anything that describes how
the reserved mapped space is managed.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-04-14 07:20 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 08:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-14 07:20 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Mon, Apr 14, 2025 at 10:40:28AM GMT, Ashutosh Bapat wrote:
>
> However, when we put back the patches to shrink buffers, we will evict
> the extra buffers, and shrink - if all the processes haven't
> participated in the barrier by then, some of them may try to access
> those buffers - re-installing them and then bad things can happen.
As I've mentioned above, I don't see how a process could try to access a
buffer, if it's on the path between receiving the ProcSignalBarrier and
attaching to the global shmem Barrier, even if we shrink buffers.
AFAICT interrupt handles should not touch buffers, and otherwise the
process doesn't have any point withing this window where it might do
this. Do you have some particular scenario in mind?
> I might have not noticed it, but are we putting two mappings one
> reserved and one allocated in the same address space, so that when the
> allocated mapping shrinks or expands, the reserved mapping continues
> to prohibit any other mapping from appearing there? I looked at some
> of the previous emails, but didn't find anything that describes how
> the reserved mapped space is managed.
I though so, but this turns out to be incorrect. Just have done a small
experiment -- looks like when reserving some space, mapping and
unmapping a small segment from it leaves a non-mapped gap. That would
mean for shrinking the new available space has to be reserved again.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-14 08:58 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-04-14 08:58 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Mon, Apr 14, 2025 at 12:50 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Mon, Apr 14, 2025 at 10:40:28AM GMT, Ashutosh Bapat wrote:
> >
> > However, when we put back the patches to shrink buffers, we will evict
> > the extra buffers, and shrink - if all the processes haven't
> > participated in the barrier by then, some of them may try to access
> > those buffers - re-installing them and then bad things can happen.
>
> As I've mentioned above, I don't see how a process could try to access a
> buffer, if it's on the path between receiving the ProcSignalBarrier and
> attaching to the global shmem Barrier, even if we shrink buffers.
> AFAICT interrupt handles should not touch buffers, and otherwise the
> process doesn't have any point withing this window where it might do
> this. Do you have some particular scenario in mind?
ProcessProcSignalBarrier() is not within an interrupt handler but it
responds to a flag set by an interrupt handler. After calling
pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
shared_gen); it will enter the loop
while (flags != 0)
where it may process many barriers before processing
PROCSIGNAL_BARRIER_SHMEM_RESIZE. Nothing stops the other barrier
processing code from touching buffers. Right now it's just smgrrelease
that gets called in the other barrier. But that's not guaranteed in
future.
>
> > I might have not noticed it, but are we putting two mappings one
> > reserved and one allocated in the same address space, so that when the
> > allocated mapping shrinks or expands, the reserved mapping continues
> > to prohibit any other mapping from appearing there? I looked at some
> > of the previous emails, but didn't find anything that describes how
> > the reserved mapped space is managed.
>
> I though so, but this turns out to be incorrect. Just have done a small
> experiment -- looks like when reserving some space, mapping and
> unmapping a small segment from it leaves a non-mapped gap. That would
> mean for shrinking the new available space has to be reserved again.
Right. That's what I thought. But I didn't see the corresponding code.
So we have to keep track of two mappings for every segment - 1 for
allocation and one for reserving space and resize those two while
shrinking and expanding buffers. Am I correct?
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-17 09:52 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 23:05 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
1 sibling, 2 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-04-17 09:52 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Hi Dmitry,
On Mon, Apr 14, 2025 at 12:50 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Mon, Apr 14, 2025 at 10:40:28AM GMT, Ashutosh Bapat wrote:
> >
> > However, when we put back the patches to shrink buffers, we will evict
> > the extra buffers, and shrink - if all the processes haven't
> > participated in the barrier by then, some of them may try to access
> > those buffers - re-installing them and then bad things can happen.
>
> As I've mentioned above, I don't see how a process could try to access a
> buffer, if it's on the path between receiving the ProcSignalBarrier and
> attaching to the global shmem Barrier, even if we shrink buffers.
> AFAICT interrupt handles should not touch buffers, and otherwise the
> process doesn't have any point withing this window where it might do
> this. Do you have some particular scenario in mind?
>
> > I might have not noticed it, but are we putting two mappings one
> > reserved and one allocated in the same address space, so that when the
> > allocated mapping shrinks or expands, the reserved mapping continues
> > to prohibit any other mapping from appearing there? I looked at some
> > of the previous emails, but didn't find anything that describes how
> > the reserved mapped space is managed.
>
> I though so, but this turns out to be incorrect. Just have done a small
> experiment -- looks like when reserving some space, mapping and
> unmapping a small segment from it leaves a non-mapped gap. That would
> mean for shrinking the new available space has to be reserved again.
In an offlist chat Thomas Munro mentioned that just ftruncate() would
be enough to resize the shared memory without touching address maps
using mmap and munmap().
ftruncate man page seems to concur with him
If the effect of ftruncate() is to decrease the size of a memory
mapped file or a shared memory object and whole pages beyond the
new end were previously mapped, then the whole pages beyond the
new end shall be discarded.
References to discarded pages shall result in the generation of a
SIGBUS signal.
If the effect of ftruncate() is to increase the size of a memory
object, it is unspecified whether the contents of any mapped pages
between the old end-of-file and the new are flushed to the
underlying object.
ftruncate() when shrinking memory will release the extra pages and
also would cause segmentation fault when memory outside the size of
file is accessed even if the actual address map is larger than the
mapped file. The expanded memory is allocated as it is written to, and
those pages also become visible in the underlying object.
I played with the attached small program under debugger observing pmap
and /proc/<pid>/status after every memory operation. The address map
always shows that it's as long as 300K memory.
00007fffd2200000 307200K rw-s- memfd:mmap_fd_exp (deleted)
Immediately after mmap()
RssShmem: 0 kB
after first memset
RssShmem: 307200 kB
after ftruncate to 100MB (we don't need to wait for memset() to see
the effect on RssShmem)
RssShmem: 102400 kB
after ftruncate to 200MB (requires memset to see effect on RssShmem)
RssShmem: 102400 kB
after memsetting upto 200MB
RssShmem: 204800 kB
All the observations concur with the man page.
[1] https://man7.org/linux/man-pages/man3/ftruncate.3p.html#:~:text=If%20the%20effect%20of%20ftruncate,g....
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-04-17 21:16 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-17 21:16 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Thu, Apr 17, 2025 at 03:22:28PM GMT, Ashutosh Bapat wrote:
>
> In an offlist chat Thomas Munro mentioned that just ftruncate() would
> be enough to resize the shared memory without touching address maps
> using mmap and munmap().
>
> ftruncate man page seems to concur with him
>
> If the effect of ftruncate() is to decrease the size of a memory
> mapped file or a shared memory object and whole pages beyond the
> new end were previously mapped, then the whole pages beyond the
> new end shall be discarded.
>
> References to discarded pages shall result in the generation of a
> SIGBUS signal.
>
> If the effect of ftruncate() is to increase the size of a memory
> object, it is unspecified whether the contents of any mapped pages
> between the old end-of-file and the new are flushed to the
> underlying object.
>
> ftruncate() when shrinking memory will release the extra pages and
> also would cause segmentation fault when memory outside the size of
> file is accessed even if the actual address map is larger than the
> mapped file. The expanded memory is allocated as it is written to, and
> those pages also become visible in the underlying object.
Thanks for sharing. I need to do more thorough tests, but after a quick
look I'm not sure about that. ftruncate will take care about the memory,
but AFAICT the memory mapping will stay the same, is that what you mean?
In that case if the segment got increased, the memory still can't be
used because it's beyond the mapping end (at least in my test that's
what happened). If the segment got shrinked, the memory couldn't be
reclaimed, because, well, there is already a mapping. Or do I miss
something?
> > > I might have not noticed it, but are we putting two mappings one
> > > reserved and one allocated in the same address space, so that when the
> > > allocated mapping shrinks or expands, the reserved mapping continues
> > > to prohibit any other mapping from appearing there? I looked at some
> > > of the previous emails, but didn't find anything that describes how
> > > the reserved mapped space is managed.
> >
> > I though so, but this turns out to be incorrect. Just have done a small
> > experiment -- looks like when reserving some space, mapping and
> > unmapping a small segment from it leaves a non-mapped gap. That would
> > mean for shrinking the new available space has to be reserved again.
>
> Right. That's what I thought. But I didn't see the corresponding code.
> So we have to keep track of two mappings for every segment - 1 for
> allocation and one for reserving space and resize those two while
> shrinking and expanding buffers. Am I correct?
Not necessarily, depending on what we want. Again, I'll do a bit more testing,
but after a quick check it seems that it's possible to "plug" the gap with a
new reservation mapping, then reallocate it to another mapping or unmap both
reservations (main and the "gap" one) at once. That would mean that for the
current functionality we don't need to track reservation in any way more than
just start and the end of the "main" reserved space. The only consequence I can
imagine is possible fragmentation of the reserved space in case of frequent
increase/decrease of a segment with even decreasing size. But since it's only
reserved space, which will not really be used, it's probably not going to be a
problem.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-18 09:17 ` Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:02 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Thomas Munro @ 2025-04-18 09:17 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Fri, Apr 18, 2025 at 7:25 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > On Thu, Apr 17, 2025 at 03:22:28PM GMT, Ashutosh Bapat wrote:
> >
> > In an offlist chat Thomas Munro mentioned that just ftruncate() would
> > be enough to resize the shared memory without touching address maps
> > using mmap and munmap().
> >
> > ftruncate man page seems to concur with him
> >
> > If the effect of ftruncate() is to decrease the size of a memory
> > mapped file or a shared memory object and whole pages beyond the
> > new end were previously mapped, then the whole pages beyond the
> > new end shall be discarded.
> >
> > References to discarded pages shall result in the generation of a
> > SIGBUS signal.
> >
> > If the effect of ftruncate() is to increase the size of a memory
> > object, it is unspecified whether the contents of any mapped pages
> > between the old end-of-file and the new are flushed to the
> > underlying object.
> >
> > ftruncate() when shrinking memory will release the extra pages and
> > also would cause segmentation fault when memory outside the size of
> > file is accessed even if the actual address map is larger than the
> > mapped file. The expanded memory is allocated as it is written to, and
> > those pages also become visible in the underlying object.
>
> Thanks for sharing. I need to do more thorough tests, but after a quick
> look I'm not sure about that. ftruncate will take care about the memory,
> but AFAICT the memory mapping will stay the same, is that what you mean?
> In that case if the segment got increased, the memory still can't be
> used because it's beyond the mapping end (at least in my test that's
> what happened). If the segment got shrinked, the memory couldn't be
> reclaimed, because, well, there is already a mapping. Or do I miss
> something?
I was imagining that you might map some maximum possible size at the
beginning to reserve the address space permanently, and then adjust
the virtual memory object's size with ftruncate as required to provide
backing. Doesn't that achieve the goal with fewer steps, using only
portable* POSIX stuff, and keeping all pointers stable? I understand
that pointer stability may not be required (I can see roughly how that
argument is constructed), but isn't it still better to avoid having to
prove that and deal with various other problems completely? Is there
a downside/cost to having a large mapping that is only partially
backed? I suppose choosing that number might offend you but at least
there is an obvious upper bound: physical memory size.
*You might also want to use fallocate after ftruncate on Linux to
avoid SIGBUS on allocation failure on first touch page fault, which
raises portability questions since it's unspecified whether you can do
that with shm fds and fails on some systems, but it let's call that an
independent topic as it's not affected by this choice.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
@ 2025-04-18 11:02 ` Andres Freund <andres@anarazel.de>
1 sibling, 0 replies; 167+ messages in thread
From: Andres Freund @ 2025-04-18 11:02 UTC (permalink / raw)
To: pgsql-hackers@lists.postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Hi,
On April 18, 2025 11:17:21 AM GMT+02:00, Thomas Munro <thomas.munro@gmail.com> wrote:
> Doesn't that achieve the goal with fewer steps, using only
>portable* POSIX stuff, and keeping all pointers stable? I understand
>that pointer stability may not be required (I can see roughly how that
>argument is constructed), but isn't it still better to avoid having to
>prove that and deal with various other problems completely?
I think we should flat out reject any approach that does not maintain pointer stability. It would restrict future optimizations a lot if we can't rely on that (e.g. not materializing tuples when transporting them from worker to leader; pointering datastructures in shared buffers).
Greetings,
Andres
--
Sent from my Android device with K-9 Mail. Please excuse my brevity.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
@ 2025-04-18 11:05 ` Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Thomas Munro @ 2025-04-18 11:05 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Fri, Apr 18, 2025 at 9:17 PM Thomas Munro <thomas.munro@gmail.com> wrote:
> On Fri, Apr 18, 2025 at 7:25 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > Thanks for sharing. I need to do more thorough tests, but after a quick
> > look I'm not sure about that. ftruncate will take care about the memory,
> > but AFAICT the memory mapping will stay the same, is that what you mean?
> > In that case if the segment got increased, the memory still can't be
> > used because it's beyond the mapping end (at least in my test that's
> > what happened). If the segment got shrinked, the memory couldn't be
> > reclaimed, because, well, there is already a mapping. Or do I miss
> > something?
>
> I was imagining that you might map some maximum possible size at the
> beginning to reserve the address space permanently, and then adjust
> the virtual memory object's size with ftruncate as required to provide
> backing. Doesn't that achieve the goal with fewer steps, using only
> portable* POSIX stuff, and keeping all pointers stable? I understand
> that pointer stability may not be required (I can see roughly how that
> argument is constructed), but isn't it still better to avoid having to
> prove that and deal with various other problems completely? Is there
> a downside/cost to having a large mapping that is only partially
> backed? I suppose choosing that number might offend you but at least
> there is an obvious upper bound: physical memory size.
TIL that mmap(size, fd) will actually extend a hugetlb memfd as a side
effect on Linux, as if you had called ftruncate on it (fully allocated
huge pages I expected up to the object's size, just not magical size
changes beyond that when I merely asked to map it). That doesn't
happen for regular page size, or for any page size on my local OS's
shm objects and doesn't seem to fit mmap's job description given an
fd*, but maybe I'm just confused. Anyway, a workaround seems to be
to start out with PROT_NONE and MAP_NORESERVE, then mprotect(PROT_READ
| PROT_WRITE) new regions after extending with ftruncate(), at least
in simple tests...
(*Hmm, wiild uninformed speculation: perhap the size-setting behaviour
needed when hugetlbfs is used secretly to implement MAP_ANONYMOUS is
being exposed also when a hugetlbfs fd is given explicitly to mmap,
generating this bizarro side effect?)
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
@ 2025-04-21 09:29 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-21 09:29 UTC (permalink / raw)
To: Thomas Munro <thomas.munro@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Fri, Apr 18, 2025 at 09:17:21PM GMT, Thomas Munro wrote:
> I was imagining that you might map some maximum possible size at the
> beginning to reserve the address space permanently, and then adjust
> the virtual memory object's size with ftruncate as required to provide
> backing. Doesn't that achieve the goal with fewer steps, using only
> portable* POSIX stuff, and keeping all pointers stable?
Ah, I see what you folks mean. So in the latest patch there is a single large
shared memory area reserved with PROT_NONE + MAP_NORESERVE. This area is
logically divided between shmem segments, and each segment is mmap'd out of it
and could be resized withing these logical boundaries. Now the suggestion is to
have one reserved area for each segment, and instead of really mmap'ing
something out of it, manage memory via ftruncate.
Yeah, that would work and will allow to avoid MAP_FIXED and mremap, which are
questionable from portability point of view. This leaves memfd_create, and I'm
still not completely clear on it's portability -- it seems to be specific to
Linux, but others provide compatible implementation as well.
Let me experiment with this idea a bit, I would like to make sure there are no
other limitations we might face.
> I understand that pointer stability may not be required
Just to clarify, the current patch maintains this property (stable pointers),
which I also see as mandatory for any possible implementation.
> *You might also want to use fallocate after ftruncate on Linux to
> avoid SIGBUS on allocation failure on first touch page fault, which
> raises portability questions since it's unspecified whether you can do
> that with shm fds and fails on some systems, but it let's call that an
> independent topic as it's not affected by this choice.
I'm afraid it would be strictly neccessary to do fallocate, otherwise we're
back where we were before reservation accounting for huge pages in Linux (lot's
of people were facing unexpected SIGBUS when dealing with cgroups).
> TIL that mmap(size, fd) will actually extend a hugetlb memfd as a side
> effect on Linux, as if you had called ftruncate on it (fully allocated
> huge pages I expected up to the object's size, just not magical size
> changes beyond that when I merely asked to map it). That doesn't
> happen for regular page size, or for any page size on my local OS's
> shm objects and doesn't seem to fit mmap's job description given an
> fd*, but maybe I'm just confused. Anyway, a workaround seems to be
> to start out with PROT_NONE and MAP_NORESERVE, then mprotect(PROT_READ
> | PROT_WRITE) new regions after extending with ftruncate(), at least
> in simple tests...
Right, it's similar to the currently implemented space reservation, which also
goes with PROT_NONE and MAP_NORESERVE. I assume it boils down to the way how
memory reservation accounting in Linux works.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-21 14:16 ` Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Thomas Munro @ 2025-04-21 14:16 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Mon, Apr 21, 2025 at 9:30 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> Yeah, that would work and will allow to avoid MAP_FIXED and mremap, which are
> questionable from portability point of view. This leaves memfd_create, and I'm
> still not completely clear on it's portability -- it seems to be specific to
> Linux, but others provide compatible implementation as well.
Something like this should work, roughly based on DSM code except here
we don't really need the name so we unlink it immediately, at the
slight risk of leaking it if the postmaster is killed between those
lines (maybe someone should go and tell POSIX to support the special
name SHM_ANON or some other way to avoid that; I can't see any
portable workaround). Not tested/compiled, just a sketch:
#ifdef HAVE_MEMFD_CREATE
/* Anonymous shared memory region. */
fd = memfd_create("foo", MFD_CLOEXEC | huge_pages_flags);
#else
/* Standard POSIX insists on a name, which we unlink immediately. */
do
{
char tmp[80];
snprintf(tmp, sizeof(tmp), "PostgreSQL.%u",
pg_prng_uint32(&pg_global_prng_state));
fd.= shm_open(tmp, O_CREAT | O_EXCL);
if (fd >= 0)
shm_unlink(tmp);
} while (fd < 0 && errno == EXIST);
#endif
> Let me experiment with this idea a bit, I would like to make sure there are no
> other limitations we might face.
One thing I'm still wondering about is whether you really need all
this multi-phase barrier stuff, or even need to stop other backends
from running at all while doing the resize. I guess that's related to
your remapping scheme, but supposing you find the simple
ftruncate()-only approach to be good, my next question is: why isn't
it enough to wait for all backends to agree to stop allocating new
buffers in the range to be truncated, and then left them continue to
run as normal? As far as they would be concerned, the in-progress
downsize has already happened, though it could be reverted later if
the eviction phase fails. Then the coordinator could start evicting
buffers and truncating the shared memory object, which are
phases/steps, sure, but it's not clear to me why they need other
backends' help.
It sounds like Windows might need a second ProcSignalBarrier poke in
order to call VirtualUnlock() in every backend. That's based on that
Usenet discussion I lobbed in here the other day; I haven't tried it
myself or fully grokked why it works, and there could well be other
ways, IDK. Assuming it's the right approach, between the first poke
to make all backends accept the new lower size and the second poke to
unlock the memory, I don't see why they need to wait. I suppose it
would be the same ProcSignalBarrier, but behave differently based on a
control variables. I suppose there could also be a third poke, if you
want to consider the operation to be fully complete only once they
have all actually done that unlock step, but it may also be OK not to
worry about that, IDK.
On the other hand, maybe it just feels less risky if you stop the
whole world, or maybe you envisage parallelising the eviction work, or
there is some correctness concern I haven't grokked yet, but what?
> > *You might also want to use fallocate after ftruncate on Linux to
> > avoid SIGBUS on allocation failure on first touch page fault, which
> > raises portability questions since it's unspecified whether you can do
> > that with shm fds and fails on some systems, but it let's call that an
> > independent topic as it's not affected by this choice.
>
> I'm afraid it would be strictly neccessary to do fallocate, otherwise we're
> back where we were before reservation accounting for huge pages in Linux (lot's
> of people were facing unexpected SIGBUS when dealing with cgroups).
Yeah. FWIW here is where we decided to gate that on __linux__ while
fixing that for DSM:
https://www.postgresql.org/message-id/flat/CAEepm%3D0euOKPaYWz0-gFv9xfG%2B8ptAjhFjiQEX0CCJaYN--sDQ%4...
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
@ 2025-06-10 11:09 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-06-10 11:09 UTC (permalink / raw)
To: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Mon, Apr 21, 2025 at 7:47 PM Thomas Munro <thomas.munro@gmail.com> wrote:
> On Mon, Apr 21, 2025 at 9:30 PM Dmitry Dolgov <9erthalion6@gmail.com>
> wrote:
> > Yeah, that would work and will allow to avoid MAP_FIXED and mremap,
> which are
> > questionable from portability point of view. This leaves memfd_create,
> and I'm
> > still not completely clear on it's portability -- it seems to be
> specific to
> > Linux, but others provide compatible implementation as well.
>
> Something like this should work, roughly based on DSM code except here
> we don't really need the name so we unlink it immediately, at the
> slight risk of leaking it if the postmaster is killed between those
> lines (maybe someone should go and tell POSIX to support the special
> name SHM_ANON or some other way to avoid that; I can't see any
> portable workaround). Not tested/compiled, just a sketch:
>
> #ifdef HAVE_MEMFD_CREATE
> /* Anonymous shared memory region. */
> fd = memfd_create("foo", MFD_CLOEXEC | huge_pages_flags);
> #else
> /* Standard POSIX insists on a name, which we unlink immediately. */
> do
> {
> char tmp[80];
> snprintf(tmp, sizeof(tmp), "PostgreSQL.%u",
> pg_prng_uint32(&pg_global_prng_state));
> fd.= shm_open(tmp, O_CREAT | O_EXCL);
> if (fd >= 0)
> shm_unlink(tmp);
> } while (fd < 0 && errno == EXIST);
> #endif
>
> > Let me experiment with this idea a bit, I would like to make sure there
> are no
> > other limitations we might face.
>
> One thing I'm still wondering about is whether you really need all
> this multi-phase barrier stuff, or even need to stop other backends
> from running at all while doing the resize. I guess that's related to
> your remapping scheme, but supposing you find the simple
> ftruncate()-only approach to be good, my next question is: why isn't
> it enough to wait for all backends to agree to stop allocating new
> buffers in the range to be truncated, and then left them continue to
> run as normal? As far as they would be concerned, the in-progress
> downsize has already happened, though it could be reverted later if
> the eviction phase fails. Then the coordinator could start evicting
> buffers and truncating the shared memory object, which are
> phases/steps, sure, but it's not clear to me why they need other
> backends' help.
>
AFAIU, we required the phased approach since mremap needed to happen in
every backend after buffer eviction but before making modifications to the
shared memory. If we don't need to call mremap in every backend and just
ftruncate + initializing memory (when expanding buffers) is enough, I think
phased approach isn't needed. But I haven't tried it myself.
Here's patchset rebased on 3feff3916ee106c084eca848527dc2d2c3ef4e89.
0001 - 0008 are same as the previous patchset
0009 adds support to shrink shared buffers. It has two changes: a. evict
the buffers outside the new buffer size b. remove buffers with buffer id
outside the new buffer size from the free list. If a buffer being evicted
is pinned, the operation is aborted and a FATAL error is raised. I think we
need to change this behaviour to be less severe like rolling back the
operation or waiting for the pinned buffer to be unpinned etc. Better even
if we could let users control the behaviour. But we need better
infrastructure to do such things. That's one TODO left in the patch.
0010 is about reinitializing the Strategy reinitialization. Once we expand
the buffers, the new buffers need to be added to the free list. Some
StrategyControl area members (not all) need to be adjusted. That's what
this patch does. But a deeper adjustment in BgBufferSync() and
ClockSweepTick() is required. Further we need to do something about the
buffer lookup table. More on that later in the email.
0011-0012 fix compilation issues in these patches but those fixes are not
correct. The patches are there so that binaries can be built without any
compilation issues and someone can experiment with buffer resizing. Good
thing is the compilation fixes are in SQL callable functions
pg_get_shmem_pagesize() and pg_get_shmem_numa(). So there's no ill-effect
because of these patches as long as those two functions are not called.
Buffer lookup table resizing
------------------------------------
The size of the buffer lookup table depends upon (number of shared
buffers + number of partitions in the shared buffer lookup table). If we
shrink the buffer pool, the buffer lookup table will become sparse but
still useful. If we expand the buffers we need to expand the buffer lookup
table too. That's not implemented in the current patchset. There are two
solutions here:
1. We map a lot of extra address space (not memory) initially to
accomodate for future expansion of shared buffer pool. Let's say that the
total address space is sufficient to accomodate Nx buffers. Simple solution
is to allocate a buffer lookup table with Nx initial entries so that we
don't have to resize the buffer lookup table ever. It will waste memory but
we might be ok with that as version 1 solution. According to my offline
discussion with David Rowley, buffer lookups in sparse hash tables are
inefficient because or more cacheline faults. Whether that translates to
any noticeable performance degradation in TPS needs to be measured.
2. Alternate solution is to resize the buffer mapping table as well. This
means that we rehash all the entries again which may take a longer time and
the partitions will remain locked for that amount of time. Not to mention
this will require non-trivial change to dynahash implementation.
Next I will look at BgBufferSync() and ClockSweepTick() adjustments and
then buffer lookup table fix with approach 1.
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] 0005-Introduce-pss_barrierReceivedGeneration-20250610.patch (7.3K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/3-0005-Introduce-pss_barrierReceivedGeneration-20250610.patch)
download | inline diff:
From fb470c019742f2e9eaa7666ab81a24f816066387 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 4 Apr 2025 21:46:14 +0200
Subject: [PATCH 05/17] Introduce pss_barrierReceivedGeneration
Currently WaitForProcSignalBarrier allows to make sure the message sent
via EmitProcSignalBarrier was processed by all ProcSignal mechanism
participants.
Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
which will be updated when a process has received the message, but not
processed it yet. This makes it possible to support a new mode of
waiting, when ProcSignal participants want to synchronize message
processing. To do that, a participant can wait via
WaitForProcSignalBarrierReceived when processing a message, effectively
making sure that all processes are going to start processing
ProcSignalBarrier simultaneously.
---
src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------
src/include/storage/procsignal.h | 1 +
2 files changed, 54 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index a9bb540b55a..c6bec9be423 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -58,7 +58,10 @@
* of it. For such use cases, we set a bit in pss_barrierCheckMask and then
* increment the current "barrier generation"; when the new barrier generation
* (or greater) appears in the pss_barrierGeneration flag of every process,
- * we know that the message has been received everywhere.
+ * we know that the message has been received and processed everywhere. In case
+ * if we only need to know only that the message was received everywhere (e.g.
+ * receiving processes need to handle the message in a coordinated fashion)
+ * use pss_barrierReceivedGeneration in the same way.
*/
typedef struct
{
@@ -70,6 +73,7 @@ typedef struct
/* Barrier-related fields (not protected by pss_mutex) */
pg_atomic_uint64 pss_barrierGeneration;
+ pg_atomic_uint64 pss_barrierReceivedGeneration;
pg_atomic_uint32 pss_barrierCheckMask;
ConditionVariable pss_barrierCV;
} ProcSignalSlot;
@@ -152,6 +156,8 @@ ProcSignalShmemInit(void)
slot->pss_cancel_key_len = 0;
MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration,
+ PG_UINT64_MAX);
pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
ConditionVariableInit(&slot->pss_barrierCV);
}
@@ -199,6 +205,8 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
barrier_generation =
pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration);
pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration,
+ barrier_generation);
if (cancel_key_len > 0)
memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len);
@@ -263,6 +271,7 @@ CleanupProcSignalState(int status, Datum arg)
* no barrier waits block on it.
*/
pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX);
SpinLockRelease(&slot->pss_mutex);
@@ -416,12 +425,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type)
return generation;
}
-/*
- * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
- * requested by a specific call to EmitProcSignalBarrier() have taken effect.
- */
-void
-WaitForProcSignalBarrier(uint64 generation)
+static void
+WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly)
{
Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration));
@@ -436,12 +441,17 @@ WaitForProcSignalBarrier(uint64 generation)
uint64 oldval;
/*
- * It's important that we check only pss_barrierGeneration here and
- * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared
- * before the barrier is actually absorbed, but pss_barrierGeneration
+ * It's important that we check only pss_barrierGeneration &
+ * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in
+ * pss_barrierCheckMask get cleared before the barrier is actually
+ * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration
* is updated only afterward.
*/
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
while (oldval < generation)
{
if (ConditionVariableTimedSleep(&slot->pss_barrierCV,
@@ -450,7 +460,11 @@ WaitForProcSignalBarrier(uint64 generation)
ereport(LOG,
(errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier",
(int) pg_atomic_read_u32(&slot->pss_pid))));
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
}
ConditionVariableCancelSleep();
}
@@ -464,12 +478,33 @@ WaitForProcSignalBarrier(uint64 generation)
* The caller is probably calling this function because it wants to read
* the shared state or perform further writes to shared state once all
* backends are known to have absorbed the barrier. However, the read of
- * pss_barrierGeneration was performed unlocked; insert a memory barrier
- * to separate it from whatever follows.
+ * pss_barrierGeneration & pss_barrierReceivedGeneration was performed
+ * unlocked; insert a memory barrier to separate it from whatever follows.
*/
pg_memory_barrier();
}
+/*
+ * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
+ * requested by a specific call to EmitProcSignalBarrier() have taken effect.
+ */
+void
+WaitForProcSignalBarrier(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, false);
+}
+
+/*
+ * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all
+ * backends have observed the message sent by a specific call to
+ * EmitProcSignalBarrier().
+ */
+void
+WaitForProcSignalBarrierReceived(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, true);
+}
+
/*
* Handle receipt of an interrupt indicating a global barrier event.
*
@@ -523,6 +558,10 @@ ProcessProcSignalBarrier(void)
if (local_gen == shared_gen)
return;
+ /* The message is observed, record that */
+ pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
+ shared_gen);
+
/*
* Get and clear the flags that are set for this backend. Note that
* pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index afeeb1ca019..2733bbb8c5b 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, const uint8 *cancel_key, int cance
extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type);
extern void WaitForProcSignalBarrier(uint64 generation);
+extern void WaitForProcSignalBarrierReceived(uint64 generation);
extern void ProcessProcSignalBarrier(void);
extern void procsignal_sigusr1_handler(SIGNAL_ARGS);
--
2.34.1
[text/x-patch] 0004-Introduce-pending-flag-for-GUC-assign-hooks-20250610.patch (12.8K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/4-0004-Introduce-pending-flag-for-GUC-assign-hooks-20250610.patch)
download | inline diff:
From 1afa2048d803e0b3372de348868b26542ccfb3cd Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:40:32 +0200
Subject: [PATCH 04/17] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 61 insertions(+), 40 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 1914859b2ee..5e204341bde 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2321,7 +2321,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 608f10d9412..e40dae2ddf2 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,12 +1161,12 @@ assign_maintenance_io_concurrency(int newval, void *extra)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra)
+assign_io_max_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(newval, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra)
+assign_io_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(io_max_combine_limit, newval);
}
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index e5171467de1..2a6a587ef76 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1952,7 +1952,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1985,7 +1985,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2008,7 +2008,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2031,7 +2031,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 2f8c3d5f918..0d1b6466d1e 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3591,7 +3591,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 667df448732..bb681f5bc60 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f619100467d..8802ad8a3cb 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 799fa7ace68..c8300cffa8e 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,14 +81,14 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
-extern void assign_io_max_combine_limit(int newval, void *extra);
-extern void assign_io_combine_limit(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
+extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -143,13 +143,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -165,7 +165,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.34.1
[text/x-patch] 0001-Allow-to-use-multiple-shared-memory-mapping-20250610.patch (30.0K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/5-0001-Allow-to-use-multiple-shared-memory-mapping-20250610.patch)
download | inline diff:
From 25a501f17a36523be0b133f992393433428d73c5 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH 01/17] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 ++++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 141 +++++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 13 ++-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
12 files changed, 278 insertions(+), 125 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 423b2b4f9d6..4ce2cfb662b 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -307,7 +307,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -328,7 +328,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..56af0231d24 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 567739b5be9..5b55bec8d9d 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2fa045e6b0f..8b38e985327 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -363,7 +368,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index c9ae3b45b76..7e1a9b43fae 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -76,19 +76,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -101,9 +101,17 @@ Datum pg_numa_available(PG_FUNCTION_ARGS);
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -114,7 +122,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -123,9 +137,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -150,11 +164,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -184,6 +204,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -203,22 +229,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -236,6 +262,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -246,19 +278,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -273,7 +305,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -334,6 +372,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -350,9 +400,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -385,6 +435,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -393,7 +450,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -416,7 +473,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -433,8 +490,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -459,7 +516,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -478,14 +535,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -542,10 +598,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -557,15 +614,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 46f44bc4511..a36b08895c8 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -618,10 +620,15 @@ LWLockNewTrancheId(void)
int *LWLockCounter;
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
- /* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ /*
+ * We use the ShmemLock spinlock to protect LWLockCounter.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
+ */
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..6ebda479ced 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 5f7d4b83a60..2348c59b5a0 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -91,4 +106,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index c1f668ded95..69663d412c3 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: 3feff3916ee106c084eca848527dc2d2c3ef4e89
--
2.34.1
[text/x-patch] 0002-Address-space-reservation-for-shared-memory-20250610.patch (21.6K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/6-0002-Address-space-reservation-for-shared-memory-20250610.patch)
download | inline diff:
From acb2308031b68862a2f1238bf7ac803210063f71 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH 02/17] Address space reservation for shared memory
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is:
* To reserve some address space via mmap'ing a large chunk of memory
with PROT_NONE and MAP_NORESERVE. This way we prepare a playground for
preparing shared memory layout without risking anything interfering
with that.
* To slice the reserved space up into sections, one to use for each
shared segment.
* Allocate shared memory segments out of corresponding slices and
leaving unclaimed space in between them. This is implemented via
mmap'ing memory at a specified address from the reserved space with
MAP_FIXED.
The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
7f444196c000-7f470a800000 # reserved space
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
Things like address space randomization should not be a problem in this
context, since the randomization is applied to the mmap base, which is
one per process.
This approach also do not impact the actual memory usage as reported by
the kernel. Here is the output of /proc/$PID/status for the master
version with shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages
// mapped in mm_struct. It corresponds to the mapped reserved space
// and is the only number that grows with it.
VmPeak: 2043192 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 22908 kB
// Size of resident anonymous memory
RssAnon: 768 kB
// Size of resident file mappings
RssFile: 10364 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 11776 kB
Here is the same for the patch when reserving 20GB of space:
VmPeak: 21250648 kB
VmRSS: 22948 kB
RssAnon: 768 kB
RssFile: 10404 kB
RssShmem: 11776 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
17637376 (~16.8 MB)
To control the amount of space reserved a new GUC max_available_memory
is introduced. Ideally it should be based on the maximum available
memory, hense the name.
---
src/backend/port/sysv_shmem.c | 284 ++++++++++++++++++++++++----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 5 +-
src/backend/utils/init/globals.c | 1 +
src/backend/utils/misc/guc_tables.c | 14 ++
src/include/storage/pg_shmem.h | 4 +-
6 files changed, 271 insertions(+), 39 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 56af0231d24..a0f03ff868f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,66 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * To achieve that we first reserve some shared memory address space by
+ * mmap'ing a segment of MaxAvailableMemory size with PROT_NONE and
+ * MAP_NORESERVE (these flags allow to make sure this space will not be used by
+ * anything else, yet do not count against memory limits). Having the reserved
+ * space, we allocate out of it actual chunks of shared memory as usual,
+ * updating a pointer to the current available reserved space for the next
+ * allocation with the gap between segments in mind.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f442e010000-7f443a800000 # reserved empty space
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f444196c000-7f470a800000 # reserved empty space
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * [...]
+ *
+ * The reserved space pointer is calculated to slice up the total reserved
+ * space into fixed fractions of address space for each segment, as specified
+ * in the SHMEM_RESIZE_RATIO array.
+ */
+static double SHMEM_RESIZE_RATIO[1] = {
+ 1.0, /* MAIN_SHMEM_SLOT */
+};
+
+/*
+ * Offset from the beginning of the reserved space, which indicates currently
+ * available range. New shared memory segments have to be allocated at this
+ * offset related to the reserved space.
+ */
+static Size reserved_offset = 0;
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -626,39 +686,198 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
*
* This function will modify mapping size to the actual size of the allocation,
* if it ends up allocating a segment that is larger than requested.
+ *
+ * Note that we do not switch from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
*/
static void
-CreateAnonymousSegment(AnonymousMapping *mapping)
+CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
{
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS;
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* ReserveAnonymousMemory should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the request size to a suitable large value */
GetHugePageSize(&hugepagesize, &mmap_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
+ }
+#endif
+
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p",
+ MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+
+ /*
+ * Try to create mapping at an address out of the reserved range, which
+ * will allow to extend it later. Use reserved_offset to allocate the
+ * segment, then update currently available reserved range.
+ *
+ * If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized without
+ * a restart.
+ */
+ ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_FIXED, -1, 0);
+ mmap_errno = errno;
+
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m, "
+ "fallback to the non-resizable allocation",
+ MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS, -1, 0);
+ mmap_errno = errno;
+ }
+ else
+ {
+ Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+
+ reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ errno = mmap_errno;
+ DebugMappings();
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
+ (mmap_errno == ENOMEM) ?
+ errhint("This error usually means that PostgreSQL's request "
+ "for a shared memory segment exceeded available memory, "
+ "swap space, or huge pages. To reduce the request size "
+ "(currently %zu bytes), reduce PostgreSQL's shared "
+ "memory usage, perhaps by reducing \"shared_buffers\" or "
+ "\"max_connections\".",
+ allocsize) : 0));
+ }
+
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
+}
+
+/*
+ * ReserveAnonymousMemory
+ *
+ * Reserve shared memory address space, from which shared memory segments are
+ * going to be sliced out. The goal of this exercise is to support segments
+ * resizing, for which we need a reserved space free of potential clashes with
+ * other mmap'd areas that are not under our control. Reservation is done via
+ * mmap, and will not allocate any memory until it will be actually used, and
+ * MAP_NORESERVE allows to make it not counting againt kernel reservation
+ * limits (e.g. in cgroups or for huge pages). Do not get confused because of
+ * MAP_NORESERVE -- we need to reserve some space, but not the actual memory,
+ * and that is that this flag is about.
+ *
+ * Note, that with MAP_NORESERVE a reservation with hugetlb will succeed even
+ * if there is actually not enough huge pages. Hence this function is
+ * responsible for deciding whether to use huge pages or not. To achieve that
+ * we need to probe first and try to allocate needed memory for all segments --
+ * if this succeeds, we unmap the probe segment and use hugetlb; if it fails,
+ * we proceed with the regular memory.
+ */
+void *
+ReserveAnonymousMemory(Size reserve_size)
+{
+ Size allocsize = reserve_size;
+ void *ptr = MAP_FAILED;
+ int mmap_errno = 0;
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ *
+ * We could actually have a mix and match of segments with and without
+ * huge pages. But in that case we need to have multiple reservation
+ * spaces to use corresponding memory (hugetlb adress space reserved
+ * for hugetlb segments, regular memory for others), and it doesn't
+ * seem to worth the complexity for now.
+ */
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ int numSemas;
+ Size segment_size = CalculateShmemSize(&numSemas, segment);
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
- DebugMappings();
- elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- MappingName(mapping->shmem_segment), allocsize);
+ /* No huge pages, we will go with the regular page size */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB "
+ "failed, huge pages disabled: %m", total_size);
+ }
+ else
+ {
+ /*
+ * All fine, unmap the temporary segment and proceed with reserving
+ * using huge pages.
+ */
+ if (munmap(ptr, total_size) < 0)
+ elog(LOG, "reservice space: munmap(%p, %zu) failed: %m",
+ ptr, total_size);
+
+ /* Round up the requested size to a suitable large value. */
+ if (allocsize % hugepagesize != 0)
+ allocsize += hugepagesize - (allocsize % hugepagesize);
+
+ elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB",
+ allocsize);
+ ptr = mmap(NULL, allocsize, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | MAP_NORESERVE | mmap_flags,
+ -1, 0);
+ mmap_errno = errno;
+
+ /* This should not happen, but handle errors anyway */
+ if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
+ {
+ elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB "
+ "failed, huge pages disabled: %m", allocsize);
+ }
}
}
#endif
@@ -666,10 +885,12 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
/*
* Report whether huge pages are in use. This needs to be tracked before
* the second mmap() call if attempting to use huge pages failed
- * previously.
+ * previously. At this point ptr is either pointing to the probe segment,
+ * if we couldn't mmap it, or the reservation space.
*/
SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
@@ -677,10 +898,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
+ allocsize = reserve_size;
+
+ elog(DEBUG1, "reserving space: mmap(%zu)", allocsize);
+ ptr = mmap(NULL, allocsize, PROT_NONE,
+ MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0);
}
if (ptr == MAP_FAILED)
@@ -688,20 +910,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
errno = mmap_errno;
DebugMappings();
ereport(FATAL,
- (errmsg("segment[%s]: could not map anonymous shared memory: %m",
- MappingName(mapping->shmem_segment)),
+ (errmsg("reserving space: could not map anonymous shared "
+ "memory: %m"),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
- "for a shared memory segment exceeded available memory, "
- "swap space, or huge pages. To reduce the request size "
- "(currently %zu bytes), reduce PostgreSQL's shared "
- "memory usage, perhaps by reducing \"shared_buffers\" or "
- "\"max_connections\".",
+ "for a reserved shared memory address space exceeded "
+ "available memory, swap space, or huge pages. To "
+ "reduce the request reservation size (currently %zu "
+ "bytes), reduce PostgreSQL's \"maximum_shared_buffers\".",
allocsize) : 0));
}
- mapping->shmem = ptr;
- mapping->shmem_size = allocsize;
+ return ptr;
}
/*
@@ -740,7 +960,7 @@ AnonymousShmemDetach(int status, Datum arg)
*/
PGShmemHeader *
PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim)
+ PGShmemHeader **shim, Pointer base)
{
IpcMemoryKey NextShmemSegID;
void *memAddress;
@@ -760,14 +980,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -782,7 +994,7 @@ PGSharedMemoryCreate(Size size,
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
/* On success, mapping data will be modified. */
- CreateAnonymousSegment(mapping);
+ CreateAnonymousSegment(mapping, base);
next_free_segment++;
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..ce719f1b412 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -205,7 +205,7 @@ EnableLockPagesPrivilege(int elevel)
*/
PGShmemHeader *
PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim)
+ PGShmemHeader **shim, Pointer base)
{
void *memAddress;
PGShmemHeader *hdr;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8b38e985327..076888c0172 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -203,9 +203,12 @@ CreateSharedMemoryAndSemaphores(void)
PGShmemHeader *seghdr;
Size size;
int numSemas;
+ void *base;
Assert(!IsUnderPostmaster);
+ base = ReserveAnonymousMemory((Size) MaxAvailableMemory * BLCKSZ);
+
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
/* Compute the size of the shared-memory block */
@@ -217,7 +220,7 @@ CreateSharedMemoryAndSemaphores(void)
*
* XXX: Do multiple shims are needed, one per segment?
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(size, &shim, base);
/*
* Make sure that huge pages are never reported as "unknown" while the
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index d31cb45a058..9ccb7d455d6 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffers = 16384;
+int MaxAvailableMemory = 131072;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index f04bfedb2fd..e63521e5a2d 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2376,6 +2376,20 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"max_available_memory", PGC_SIGHUP, RESOURCES_MEM,
+ gettext_noop("Sets the upper limit for the shared_buffers value."),
+ gettext_noop("Shared memory could be resized at runtime, this "
+ "parameters sets the upper limit for it, beyond which "
+ "resizing would not be supported. Normally this value "
+ "would be the same as the total available memory."),
+ GUC_UNIT_BLOCKS
+ },
+ &MaxAvailableMemory,
+ 131072, 16, INT_MAX / 2,
+ NULL, NULL, NULL
+ },
+
{
{"vacuum_buffer_usage_limit", PGC_USERSET, RESOURCES_MEM,
gettext_noop("Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum."),
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 2348c59b5a0..8cb1e159917 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -61,6 +61,7 @@ extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT int MaxAvailableMemory;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -101,10 +102,11 @@ extern void PGSharedMemoryNoReAttach(void);
#endif
extern PGShmemHeader *PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim);
+ PGShmemHeader **shim, Pointer base);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+void *ReserveAnonymousMemory(Size reserve_size);
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
--
2.34.1
[text/x-patch] 0003-Introduce-multiple-shmem-segments-for-share-20250610.patch (11.7K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/7-0003-Introduce-multiple-shmem-segments-for-share-20250610.patch)
download | inline diff:
From 92e4e639547207e883bc010c485177afa32d5c72 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 15 Mar 2025 16:38:59 +0100
Subject: [PATCH 03/17] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 24 +++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 105 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a0f03ff868f..f46d9d5d9cd 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -147,10 +147,18 @@ static int next_free_segment = 0;
*
* The reserved space pointer is calculated to slice up the total reserved
* space into fixed fractions of address space for each segment, as specified
- * in the SHMEM_RESIZE_RATIO array.
+ * in the SHMEM_RESIZE_RATIO array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take
+ * up to 60% of the whole space when resizing, based on the fact that it most
+ * likely will be the main consumer of this memory. Those numbers are pulled
+ * out of thin air for now, makes sense to evaluate them more precise.
*/
-static double SHMEM_RESIZE_RATIO[1] = {
- 1.0, /* MAIN_SHMEM_SLOT */
+static double SHMEM_RESIZE_RATIO[6] = {
+ 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.6, /* BUFFERS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -182,6 +190,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1dc488a42..bd68b69ee98 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +77,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -156,33 +160,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..a9952b36eba 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -22,6 +22,7 @@
#include "postgres.h"
#include "storage/buf_internals.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -59,10 +60,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 01909be0272..bd390f2709d 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -491,9 +492,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 076888c0172..9d00b80b4f8 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 41fdc1e7693..edac9db6a12 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -318,7 +318,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 8cb1e159917..f8459a5a421 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -108,7 +108,29 @@ extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
void *ReserveAnonymousMemory(Size reserve_size);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.34.1
[text/x-patch] 0007-Use-anonymous-files-to-back-shared-memory-s-20250610.patch (10.7K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/8-0007-Use-anonymous-files-to-back-shared-memory-s-20250610.patch)
download | inline diff:
From 441f537b64b6bc8f0f00fa0de7850911acff621c Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 15 Mar 2025 16:39:45 +0100
Subject: [PATCH 07/17] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f90cde00000-7f90d5126000 rw-s 00000000 00:01 5463 /memfd:main (deleted)
7f90d5126000-7f914de00000 ---p 00000000 00:00 0
7f914de00000-7f9175128000 rw-s 00000000 00:01 5466 /memfd:buffers (deleted)
7f9175128000-7f944de00000 ---p 00000000 00:00 0
7f944de00000-7f9455528000 rw-s 00000000 00:01 5469 /memfd:descriptors (deleted)
7f9455528000-7f94cde00000 ---p 00000000 00:00 0
7f94cde00000-7f94d5228000 rw-s 00000000 00:01 5472 /memfd:iocv (deleted)
7f94d5228000-7f954de00000 ---p 00000000 00:00 0
7f954de00000-7f9555266000 rw-s 00000000 00:01 5475 /memfd:checkpoint (deleted)
7f9555266000-7f958de00000 ---p 00000000 00:00 0
7f958de00000-7f95954aa000 rw-s 00000000 00:01 5478 /memfd:strategy (deleted)
7f95954aa000-7f95cde00000 ---p 00000000 00:00 0
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 73 +++++++++++++++++++++++++++++-----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/portability/mem.h | 2 +-
src/include/storage/pg_shmem.h | 3 +-
5 files changed, 68 insertions(+), 14 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a3437973784..87000a24eea 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -107,6 +107,7 @@ typedef struct AnonymousMapping
Pointer shmem; /* Pointer to the start of the mapped memory */
Pointer seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -127,7 +128,7 @@ static int next_free_segment = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -150,9 +151,9 @@ static int next_free_segment = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* 7f442e010000-7f443a800000 # reserved empty space
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* 7f444196c000-7f470a800000 # reserved empty space
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -643,13 +644,14 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* *hugepagesize and *mmap_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -708,6 +710,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -718,7 +721,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -727,6 +739,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -734,6 +748,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -771,7 +787,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
- int mmap_flags = PG_MMAP_FLAGS;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
#ifndef MAP_HUGETLB
/* ReserveAnonymousMemory should have dealt with this case */
@@ -785,7 +801,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
/* Round up the request size to a suitable large value */
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
@@ -794,6 +810,29 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
}
#endif
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment),
+ memfd_flags);
+
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(mapping->shmem_segment))));
+
elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p",
MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
@@ -807,7 +846,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
* a restart.
*/
ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | MAP_FIXED, -1, 0);
+ mmap_flags | MAP_FIXED, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
@@ -817,8 +856,15 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
"fallback to the non-resizable allocation",
MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+ /* Specify the segment file size using allocsize. */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(mapping->shmem_segment))));
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
else
@@ -889,7 +935,7 @@ ReserveAnonymousMemory(Size reserve_size)
Size hugepagesize, total_size = 0;
int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
/*
* Figure out how much memory is needed for all segments, keeping in
@@ -1070,6 +1116,13 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ if(ftruncate(m->segment_fd, new_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
/* Clean up some reserved space to resize into */
if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
ereport(FATAL,
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index ce719f1b412..ba972106de1 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index abeb91e24fd..dc2b4becf4a 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -396,7 +396,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 19ad2e2f788..192b637cc65 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -125,7 +125,8 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size,
PGShmemHeader **shim, Pointer base);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
void *ReserveAnonymousMemory(Size reserve_size);
bool ProcessBarrierShmemResize(Barrier *barrier);
--
2.34.1
[text/x-patch] 0008-Support-resize-for-hugetlb-20250610.patch (4.3K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/9-0008-Support-resize-for-hugetlb-20250610.patch)
download | inline diff:
From 2ebc737cd5b22c4cb3fbcafb583c0bbd61fe93d0 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 5 Apr 2025 19:51:33 +0200
Subject: [PATCH 08/17] Support resize for hugetlb
Linux kernel has a set of limitations on remapping hugetlb segments: it
can't increase size of such segment [1], and shrinking it will not
release the memory back. In fact support for hugetlb mremap was
implemented no so long time ago [2].
As a workaround, avoid mremap for resizing shared memory. Instead unmap
the whole segment and map it back at the same address with the new size,
relying on the fact that fd for the anon file behind the segment is
still open and will keep the memory content.
[1]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/mremap.c?id=f4d2ef48250ad057e4f00087967b5ff366da9f39#n1593
[2]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/mm/mremap.c?id=550a7d60bd5e35a56942dba6d8a26752beb26c9f
---
src/backend/port/sysv_shmem.c | 60 +++++++++++++++++++++++++----------
1 file changed, 44 insertions(+), 16 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 87000a24eea..f0b53ce1d7c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1109,6 +1109,7 @@ AnonymousShmemResize(void)
/* Note that CalculateShmemSize indirectly depends on NBuffers */
Size new_size = CalculateShmemSize(&numSemas, i);
AnonymousMapping *m = &Mappings[i];
+ int mmap_flags = PG_MMAP_FLAGS;
if (m->shmem == NULL)
continue;
@@ -1116,6 +1117,44 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+#ifndef MAP_HUGETLB
+ /* ReserveAnonymousMemory should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ Size hugepagesize;
+
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the new size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ if (new_size % hugepagesize != 0)
+ new_size += hugepagesize - (new_size % hugepagesize);
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
+ }
+#endif
+
+ /*
+ * Linux limitations do not allow us to mremap hugetlb in the way we
+ * want. E.g. no size increase is allowed, and for shrinking the memory
+ * will not be released back. To work around this unmap the segment and
+ * create a new one at the same address. Thanks for the backing anon
+ * file the content will still be kept in memory.
+ */
+ elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ if (munmap(m->shmem, m->shmem_size) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not unmap shared memory segment %s [%p]: %m",
+ MappingName(m->shmem_segment), m->shmem)));
+
/* Resize the backing anon file. */
if(ftruncate(m->segment_fd, new_size) == -1)
ereport(FATAL,
@@ -1123,25 +1162,14 @@ AnonymousShmemResize(void)
errmsg("could not truncase anonymous file for \"%s\": %m",
MappingName(m->shmem_segment))));
- /* Clean up some reserved space to resize into */
- if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
- ereport(FATAL,
- (errcode(ERRCODE_SYSTEM_ERROR),
- errmsg("could not unmap %zu from reserved shared memory %p: %m",
- new_size - m->shmem_size, m->shmem)));
-
- /* Claim the unused space */
- elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
- MappingName(m->shmem_segment), m->shmem_size,
- new_size, m->shmem);
-
- ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
+ /* Reclaim the space */
+ ptr = mmap(m->shmem, new_size, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_FIXED, m->segment_fd, 0);
if (ptr == MAP_FAILED)
ereport(FATAL,
(errcode(ERRCODE_SYSTEM_ERROR),
- errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m",
- MappingName(m->shmem_segment), m->shmem, NBuffers,
- new_size)));
+ errmsg("could not map shared memory segment %s [%p] with size %zu: %m",
+ MappingName(m->shmem_segment), m->shmem, new_size)));
reinit = true;
m->shmem_size = new_size;
--
2.34.1
[text/x-patch] 0006-Allow-to-resize-shared-memory-without-resta-20250610.patch (39.8K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/10-0006-Allow-to-resize-shared-memory-without-resta-20250610.patch)
download | inline diff:
From 12a39ceb67b438c2fdf3727b51f5e4a9c105b3b0 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:47:16 +0200
Subject: [PATCH 06/17] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers and extend it using mremap. One elected process signals the
postmaster to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f90cde00000-7f90d4fa6000 /dev/zero (deleted)
7f90d4fa6000-7f914de00000
7f914de00000-7f915cfa8000 /dev/zero (deleted)
^ buffers mapping, ~241 MB
7f915cfa8000-7f944de00000
7f944de00000-7f94550a8000 /dev/zero (deleted)
7f94550a8000-7f94cde00000
7f94cde00000-7f94d4fe8000 /dev/zero (deleted)
7f94d4fe8000-7f954de00000
7f954de00000-7f9554ff6000 /dev/zero (deleted)
7f9554ff6000-7f958de00000
7f958de00000-7f959508a000 /dev/zero (deleted)
7f959508a000-7f95cde00000
-- 512 MB
7f90cde00000-7f90d5126000 /dev/zero (deleted)
7f90d5126000-7f914de00000
7f914de00000-7f9175128000 /dev/zero (deleted)
^ buffers mapping, ~627 MB
7f9175128000-7f944de00000
7f944de00000-7f9455528000 /dev/zero (deleted)
7f9455528000-7f94cde00000
7f94cde00000-7f94d5228000 /dev/zero (deleted)
7f94d5228000-7f954de00000
7f954de00000-7f9555266000 /dev/zero (deleted)
7f9555266000-7f958de00000
7f958de00000-7f95954aa000 /dev/zero (deleted)
7f95954aa000-7f95cde00000
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it receives something.
Note, that mremap is Linux specific, thus the implementation not very
portable.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 413 ++++++++++++++++++
src/backend/postmaster/postmaster.c | 18 +
src/backend/storage/buffer/buf_init.c | 75 ++--
src/backend/storage/ipc/ipci.c | 18 +-
src/backend/storage/ipc/procsignal.c | 46 ++
src/backend/storage/ipc/shmem.c | 23 +-
src/backend/tcop/postgres.c | 10 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/miscadmin.h | 1 +
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 26 ++
src/include/storage/pmsignal.h | 1 +
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
17 files changed, 603 insertions(+), 43 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index f46d9d5d9cd..a3437973784 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -105,6 +111,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -176,6 +189,49 @@ static Size reserved_offset = 0;
*/
static bool huge_pages_on = false;
+/*
+ * Flag telling that we have prepared the memory layout to be resizable. If
+ * false after all shared memory segments creation, it means we failed to setup
+ * needed layout and falled back to the regular non-resizable approach.
+ */
+static bool shmem_resizable = false;
+
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -769,6 +825,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
{
Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+ shmem_resizable = true;
reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
}
@@ -964,6 +1021,315 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * mremap, which will fail if is not sufficient space to expand the mapping.
+ * When finished, based on the new and old values initialize new buffer blocks
+ * if any.
+ *
+ * If reinitializing took place, as the last step this function does buffers
+ * reinitialization as well and broadcasts the new value of NSharedBuffers. All
+ * of that needs to be done only by one backend, the first one that managed to
+ * grab the ShmemResizeLock.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ void *ptr = MAP_FAILED;
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ /* Clean up some reserved space to resize into */
+ if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not unmap %zu from reserved shared memory %p: %m",
+ new_size - m->shmem_size, m->shmem)));
+
+ /* Claim the unused space */
+ elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
+ if (ptr == MAP_FAILED)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m",
+ MappingName(m->shmem_segment), m->shmem, NBuffers,
+ new_size)));
+
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * Note that it's enough when only one backend will do that,
+ * even the ShmemInitStruct part. The reason is that resized
+ * shared memory will maintain the same addresses, meaning that
+ * all the pointers are still valid, and we only need to update
+ * structures size in the ShmemIndex once -- any other backend
+ * will pick up this shared structure from the index.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ BufferManagerShmemInit(NBuffersOld);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Wait for all ProcSignal participants
+ * to join the barrier, then do the resize and wait on the barrier until all
+ * participating finish resizing as well -- otherwise we face danger of
+ * inconsistency between backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * First thing to do after attaching to the barrier is to wait for others.
+ * We can't simply use BarrierArriveAndWait, because backends might arrive
+ * here in disjoint groups, e.g. first two backends, pause, then second two
+ * backends. If the resize is quick enough that can lead to a situation
+ * when the first group is already finished before the second has appeared,
+ * and the barrier will only synchonize withing those groups.
+ */
+ if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
+ WaitForProcSignalBarrierReceived(
+ pg_atomic_read_u64(&ShmemCtrl->Generation));
+
+ /*
+ * Now start the procedure, and elect one backend to ping postmaster to do
+ * the same.
+ *
+ * XXX: If we need to be able to abort resizing, this has to be done later,
+ * after the SHMEM_RESIZE_DONE.
+ */
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ {
+ Assert(IsUnderPostmaster);
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ }
+
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ BarrierDetach(barrier);
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ elog(DEBUG1, "Set pending signal");
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Start resizing procedure, making sure all existing processes will have
+ * consistent view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * We use dynamic barrier to help dealing with backends that were spawned
+ * during the resize.
+ */
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ *
+ * Those phases are ensured by joining the shared barrier associated with
+ * the procedure. Since resizing takes time, we need to take into account
+ * that during that time:
+ *
+ * - New backends can be spawned. They will check status of the barrier
+ * early during the bootstrap, and wait until everything is over to work
+ * with the new NBuffers value.
+ *
+ * - Old backends can exit before attempting to resize. Synchronization
+ * used between backends relies on ProcSignalBarrier and waits for all
+ * participants received the message at the beginning to gather all
+ * existing backends.
+ *
+ * - Some backends might be blocked and not responsing either before or
+ * after receiving the message. In the first case such backend still
+ * have ProcSignalSlot and should be waited for, in the second case
+ * shared barrier will make sure we still waiting for those backends. In
+ * any case there is an unbounded wait.
+ *
+ * - Backends might join barrier in disjoint groups with some time in
+ * between. That means that relying only on the shared dynamic barrier is
+ * not enough -- it will only synchronize resize procedure withing those
+ * groups. That's why we wait first for all participants of ProcSignal
+ * mechanism who received the message.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ pg_atomic_init_u64(&ShmemCtrl->Generation,
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
+
+ /* To order everything after setting Generation value */
+ pg_memory_barrier();
+
+ /*
+ * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
+ * the rest of the pack has started the procedure and it can resize shared
+ * memory as well.
+ *
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier. In fact even if we would like to
+ * wait here, it wouldn't be possible -- we're in the postmaster, without
+ * any waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1271,3 +1637,50 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier()
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ /* Nothing to do if resizing is not started */
+ if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
+ return;
+
+ BarrierAttach(barrier);
+
+ /* Otherwise wait through all available phases */
+ while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
+ {
+ ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
+ BarrierPhase(barrier))));
+
+ BarrierArriveAndWait(barrier, 0);
+ }
+
+ BarrierDetach(barrier);
+}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /*
+ * The barrier is missing here, it will be initialized right before
+ * starting the resizing process as a convenient way to reset it.
+ */
+
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+
+ /* shmem_resizable should be initialized by now */
+ ShmemCtrl->Resizable = shmem_resizable;
+ }
+}
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 490f7ce3664..f0cb0098dcd 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -426,6 +426,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1694,6 +1695,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2039,6 +2043,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
@@ -3852,6 +3867,9 @@ process_pm_pmsignal(void)
request_state_update = true;
}
+ if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
+ AnonymousShmemResize();
+
/*
* Try to advance postmaster's state machine, if a child requests it.
*/
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index bd68b69ee98..ac844b114bd 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -24,7 +25,6 @@ ConditionVariableMinimallyPadded *BufferIOCVArray;
WritebackContext BackendWritebackContext;
CkptSortItem *CkptBufferIds;
-
/*
* Data Structures:
* buffers live in a freelist and a lookup data structure.
@@ -62,18 +62,28 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * postmaster, or in a standalone backend) or during shared-memory resize. Size
+ * of data structures initialized here depends on NBuffers, and to be able to
+ * change NBuffers without a restart we store each structure into a separate
+ * shared memory segment, which could be resized on demand.
+ *
+ * FirstBufferToInit tells where to start initializing buffers. For
+ * initialization it always will be zero, but when resizing shared-memory it
+ * indicates the number of already initialized buffers.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(void)
+BufferManagerShmemInit(int FirstBufferToInit)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
+ elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
+ FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -110,43 +120,44 @@ BufferManagerShmemInit(void)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory, or in all cases when resizing shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
- ClearBufferTag(&buf->tag);
+#ifndef EXEC_BACKEND
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = FirstBufferToInit; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ ClearBufferTag(&buf->tag);
- buf->buf_id = i;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- pgaio_wref_clear(&buf->io_wref);
+ buf->buf_id = i;
- /*
- * Initially link all the buffers together as unused. Subsequent
- * management of this list is done by freelist.c.
- */
- buf->freeNext = i + 1;
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
+#endif
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 9d00b80b4f8..abeb91e24fd 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -84,6 +84,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -151,6 +154,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, AioShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -298,7 +309,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit();
+ BufferManagerShmemInit(0);
/*
* Set up lock manager
@@ -310,6 +321,11 @@ CreateOrAttachShmemStructs(void)
*/
PredicateLockShmemInit();
+ /*
+ * Set up shared memory resize manager
+ */
+ ShmemControlInit();
+
/*
* Set up process table
*/
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index c6bec9be423..d7b56a18b24 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -113,6 +114,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -176,6 +181,43 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -615,6 +657,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 7e1a9b43fae..c07572d6f89 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -498,17 +498,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
+ *
+ * XXX: There is an implicit assumption this can only happen in
+ * "resizable" segments, where only one shared structure is allowed.
+ * This has to be implemented more cleanly.
*/
if (result->size != size)
{
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
+ Size delta = size - result->size;
+
+ result->size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
}
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 0d1b6466d1e..0942d2bffe2 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4309,6 +4310,15 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /* Verify the shared barrier, if it's still active: join and wait. */
+ WaitOnShmemBarrier();
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 4da68312b5f..691fa14e9e3 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
@@ -352,6 +354,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index e63521e5a2d..9f00608f508 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2366,14 +2366,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 1bef98471c3..a0c37a7749e 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int MaxAvailableMemory;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index edac9db6a12..4239ebe640b 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -317,7 +317,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_skipped);
/* in buf_init.c */
-extern void BufferManagerShmemInit(void);
+extern void BufferManagerShmemInit(int);
extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6ebda479ced..bb7ae4d33b3 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index a9681738146..558da6fdd55 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
+PG_LWLOCK(54, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index f8459a5a421..19ad2e2f788 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,25 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+ pg_atomic_uint64 Generation;
+ bool Resizable;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -108,6 +128,12 @@ extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
void *ReserveAnonymousMemory(Size reserve_size);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 428aa3fd68a..1a55bf57a70 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,6 +42,7 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 2733bbb8c5b..97033f84dce 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index a8346cda633..b026a275c38 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2745,6 +2745,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.34.1
[text/x-patch] 0009-Support-shrinking-shared-buffers-20250610.patch (13.2K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/11-0009-Support-shrinking-shared-buffers-20250610.patch)
download | inline diff:
From 44dd06152fd9b8f65f80c81420974cb77e12e237 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 9 Jun 2025 14:40:34 +0530
Subject: [PATCH 09/17] Support shrinking shared buffers
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
Buffer eviction requires a separate barrier phase for two reasons:
1. No other backend should map a new page to any of buffers being
evicted when eviction is in progress. So they wait while eviction is
in progress.
2. Since a pinned buffer has the pin recorded in the backend local
memory as well as the buffer descriptor (which is in shared memory),
eviction should not coincide with remapping the shared memory of a
backend. Otherwise we might loose consistency of local and shared
pinning records. Hence it needs to be carried out in
ProcessBarrierShmemResize() and not in AnonymousShmemResize() as
indicated by now removed comment.
If a buffer being evicted is pinned, we raise a FATAL error but this should
improve. There are multiple options 1. to wait for the pinned buffer to get
unpinned, 2. the backend is killed or it itself cancels the query or 3.
rollback the operation. Note that option 1 and 2 would require the pinning
related local and shared records to be accessed. But we need infrastructure to
do either of this right now.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 30 ++++---
src/backend/storage/buffer/buf_init.c | 8 +-
src/backend/storage/buffer/bufmgr.c | 89 +++++++++++++++++++
src/backend/storage/buffer/freelist.c | 68 ++++++++++++++
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 1 +
src/include/storage/pg_shmem.h | 1 +
8 files changed, 187 insertions(+), 12 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index f0b53ce1d7c..03aa6a41828 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1096,14 +1096,6 @@ AnonymousShmemResize(void)
*/
pending_pm_shmem_resize = false;
- /*
- * XXX: Currently only increasing of shared_buffers is supported. For
- * decreasing something similar has to be done, but buffer blocks with
- * data have to be drained first.
- */
- if(NBuffersOld > NBuffers)
- return false;
-
for(int i = 0; i < next_free_segment; i++)
{
/* Note that CalculateShmemSize indirectly depends on NBuffers */
@@ -1201,8 +1193,6 @@ AnonymousShmemResize(void)
* all the pointers are still valid, and we only need to update
* structures size in the ShmemIndex once -- any other backend
* will pick up this shared structure from the index.
- *
- * XXX: This is the right place for buffer eviction as well.
*/
BufferManagerShmemInit(NBuffersOld);
@@ -1262,6 +1252,25 @@ ProcessBarrierShmemResize(Barrier *barrier)
SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
}
+ /*
+ * Evict extra buffers when shrinking shared buffers. We need to do this
+ * while the memory for extra buffers is still mapped i.e. before remapping
+ * the shared memory segments to a smaller memory area.
+ */
+ if (NBuffersOld > NBuffersPending)
+ {
+ /*
+ * TODO: If the buffer eviction fails for any reason, we should
+ * gracefully rollback the shared buffer resizing and try again. But the
+ * infrastructure to do so is not available right now. Hence just raise
+ * a FATAL so that the system restarts.
+ */
+ if (!EvictExtraBuffers(NBuffersPending, NBuffersOld))
+ elog(FATAL, "buffer eviction failed");
+
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_EVICT);
+ }
+
AnonymousShmemResize();
/* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
@@ -1763,5 +1772,6 @@ ShmemControlInit(void)
/* shmem_resizable should be initialized by now */
ShmemCtrl->Resizable = shmem_resizable;
+ ShmemCtrl->evictor_pid = 0;
}
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ac844b114bd..f78be4700df 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -156,8 +156,12 @@ BufferManagerShmemInit(int FirstBufferToInit)
}
#endif
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ /*
+ * Correct last entry of linked list, when initializing the buffers or when
+ * expanding the buffers.
+ */
+ if (FirstBufferToInit < NBuffers)
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 667aa0c0c78..57d78c482bb 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/read_stream.h"
#include "storage/smgr.h"
@@ -7453,3 +7454,91 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int newBufSize, int oldBufSize)
+{
+ bool result = true;
+
+ /*
+ * If the buffer being evicated is locked, this function will need to wait.
+ * This function should not be called from a Postmaster since it can not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * Let only one backend perform eviction. We could split the work across all
+ * the backends but that doesn't seem necessary.
+ *
+ * The first backend to acquire ShmemResizeLock, sets its own PID as the
+ * evictor PID for other backends to know that the eviction is in progress or
+ * has already been performed. The evictor backend releases the lock when it
+ * finishes eviction. While the eviction is in progress, backends other than
+ * evictor backend won't be able to take the lock. They won't perform
+ * eviction. A backend may acquire the lock after eviction has completed, but
+ * it will not perform eviction since the evictor PID is already set. Evictor
+ * PID is reset only when the buffer resizing finishes. Thus only one backend
+ * will perform eviction in a given instance of shared buffers resizing.
+ *
+ * Any backend which acquires this lock will release it before the eviction
+ * phase finishes, hence the same lock can be reused for the next phase of
+ * resizing buffers.
+ */
+ if (LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ if (ShmemCtrl->evictor_pid == 0)
+ {
+ ShmemCtrl->evictor_pid = MyProcPid;
+
+ StrategyPurgeFreeList(newBufSize);
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = newBufSize + 1; buf <= oldBufSize; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u32(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is
+ * going one hence unlocked precheck should be safe and saves
+ * some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+ }
+ LWLockRelease(ShmemResizeLock);
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index bd390f2709d..e384e46c779 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -527,6 +527,74 @@ StrategyInitialize(bool init)
}
+/*
+ * StrategyPurgeFreeList -- remove all buffers with id higher than the number of
+ * buffers in the buffer pool.
+ *
+ * This is called before evicting buffers while shrinking shared buffers, so that
+ * the free list does not reference a buffer that will be removed.
+ *
+ * The function is called after resizing has started and thus nobody should be
+ * traversing the free list and also not touching the buffers.
+ */
+void
+StrategyPurgeFreeList(int numBuffers)
+{
+ int firstBuffer = FREENEXT_END_OF_LIST;
+ int nextFree = StrategyControl->firstFreeBuffer;
+ BufferDesc *prevValidBuf = NULL;
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ while (nextFree != FREENEXT_END_OF_LIST)
+ {
+ BufferDesc *buf = GetBufferDescriptor(nextFree);
+
+ /* nextFree should be id of buffer being examined. */
+ Assert(nextFree == buf->buf_id);
+ /* The buffer should not be marked as not in the list. */
+ Assert(buf->freeNext != FREENEXT_NOT_IN_LIST);
+
+ /*
+ * If the buffer is within the new size of pool, keep it in the free list
+ * otherwise discard it.
+ */
+ if (buf->buf_id < numBuffers)
+ {
+ if (prevValidBuf != NULL)
+ prevValidBuf->freeNext = buf->buf_id;
+ prevValidBuf = buf;
+
+ /* Save the first free buffer in the list if not already known. */
+ if (firstBuffer == FREENEXT_NOT_IN_LIST)
+ firstBuffer = nextFree;
+ }
+ /* Examine the next buffer in the free list. */
+ nextFree = buf->freeNext;
+ }
+
+ /* Update the last valid free buffer, if there's any. */
+ if (prevValidBuf != NULL)
+ {
+ StrategyControl->lastFreeBuffer = prevValidBuf->buf_id;
+ prevValidBuf->freeNext = FREENEXT_END_OF_LIST;
+ }
+ else
+ StrategyControl->lastFreeBuffer = FREENEXT_END_OF_LIST;
+
+ /* Update first valid free buffer, if there's any. */
+ StrategyControl->firstFreeBuffer = firstBuffer;
+
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+
+ /*
+ * TODO: following was suggested by AI. Check whether it is required.
+ * If we removed all buffers from the freelist, reset the clock sweep
+ * pointer to zero. This is not strictly necessary, but it seems like a
+ * good idea to avoid confusion.
+ */
+}
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
* ----------------------------------------------------------------
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 691fa14e9e3..0c588b69a90 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -156,6 +156,7 @@ REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it c
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_EVICT "Waiting for other backends to finish buffer evication phase."
SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 0dec7d93b3b..add15e3723b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -453,6 +453,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyPurgeFreeList(int numBuffers);
extern bool have_free_buffer(void);
/* buf_table.c */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 4239ebe640b..0c554f0b130 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -315,6 +315,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_evicted,
int32 *buffers_flushed,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int fromBuf, int toBuf);
/* in buf_init.c */
extern void BufferManagerShmemInit(int);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 192b637cc65..23998f5469d 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -64,6 +64,7 @@ extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
typedef struct
{
pg_atomic_uint32 NSharedBuffers;
+ pid_t evictor_pid;
Barrier Barrier;
pg_atomic_uint64 Generation;
bool Resizable;
--
2.34.1
[text/x-patch] 0010-Reinitialize-StrategyControl-after-resizing-20250610.patch (7.8K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/12-0010-Reinitialize-StrategyControl-after-resizing-20250610.patch)
download | inline diff:
From 0f7d58f7386d0e3a55bdcf931492b62ee883a998 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 10 Jun 2025 11:00:36 +0530
Subject: [PATCH 10/17] Reinitialize StrategyControl after resizing buffers
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. When expanding the buffer pool add new buffers to the free list.
2. When shrinking buffers, we remove any buffers, in the area being
shrunk, from the freelist. While doing so we adjust the first and
last free buffer pointers in the StrategyControl area. Hence nothing
more needed after resizing.
3. Check the sanity of the free buffer list is added after resizing.
4. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
5. &StrategyControl->buffer_strategy_lock need not be initialized again.
6. completePasses and numBufferAllocs need not be cleared since the
server is still running and the previous statistics is still valid.
TODO: Since StrategyControl plays a crucial role in background writer as
well as clock tick algorithm, the impact of resizing the buffers on
those two needs to be assessed. We may require more adjustments to those
two as well as StrategyControl based on the assessment.
---
src/backend/storage/buffer/buf_init.c | 11 ++-
src/backend/storage/buffer/freelist.c | 120 ++++++++++++++++++++++++++
src/include/storage/buf_internals.h | 1 +
3 files changed, 130 insertions(+), 2 deletions(-)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index f78be4700df..7b8bc577bd5 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -163,8 +163,15 @@ BufferManagerShmemInit(int FirstBufferToInit)
if (FirstBufferToInit < NBuffers)
GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
- /* Init other shared buffer-management stuff */
- StrategyInitialize(!foundDescs);
+ /*
+ * Init other shared buffer-management stuff from scratch configuring buffer
+ * pool the first time. If we are just resizing buffer pool adjust only the
+ * required structures.
+ */
+ if (FirstBufferToInit == 0)
+ StrategyInitialize(!foundDescs);
+ else
+ StrategyReInitialize(FirstBufferToInit);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index e384e46c779..c5277e090a5 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -98,6 +98,9 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
uint32 *buf_state);
static void AddBufferToRing(BufferAccessStrategy strategy,
BufferDesc *buf);
+#ifdef USE_ASSERT_CHECKING
+static void StrategyValidateFreeList(void);
+#endif /* USE_ASSERT_CHECKING */
/*
* ClockSweepTick - Helper routine for StrategyGetBuffer()
@@ -526,6 +529,75 @@ StrategyInitialize(bool init)
Assert(!init);
}
+/*
+ * StrategyReInitialize -- re-initialize the buffer cache replacement
+ * strategy.
+ *
+ * To be called when resizing buffer manager and only from the coordinator.
+ * TODO: Assess the differences between this function and StrategyInitialize().
+ */
+void
+StrategyReInitialize(int FirstBufferIdToInit)
+{
+ bool found;
+
+ /*
+ * Resizing memory for buffer pools should not affect the address of
+ * StrategyControl.
+ */
+ if (StrategyControl != (BufferStrategyControl *)
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, STRATEGY_SHMEM_SEGMENT))
+ elog(FATAL, "something went wrong while re-initializing the buffer strategy");
+
+ /* TODO: Buffer lookup table adjustment: There are two options:
+ *
+ * 1. Resize the buffer lookup table to match the new number of buffers. But
+ * this requires rehashing all the entries in the buffer lookup table with
+ * the new table size.
+ *
+ * 2. Allocate maximum size of the buffer lookup table at the beginning and
+ * never resize it. This leaves sparse buffer lookup table which is
+ * inefficient from both memory and time perspective. According to David
+ * Rowley, the sparse entries in the buffer look up table cause frequent
+ * cacheline reload which affect performance. If the impact of that
+ * inefficiency in a benchmark is significant, we will need to consider first
+ * option.
+ */
+
+ /*
+ * When shrinking buffers, we must have adjusted the first and the last free
+ * buffer when removing the buffers being shrunk from the free list. Nothing
+ * to be done here.
+ *
+ * When expanding the shared buffers, new buffers are added at the end of the
+ * freelist or they form the new free list if there are no free buffers.
+ */
+ if (FirstBufferIdToInit < NBuffers)
+ {
+ if (StrategyControl->firstFreeBuffer == FREENEXT_END_OF_LIST)
+ StrategyControl->firstFreeBuffer = FirstBufferIdToInit;
+ else
+ {
+ Assert(StrategyControl->lastFreeBuffer >= 0);
+ GetBufferDescriptor(StrategyControl->lastFreeBuffer - 1)->freeNext = FirstBufferIdToInit;
+ }
+
+ StrategyControl->lastFreeBuffer = NBuffers - 1;
+ }
+
+ /* Check free list sanity after resizing. */
+#ifdef USE_ASSERT_CHECKING
+ StrategyValidateFreeList();
+#endif /* USE_ASSERT_CHECKING */
+
+ /* Initialize the clock sweep pointer */
+ pg_atomic_init_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /* No pending notification */
+ StrategyControl->bgwprocno = -1;
+}
/*
* StrategyPurgeFreeList -- remove all buffers with id higher than the number of
@@ -595,6 +667,54 @@ StrategyPurgeFreeList(int numBuffers)
*/
}
+#ifdef USE_ASSERT_CHECKING
+/*
+ * StrategyValidateFreeList-- check sanity of free buffer list.
+ */
+static void
+StrategyValidateFreeList(void)
+{
+ int nextFree = StrategyControl->firstFreeBuffer;
+ int numFreeBuffers = 0;
+ int lastFreeBuffer = FREENEXT_END_OF_LIST;
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ while (nextFree != FREENEXT_END_OF_LIST)
+ {
+ BufferDesc *buf = GetBufferDescriptor(nextFree);
+
+ /* nextFree should be id of buffer being examined. */
+ Assert(nextFree == buf->buf_id);
+ Assert(buf->buf_id < NBuffers);
+ /* The buffer should not be marked as not in the list. */
+ Assert(buf->freeNext != FREENEXT_NOT_IN_LIST);
+
+ /* Update our knowledge of last buffer in the free list. */
+ lastFreeBuffer = buf->buf_id;
+
+ numFreeBuffers++;
+
+ /* Avoid infinite recursion in case there are cycles in free list. */
+ if (numFreeBuffers > NBuffers)
+ break;
+
+ nextFree = buf->freeNext;
+ }
+
+ Assert(numFreeBuffers <= NBuffers);
+
+ /*
+ * Make sure that the StrategyControl's knowledge of last free buffer
+ * agrees with what's there in the free list.
+ */
+ if (StrategyControl->firstFreeBuffer != FREENEXT_END_OF_LIST)
+ Assert(StrategyControl->lastFreeBuffer == lastFreeBuffer);
+
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+}
+#endif /* USE_ASSERT_CHECKING */
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
* ----------------------------------------------------------------
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index add15e3723b..46949e9d90e 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -454,6 +454,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
extern void StrategyPurgeFreeList(int numBuffers);
+extern void StrategyReInitialize(int FirstBufferToInit);
extern bool have_free_buffer(void);
/* buf_table.c */
--
2.34.1
[text/x-patch] 0012-Fix-compilation-failure-in-pg_get_shmem_pag-20250610.patch (959B, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/13-0012-Fix-compilation-failure-in-pg_get_shmem_pag-20250610.patch)
download | inline diff:
From bf5b3cca2e8b26ce6ecf634905074ae930e3745a Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 5 Jun 2025 14:42:53 +0530
Subject: [PATCH 12/17] Fix compilation failure in pg_get_shmem_pagesize()
Fix compilation failure in pg_get_shmem_pagesize() due to incorrect call to
GetHugePageSize(). This is a temporary fix to allow compilation to proceed.
Ashutosh Bapat
---
src/backend/storage/ipc/shmem.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index b411fbce37e..4c2bddfe6ca 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -826,7 +826,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
--
2.34.1
[text/x-patch] 0011-Fix-compilation-failure-in-pg_get_shmem_all-20250610.patch (1.4K, ../../CAExHW5v9cE+ETusTafZyvy+eVpnVvoayVm7ZOk4Ddq8fxY270A@mail.gmail.com/14-0011-Fix-compilation-failure-in-pg_get_shmem_all-20250610.patch)
download | inline diff:
From f079821259a3841637b794d2d8ad1153e14eb4b3 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 5 Jun 2025 11:34:12 +0530
Subject: [PATCH 11/17] Fix compilation failure in
pg_get_shmem_allocations_numa()
The compilation failure is caused by
5cefa489760e34d947dbe67b4a922468b2e43668. Ideal fix should be compute
the total page count across all the shared memory segments. This commit
just fixes the compilation failure. # Please enter the commit message
for your changes. Lines starting
Ashutosh Bapat
---
src/backend/storage/ipc/shmem.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index c07572d6f89..b411fbce37e 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -696,7 +696,12 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ /*
+ * TODO: We should loop through all the Shm segments, instead of just the
+ * main segment, to find the total page count.
+ */
+ shm_total_page_count = (Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize
+ / os_page_size) + 1;
page_ptrs = palloc0(sizeof(void *) * shm_total_page_count);
pages_status = palloc(sizeof(int) * shm_total_page_count);
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-06-16 12:39 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-06-16 12:39 UTC (permalink / raw)
To: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Tue, Jun 10, 2025 at 4:39 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
Here's patchset rebased on f85f6ab051b7cf6950247e5fa6072c4130613555
with some more fixes as described below.
> 0001 - 0008 are same as the previous patchset
>
> 0009 adds support to shrink shared buffers. It has two changes: a. evict the buffers outside the new buffer size b. remove buffers with buffer id outside the new buffer size from the free list. If a buffer being evicted is pinned, the operation is aborted and a FATAL error is raised. I think we need to change this behaviour to be less severe like rolling back the operation or waiting for the pinned buffer to be unpinned etc. Better even if we could let users control the behaviour. But we need better infrastructure to do such things. That's one TODO left in the patch.
>
Patches upto 0009 are same as the previous patch set.
> 0010 is about reinitializing the Strategy reinitialization. Once we expand the buffers, the new buffers need to be added to the free list. Some StrategyControl area members (not all) need to be adjusted. That's what this patch does. But a deeper adjustment in BgBufferSync() and ClockSweepTick() is required. Further we need to do something about the buffer lookup table. More on that later in the email.
0010 is improved with fixes for background writer and clocksweeptick.
Now we just reset the information saved between calls to BgBufferSync
since it doesn't make sense after NBuffers has changed. Also the
members in StrategyControl related to ClockSweepTick are reset for the
same reason. More details in the commit message.
0011: GetBufferFromRing() invalidates the buffers beyond NBuffers
since those may have been added before resizing and are not valid
anymore. Details in commit message.
>
> 0011-0012 fix compilation issues in these patches but those fixes are not correct. The patches are there so that binaries can be built without any compilation issues and someone can experiment with buffer resizing. Good thing is the compilation fixes are in SQL callable functions pg_get_shmem_pagesize() and pg_get_shmem_numa(). So there's no ill-effect because of these patches as long as those two functions are not called.
These patches are now 0012 and 0013 respectively.
>
> Buffer lookup table resizing
> ------------------------------------
> The size of the buffer lookup table depends upon (number of shared buffers + number of partitions in the shared buffer lookup table). If we shrink the buffer pool, the buffer lookup table will become sparse but still useful. If we expand the buffers we need to expand the buffer lookup table too. That's not implemented in the current patchset. There are two solutions here:
>
> 1. We map a lot of extra address space (not memory) initially to accomodate for future expansion of shared buffer pool. Let's say that the total address space is sufficient to accomodate Nx buffers. Simple solution is to allocate a buffer lookup table with Nx initial entries so that we don't have to resize the buffer lookup table ever. It will waste memory but we might be ok with that as version 1 solution. According to my offline discussion with David Rowley, buffer lookups in sparse hash tables are inefficient because or more cacheline faults. Whether that translates to any noticeable performance degradation in TPS needs to be measured.
>
> 2. Alternate solution is to resize the buffer mapping table as well. This means that we rehash all the entries again which may take a longer time and the partitions will remain locked for that amount of time. Not to mention this will require non-trivial change to dynahash implementation.
I haven't spent time on this yet.
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] 0001-Allow-to-use-multiple-shared-memory-mapping-20250616.patch (30.0K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/2-0001-Allow-to-use-multiple-shared-memory-mapping-20250616.patch)
download | inline diff:
From 25a501f17a36523be0b133f992393433428d73c5 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH 01/17] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 ++++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 141 +++++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 13 ++-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
12 files changed, 278 insertions(+), 125 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 423b2b4f9d6..4ce2cfb662b 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -307,7 +307,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -328,7 +328,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..56af0231d24 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 567739b5be9..5b55bec8d9d 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2fa045e6b0f..8b38e985327 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -363,7 +368,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index c9ae3b45b76..7e1a9b43fae 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -76,19 +76,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -101,9 +101,17 @@ Datum pg_numa_available(PG_FUNCTION_ARGS);
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -114,7 +122,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -123,9 +137,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -150,11 +164,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -184,6 +204,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -203,22 +229,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -236,6 +262,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -246,19 +278,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -273,7 +305,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -334,6 +372,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -350,9 +400,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -385,6 +435,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -393,7 +450,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -416,7 +473,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -433,8 +490,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -459,7 +516,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -478,14 +535,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -542,10 +598,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -557,15 +614,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 46f44bc4511..a36b08895c8 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -618,10 +620,15 @@ LWLockNewTrancheId(void)
int *LWLockCounter;
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
- /* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ /*
+ * We use the ShmemLock spinlock to protect LWLockCounter.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
+ */
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..6ebda479ced 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 5f7d4b83a60..2348c59b5a0 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -91,4 +106,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index c1f668ded95..69663d412c3 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
base-commit: 3feff3916ee106c084eca848527dc2d2c3ef4e89
--
2.34.1
[text/x-patch] 0002-Address-space-reservation-for-shared-memory-20250616.patch (21.6K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/3-0002-Address-space-reservation-for-shared-memory-20250616.patch)
download | inline diff:
From acb2308031b68862a2f1238bf7ac803210063f71 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Wed, 16 Oct 2024 20:21:33 +0200
Subject: [PATCH 02/17] Address space reservation for shared memory
Currently the kernel is responsible to chose an address, where to place each
shared memory mapping, which is the lowest possible address that do not clash
with any other mappings. This is considered to be the most portable approach,
but one of the downsides is that there is no place to resize allocated mappings
anymore. Here is how it looks like for one mapping in /proc/$PID/maps,
/dev/zero represents the anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
...
7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
By specifying the mapping address directly it's possible to place the
mapping in a way that leaves room for resizing. The idea is:
* To reserve some address space via mmap'ing a large chunk of memory
with PROT_NONE and MAP_NORESERVE. This way we prepare a playground for
preparing shared memory layout without risking anything interfering
with that.
* To slice the reserved space up into sections, one to use for each
shared segment.
* Allocate shared memory segments out of corresponding slices and
leaving unclaimed space in between them. This is implemented via
mmap'ing memory at a specified address from the reserved space with
MAP_FIXED.
The result looks like this:
012d9000-0133e000 [heap]
7f443a800000-7f444196c000 /dev/zero (deleted)
7f444196c000-7f470a800000 # reserved space
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
Things like address space randomization should not be a problem in this
context, since the randomization is applied to the mmap base, which is
one per process.
This approach also do not impact the actual memory usage as reported by
the kernel. Here is the output of /proc/$PID/status for the master
version with shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages
// mapped in mm_struct. It corresponds to the mapped reserved space
// and is the only number that grows with it.
VmPeak: 2043192 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 22908 kB
// Size of resident anonymous memory
RssAnon: 768 kB
// Size of resident file mappings
RssFile: 10364 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 11776 kB
Here is the same for the patch when reserving 20GB of space:
VmPeak: 21250648 kB
VmRSS: 22948 kB
RssAnon: 768 kB
RssFile: 10404 kB
RssShmem: 11776 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
17637376 (~16.8 MB)
To control the amount of space reserved a new GUC max_available_memory
is introduced. Ideally it should be based on the maximum available
memory, hense the name.
---
src/backend/port/sysv_shmem.c | 284 ++++++++++++++++++++++++----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 5 +-
src/backend/utils/init/globals.c | 1 +
src/backend/utils/misc/guc_tables.c | 14 ++
src/include/storage/pg_shmem.h | 4 +-
6 files changed, 271 insertions(+), 39 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 56af0231d24..a0f03ff868f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -108,6 +108,66 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping placing (/dev/zero (deleted) below) looks like this:
+ *
+ * 00400000-00490000 /path/bin/postgres
+ * ...
+ * 012d9000-0133e000 [heap]
+ * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * ...
+ * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842
+ * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted)
+ * ...
+ *
+ * We would like to place multiple mappings in such a way, that there will be
+ * enough space between them in the address space to be able to resize up to
+ * certain size, but without counting towards the total memory consumption.
+ *
+ * To achieve that we first reserve some shared memory address space by
+ * mmap'ing a segment of MaxAvailableMemory size with PROT_NONE and
+ * MAP_NORESERVE (these flags allow to make sure this space will not be used by
+ * anything else, yet do not count against memory limits). Having the reserved
+ * space, we allocate out of it actual chunks of shared memory as usual,
+ * updating a pointer to the current available reserved space for the next
+ * allocation with the gap between segments in mind.
+ *
+ * The result would look like this:
+ *
+ * 012d9000-0133e000 [heap]
+ * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f442e010000-7f443a800000 # reserved empty space
+ * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f444196c000-7f470a800000 # reserved empty space
+ * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
+ * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
+ * [...]
+ *
+ * The reserved space pointer is calculated to slice up the total reserved
+ * space into fixed fractions of address space for each segment, as specified
+ * in the SHMEM_RESIZE_RATIO array.
+ */
+static double SHMEM_RESIZE_RATIO[1] = {
+ 1.0, /* MAIN_SHMEM_SLOT */
+};
+
+/*
+ * Offset from the beginning of the reserved space, which indicates currently
+ * available range. New shared memory segments have to be allocated at this
+ * offset related to the reserved space.
+ */
+static Size reserved_offset = 0;
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -626,39 +686,198 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
*
* This function will modify mapping size to the actual size of the allocation,
* if it ends up allocating a segment that is larger than requested.
+ *
+ * Note that we do not switch from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
*/
static void
-CreateAnonymousSegment(AnonymousMapping *mapping)
+CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
{
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS;
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* ReserveAnonymousMemory should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the request size to a suitable large value */
GetHugePageSize(&hugepagesize, &mmap_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
+ }
+#endif
+
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p",
+ MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+
+ /*
+ * Try to create mapping at an address out of the reserved range, which
+ * will allow to extend it later. Use reserved_offset to allocate the
+ * segment, then update currently available reserved range.
+ *
+ * If the last step has failed, fallback to the regular mapping
+ * creation and signal that shared buffers could not be resized without
+ * a restart.
+ */
+ ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_FIXED, -1, 0);
+ mmap_errno = errno;
+
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m, "
+ "fallback to the non-resizable allocation",
+ MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS, -1, 0);
+ mmap_errno = errno;
+ }
+ else
+ {
+ Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+
+ reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+ }
+
+ if (ptr == MAP_FAILED)
+ {
+ errno = mmap_errno;
+ DebugMappings();
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
+ (mmap_errno == ENOMEM) ?
+ errhint("This error usually means that PostgreSQL's request "
+ "for a shared memory segment exceeded available memory, "
+ "swap space, or huge pages. To reduce the request size "
+ "(currently %zu bytes), reduce PostgreSQL's shared "
+ "memory usage, perhaps by reducing \"shared_buffers\" or "
+ "\"max_connections\".",
+ allocsize) : 0));
+ }
+
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
+}
+
+/*
+ * ReserveAnonymousMemory
+ *
+ * Reserve shared memory address space, from which shared memory segments are
+ * going to be sliced out. The goal of this exercise is to support segments
+ * resizing, for which we need a reserved space free of potential clashes with
+ * other mmap'd areas that are not under our control. Reservation is done via
+ * mmap, and will not allocate any memory until it will be actually used, and
+ * MAP_NORESERVE allows to make it not counting againt kernel reservation
+ * limits (e.g. in cgroups or for huge pages). Do not get confused because of
+ * MAP_NORESERVE -- we need to reserve some space, but not the actual memory,
+ * and that is that this flag is about.
+ *
+ * Note, that with MAP_NORESERVE a reservation with hugetlb will succeed even
+ * if there is actually not enough huge pages. Hence this function is
+ * responsible for deciding whether to use huge pages or not. To achieve that
+ * we need to probe first and try to allocate needed memory for all segments --
+ * if this succeeds, we unmap the probe segment and use hugetlb; if it fails,
+ * we proceed with the regular memory.
+ */
+void *
+ReserveAnonymousMemory(Size reserve_size)
+{
+ Size allocsize = reserve_size;
+ void *ptr = MAP_FAILED;
+ int mmap_errno = 0;
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ *
+ * We could actually have a mix and match of segments with and without
+ * huge pages. But in that case we need to have multiple reservation
+ * spaces to use corresponding memory (hugetlb adress space reserved
+ * for hugetlb segments, regular memory for others), and it doesn't
+ * seem to worth the complexity for now.
+ */
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ int numSemas;
+ Size segment_size = CalculateShmemSize(&numSemas, segment);
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
{
- DebugMappings();
- elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- MappingName(mapping->shmem_segment), allocsize);
+ /* No huge pages, we will go with the regular page size */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB "
+ "failed, huge pages disabled: %m", total_size);
+ }
+ else
+ {
+ /*
+ * All fine, unmap the temporary segment and proceed with reserving
+ * using huge pages.
+ */
+ if (munmap(ptr, total_size) < 0)
+ elog(LOG, "reservice space: munmap(%p, %zu) failed: %m",
+ ptr, total_size);
+
+ /* Round up the requested size to a suitable large value. */
+ if (allocsize % hugepagesize != 0)
+ allocsize += hugepagesize - (allocsize % hugepagesize);
+
+ elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB",
+ allocsize);
+ ptr = mmap(NULL, allocsize, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | MAP_NORESERVE | mmap_flags,
+ -1, 0);
+ mmap_errno = errno;
+
+ /* This should not happen, but handle errors anyway */
+ if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
+ {
+ elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB "
+ "failed, huge pages disabled: %m", allocsize);
+ }
}
}
#endif
@@ -666,10 +885,12 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
/*
* Report whether huge pages are in use. This needs to be tracked before
* the second mmap() call if attempting to use huge pages failed
- * previously.
+ * previously. At this point ptr is either pointing to the probe segment,
+ * if we couldn't mmap it, or the reservation space.
*/
SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
{
@@ -677,10 +898,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
+ allocsize = reserve_size;
+
+ elog(DEBUG1, "reserving space: mmap(%zu)", allocsize);
+ ptr = mmap(NULL, allocsize, PROT_NONE,
+ MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0);
}
if (ptr == MAP_FAILED)
@@ -688,20 +910,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
errno = mmap_errno;
DebugMappings();
ereport(FATAL,
- (errmsg("segment[%s]: could not map anonymous shared memory: %m",
- MappingName(mapping->shmem_segment)),
+ (errmsg("reserving space: could not map anonymous shared "
+ "memory: %m"),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
- "for a shared memory segment exceeded available memory, "
- "swap space, or huge pages. To reduce the request size "
- "(currently %zu bytes), reduce PostgreSQL's shared "
- "memory usage, perhaps by reducing \"shared_buffers\" or "
- "\"max_connections\".",
+ "for a reserved shared memory address space exceeded "
+ "available memory, swap space, or huge pages. To "
+ "reduce the request reservation size (currently %zu "
+ "bytes), reduce PostgreSQL's \"maximum_shared_buffers\".",
allocsize) : 0));
}
- mapping->shmem = ptr;
- mapping->shmem_size = allocsize;
+ return ptr;
}
/*
@@ -740,7 +960,7 @@ AnonymousShmemDetach(int status, Datum arg)
*/
PGShmemHeader *
PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim)
+ PGShmemHeader **shim, Pointer base)
{
IpcMemoryKey NextShmemSegID;
void *memAddress;
@@ -760,14 +980,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -782,7 +994,7 @@ PGSharedMemoryCreate(Size size,
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
/* On success, mapping data will be modified. */
- CreateAnonymousSegment(mapping);
+ CreateAnonymousSegment(mapping, base);
next_free_segment++;
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..ce719f1b412 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -205,7 +205,7 @@ EnableLockPagesPrivilege(int elevel)
*/
PGShmemHeader *
PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim)
+ PGShmemHeader **shim, Pointer base)
{
void *memAddress;
PGShmemHeader *hdr;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8b38e985327..076888c0172 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -203,9 +203,12 @@ CreateSharedMemoryAndSemaphores(void)
PGShmemHeader *seghdr;
Size size;
int numSemas;
+ void *base;
Assert(!IsUnderPostmaster);
+ base = ReserveAnonymousMemory((Size) MaxAvailableMemory * BLCKSZ);
+
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
/* Compute the size of the shared-memory block */
@@ -217,7 +220,7 @@ CreateSharedMemoryAndSemaphores(void)
*
* XXX: Do multiple shims are needed, one per segment?
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(size, &shim, base);
/*
* Make sure that huge pages are never reported as "unknown" while the
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index d31cb45a058..9ccb7d455d6 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffers = 16384;
+int MaxAvailableMemory = 131072;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index f04bfedb2fd..e63521e5a2d 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2376,6 +2376,20 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"max_available_memory", PGC_SIGHUP, RESOURCES_MEM,
+ gettext_noop("Sets the upper limit for the shared_buffers value."),
+ gettext_noop("Shared memory could be resized at runtime, this "
+ "parameters sets the upper limit for it, beyond which "
+ "resizing would not be supported. Normally this value "
+ "would be the same as the total available memory."),
+ GUC_UNIT_BLOCKS
+ },
+ &MaxAvailableMemory,
+ 131072, 16, INT_MAX / 2,
+ NULL, NULL, NULL
+ },
+
{
{"vacuum_buffer_usage_limit", PGC_USERSET, RESOURCES_MEM,
gettext_noop("Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum."),
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 2348c59b5a0..8cb1e159917 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -61,6 +61,7 @@ extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT int MaxAvailableMemory;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -101,10 +102,11 @@ extern void PGSharedMemoryNoReAttach(void);
#endif
extern PGShmemHeader *PGSharedMemoryCreate(Size size,
- PGShmemHeader **shim);
+ PGShmemHeader **shim, Pointer base);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+void *ReserveAnonymousMemory(Size reserve_size);
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
--
2.34.1
[text/x-patch] 0004-Introduce-pending-flag-for-GUC-assign-hooks-20250616.patch (12.8K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/4-0004-Introduce-pending-flag-for-GUC-assign-hooks-20250616.patch)
download | inline diff:
From 1afa2048d803e0b3372de348868b26542ccfb3cd Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:40:32 +0200
Subject: [PATCH 04/17] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 61 insertions(+), 40 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 1914859b2ee..5e204341bde 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2321,7 +2321,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 608f10d9412..e40dae2ddf2 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,12 +1161,12 @@ assign_maintenance_io_concurrency(int newval, void *extra)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra)
+assign_io_max_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(newval, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra)
+assign_io_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(io_max_combine_limit, newval);
}
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index e5171467de1..2a6a587ef76 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1952,7 +1952,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1985,7 +1985,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2008,7 +2008,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2031,7 +2031,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 2f8c3d5f918..0d1b6466d1e 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3591,7 +3591,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 667df448732..bb681f5bc60 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f619100467d..8802ad8a3cb 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 799fa7ace68..c8300cffa8e 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,14 +81,14 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
-extern void assign_io_max_combine_limit(int newval, void *extra);
-extern void assign_io_combine_limit(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
+extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -143,13 +143,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -165,7 +165,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.34.1
[text/x-patch] 0005-Introduce-pss_barrierReceivedGeneration-20250616.patch (7.3K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/5-0005-Introduce-pss_barrierReceivedGeneration-20250616.patch)
download | inline diff:
From fb470c019742f2e9eaa7666ab81a24f816066387 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 4 Apr 2025 21:46:14 +0200
Subject: [PATCH 05/17] Introduce pss_barrierReceivedGeneration
Currently WaitForProcSignalBarrier allows to make sure the message sent
via EmitProcSignalBarrier was processed by all ProcSignal mechanism
participants.
Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
which will be updated when a process has received the message, but not
processed it yet. This makes it possible to support a new mode of
waiting, when ProcSignal participants want to synchronize message
processing. To do that, a participant can wait via
WaitForProcSignalBarrierReceived when processing a message, effectively
making sure that all processes are going to start processing
ProcSignalBarrier simultaneously.
---
src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------
src/include/storage/procsignal.h | 1 +
2 files changed, 54 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index a9bb540b55a..c6bec9be423 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -58,7 +58,10 @@
* of it. For such use cases, we set a bit in pss_barrierCheckMask and then
* increment the current "barrier generation"; when the new barrier generation
* (or greater) appears in the pss_barrierGeneration flag of every process,
- * we know that the message has been received everywhere.
+ * we know that the message has been received and processed everywhere. In case
+ * if we only need to know only that the message was received everywhere (e.g.
+ * receiving processes need to handle the message in a coordinated fashion)
+ * use pss_barrierReceivedGeneration in the same way.
*/
typedef struct
{
@@ -70,6 +73,7 @@ typedef struct
/* Barrier-related fields (not protected by pss_mutex) */
pg_atomic_uint64 pss_barrierGeneration;
+ pg_atomic_uint64 pss_barrierReceivedGeneration;
pg_atomic_uint32 pss_barrierCheckMask;
ConditionVariable pss_barrierCV;
} ProcSignalSlot;
@@ -152,6 +156,8 @@ ProcSignalShmemInit(void)
slot->pss_cancel_key_len = 0;
MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration,
+ PG_UINT64_MAX);
pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
ConditionVariableInit(&slot->pss_barrierCV);
}
@@ -199,6 +205,8 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
barrier_generation =
pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration);
pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration,
+ barrier_generation);
if (cancel_key_len > 0)
memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len);
@@ -263,6 +271,7 @@ CleanupProcSignalState(int status, Datum arg)
* no barrier waits block on it.
*/
pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX);
SpinLockRelease(&slot->pss_mutex);
@@ -416,12 +425,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type)
return generation;
}
-/*
- * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
- * requested by a specific call to EmitProcSignalBarrier() have taken effect.
- */
-void
-WaitForProcSignalBarrier(uint64 generation)
+static void
+WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly)
{
Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration));
@@ -436,12 +441,17 @@ WaitForProcSignalBarrier(uint64 generation)
uint64 oldval;
/*
- * It's important that we check only pss_barrierGeneration here and
- * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared
- * before the barrier is actually absorbed, but pss_barrierGeneration
+ * It's important that we check only pss_barrierGeneration &
+ * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in
+ * pss_barrierCheckMask get cleared before the barrier is actually
+ * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration
* is updated only afterward.
*/
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
while (oldval < generation)
{
if (ConditionVariableTimedSleep(&slot->pss_barrierCV,
@@ -450,7 +460,11 @@ WaitForProcSignalBarrier(uint64 generation)
ereport(LOG,
(errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier",
(int) pg_atomic_read_u32(&slot->pss_pid))));
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
}
ConditionVariableCancelSleep();
}
@@ -464,12 +478,33 @@ WaitForProcSignalBarrier(uint64 generation)
* The caller is probably calling this function because it wants to read
* the shared state or perform further writes to shared state once all
* backends are known to have absorbed the barrier. However, the read of
- * pss_barrierGeneration was performed unlocked; insert a memory barrier
- * to separate it from whatever follows.
+ * pss_barrierGeneration & pss_barrierReceivedGeneration was performed
+ * unlocked; insert a memory barrier to separate it from whatever follows.
*/
pg_memory_barrier();
}
+/*
+ * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
+ * requested by a specific call to EmitProcSignalBarrier() have taken effect.
+ */
+void
+WaitForProcSignalBarrier(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, false);
+}
+
+/*
+ * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all
+ * backends have observed the message sent by a specific call to
+ * EmitProcSignalBarrier().
+ */
+void
+WaitForProcSignalBarrierReceived(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, true);
+}
+
/*
* Handle receipt of an interrupt indicating a global barrier event.
*
@@ -523,6 +558,10 @@ ProcessProcSignalBarrier(void)
if (local_gen == shared_gen)
return;
+ /* The message is observed, record that */
+ pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
+ shared_gen);
+
/*
* Get and clear the flags that are set for this backend. Note that
* pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index afeeb1ca019..2733bbb8c5b 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, const uint8 *cancel_key, int cance
extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type);
extern void WaitForProcSignalBarrier(uint64 generation);
+extern void WaitForProcSignalBarrierReceived(uint64 generation);
extern void ProcessProcSignalBarrier(void);
extern void procsignal_sigusr1_handler(SIGNAL_ARGS);
--
2.34.1
[text/x-patch] 0003-Introduce-multiple-shmem-segments-for-share-20250616.patch (11.7K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/6-0003-Introduce-multiple-shmem-segments-for-share-20250616.patch)
download | inline diff:
From 92e4e639547207e883bc010c485177afa32d5c72 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 15 Mar 2025 16:38:59 +0100
Subject: [PATCH 03/17] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 24 +++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 105 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a0f03ff868f..f46d9d5d9cd 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -147,10 +147,18 @@ static int next_free_segment = 0;
*
* The reserved space pointer is calculated to slice up the total reserved
* space into fixed fractions of address space for each segment, as specified
- * in the SHMEM_RESIZE_RATIO array.
+ * in the SHMEM_RESIZE_RATIO array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take
+ * up to 60% of the whole space when resizing, based on the fact that it most
+ * likely will be the main consumer of this memory. Those numbers are pulled
+ * out of thin air for now, makes sense to evaluate them more precise.
*/
-static double SHMEM_RESIZE_RATIO[1] = {
- 1.0, /* MAIN_SHMEM_SLOT */
+static double SHMEM_RESIZE_RATIO[6] = {
+ 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.6, /* BUFFERS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -182,6 +190,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1dc488a42..bd68b69ee98 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +77,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -156,33 +160,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..a9952b36eba 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -22,6 +22,7 @@
#include "postgres.h"
#include "storage/buf_internals.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -59,10 +60,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 01909be0272..bd390f2709d 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -491,9 +492,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 076888c0172..9d00b80b4f8 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 41fdc1e7693..edac9db6a12 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -318,7 +318,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 8cb1e159917..f8459a5a421 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -108,7 +108,29 @@ extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
void *ReserveAnonymousMemory(Size reserve_size);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.34.1
[text/x-patch] 0008-Support-resize-for-hugetlb-20250616.patch (4.3K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/7-0008-Support-resize-for-hugetlb-20250616.patch)
download | inline diff:
From 2ebc737cd5b22c4cb3fbcafb583c0bbd61fe93d0 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 5 Apr 2025 19:51:33 +0200
Subject: [PATCH 08/17] Support resize for hugetlb
Linux kernel has a set of limitations on remapping hugetlb segments: it
can't increase size of such segment [1], and shrinking it will not
release the memory back. In fact support for hugetlb mremap was
implemented no so long time ago [2].
As a workaround, avoid mremap for resizing shared memory. Instead unmap
the whole segment and map it back at the same address with the new size,
relying on the fact that fd for the anon file behind the segment is
still open and will keep the memory content.
[1]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/mremap.c?id=f4d2ef48250ad057e4f00087967b5ff366da9f39#n1593
[2]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/mm/mremap.c?id=550a7d60bd5e35a56942dba6d8a26752beb26c9f
---
src/backend/port/sysv_shmem.c | 60 +++++++++++++++++++++++++----------
1 file changed, 44 insertions(+), 16 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 87000a24eea..f0b53ce1d7c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1109,6 +1109,7 @@ AnonymousShmemResize(void)
/* Note that CalculateShmemSize indirectly depends on NBuffers */
Size new_size = CalculateShmemSize(&numSemas, i);
AnonymousMapping *m = &Mappings[i];
+ int mmap_flags = PG_MMAP_FLAGS;
if (m->shmem == NULL)
continue;
@@ -1116,6 +1117,44 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+#ifndef MAP_HUGETLB
+ /* ReserveAnonymousMemory should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ Size hugepagesize;
+
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the new size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ if (new_size % hugepagesize != 0)
+ new_size += hugepagesize - (new_size % hugepagesize);
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
+ }
+#endif
+
+ /*
+ * Linux limitations do not allow us to mremap hugetlb in the way we
+ * want. E.g. no size increase is allowed, and for shrinking the memory
+ * will not be released back. To work around this unmap the segment and
+ * create a new one at the same address. Thanks for the backing anon
+ * file the content will still be kept in memory.
+ */
+ elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ if (munmap(m->shmem, m->shmem_size) < 0)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not unmap shared memory segment %s [%p]: %m",
+ MappingName(m->shmem_segment), m->shmem)));
+
/* Resize the backing anon file. */
if(ftruncate(m->segment_fd, new_size) == -1)
ereport(FATAL,
@@ -1123,25 +1162,14 @@ AnonymousShmemResize(void)
errmsg("could not truncase anonymous file for \"%s\": %m",
MappingName(m->shmem_segment))));
- /* Clean up some reserved space to resize into */
- if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
- ereport(FATAL,
- (errcode(ERRCODE_SYSTEM_ERROR),
- errmsg("could not unmap %zu from reserved shared memory %p: %m",
- new_size - m->shmem_size, m->shmem)));
-
- /* Claim the unused space */
- elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
- MappingName(m->shmem_segment), m->shmem_size,
- new_size, m->shmem);
-
- ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
+ /* Reclaim the space */
+ ptr = mmap(m->shmem, new_size, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_FIXED, m->segment_fd, 0);
if (ptr == MAP_FAILED)
ereport(FATAL,
(errcode(ERRCODE_SYSTEM_ERROR),
- errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m",
- MappingName(m->shmem_segment), m->shmem, NBuffers,
- new_size)));
+ errmsg("could not map shared memory segment %s [%p] with size %zu: %m",
+ MappingName(m->shmem_segment), m->shmem, new_size)));
reinit = true;
m->shmem_size = new_size;
--
2.34.1
[text/x-patch] 0006-Allow-to-resize-shared-memory-without-resta-20250616.patch (39.8K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/8-0006-Allow-to-resize-shared-memory-without-resta-20250616.patch)
download | inline diff:
From 12a39ceb67b438c2fdf3727b51f5e4a9c105b3b0 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:47:16 +0200
Subject: [PATCH 06/17] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers and extend it using mremap. One elected process signals the
postmaster to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f90cde00000-7f90d4fa6000 /dev/zero (deleted)
7f90d4fa6000-7f914de00000
7f914de00000-7f915cfa8000 /dev/zero (deleted)
^ buffers mapping, ~241 MB
7f915cfa8000-7f944de00000
7f944de00000-7f94550a8000 /dev/zero (deleted)
7f94550a8000-7f94cde00000
7f94cde00000-7f94d4fe8000 /dev/zero (deleted)
7f94d4fe8000-7f954de00000
7f954de00000-7f9554ff6000 /dev/zero (deleted)
7f9554ff6000-7f958de00000
7f958de00000-7f959508a000 /dev/zero (deleted)
7f959508a000-7f95cde00000
-- 512 MB
7f90cde00000-7f90d5126000 /dev/zero (deleted)
7f90d5126000-7f914de00000
7f914de00000-7f9175128000 /dev/zero (deleted)
^ buffers mapping, ~627 MB
7f9175128000-7f944de00000
7f944de00000-7f9455528000 /dev/zero (deleted)
7f9455528000-7f94cde00000
7f94cde00000-7f94d5228000 /dev/zero (deleted)
7f94d5228000-7f954de00000
7f954de00000-7f9555266000 /dev/zero (deleted)
7f9555266000-7f958de00000
7f958de00000-7f95954aa000 /dev/zero (deleted)
7f95954aa000-7f95cde00000
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it receives something.
Note, that mremap is Linux specific, thus the implementation not very
portable.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 413 ++++++++++++++++++
src/backend/postmaster/postmaster.c | 18 +
src/backend/storage/buffer/buf_init.c | 75 ++--
src/backend/storage/ipc/ipci.c | 18 +-
src/backend/storage/ipc/procsignal.c | 46 ++
src/backend/storage/ipc/shmem.c | 23 +-
src/backend/tcop/postgres.c | 10 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/miscadmin.h | 1 +
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 26 ++
src/include/storage/pmsignal.h | 1 +
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
17 files changed, 603 insertions(+), 43 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index f46d9d5d9cd..a3437973784 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -105,6 +111,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -176,6 +189,49 @@ static Size reserved_offset = 0;
*/
static bool huge_pages_on = false;
+/*
+ * Flag telling that we have prepared the memory layout to be resizable. If
+ * false after all shared memory segments creation, it means we failed to setup
+ * needed layout and falled back to the regular non-resizable approach.
+ */
+static bool shmem_resizable = false;
+
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -769,6 +825,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
{
Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+ shmem_resizable = true;
reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
}
@@ -964,6 +1021,315 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * mremap, which will fail if is not sufficient space to expand the mapping.
+ * When finished, based on the new and old values initialize new buffer blocks
+ * if any.
+ *
+ * If reinitializing took place, as the last step this function does buffers
+ * reinitialization as well and broadcasts the new value of NSharedBuffers. All
+ * of that needs to be done only by one backend, the first one that managed to
+ * grab the ShmemResizeLock.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ void *ptr = MAP_FAILED;
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ /* Clean up some reserved space to resize into */
+ if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not unmap %zu from reserved shared memory %p: %m",
+ new_size - m->shmem_size, m->shmem)));
+
+ /* Claim the unused space */
+ elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ ptr = mremap(m->shmem, m->shmem_size, new_size, 0);
+ if (ptr == MAP_FAILED)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m",
+ MappingName(m->shmem_segment), m->shmem, NBuffers,
+ new_size)));
+
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * Note that it's enough when only one backend will do that,
+ * even the ShmemInitStruct part. The reason is that resized
+ * shared memory will maintain the same addresses, meaning that
+ * all the pointers are still valid, and we only need to update
+ * structures size in the ShmemIndex once -- any other backend
+ * will pick up this shared structure from the index.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ BufferManagerShmemInit(NBuffersOld);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Wait for all ProcSignal participants
+ * to join the barrier, then do the resize and wait on the barrier until all
+ * participating finish resizing as well -- otherwise we face danger of
+ * inconsistency between backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * First thing to do after attaching to the barrier is to wait for others.
+ * We can't simply use BarrierArriveAndWait, because backends might arrive
+ * here in disjoint groups, e.g. first two backends, pause, then second two
+ * backends. If the resize is quick enough that can lead to a situation
+ * when the first group is already finished before the second has appeared,
+ * and the barrier will only synchonize withing those groups.
+ */
+ if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
+ WaitForProcSignalBarrierReceived(
+ pg_atomic_read_u64(&ShmemCtrl->Generation));
+
+ /*
+ * Now start the procedure, and elect one backend to ping postmaster to do
+ * the same.
+ *
+ * XXX: If we need to be able to abort resizing, this has to be done later,
+ * after the SHMEM_RESIZE_DONE.
+ */
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ {
+ Assert(IsUnderPostmaster);
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ }
+
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ BarrierDetach(barrier);
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ elog(DEBUG1, "Set pending signal");
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Start resizing procedure, making sure all existing processes will have
+ * consistent view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * We use dynamic barrier to help dealing with backends that were spawned
+ * during the resize.
+ */
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ *
+ * Those phases are ensured by joining the shared barrier associated with
+ * the procedure. Since resizing takes time, we need to take into account
+ * that during that time:
+ *
+ * - New backends can be spawned. They will check status of the barrier
+ * early during the bootstrap, and wait until everything is over to work
+ * with the new NBuffers value.
+ *
+ * - Old backends can exit before attempting to resize. Synchronization
+ * used between backends relies on ProcSignalBarrier and waits for all
+ * participants received the message at the beginning to gather all
+ * existing backends.
+ *
+ * - Some backends might be blocked and not responsing either before or
+ * after receiving the message. In the first case such backend still
+ * have ProcSignalSlot and should be waited for, in the second case
+ * shared barrier will make sure we still waiting for those backends. In
+ * any case there is an unbounded wait.
+ *
+ * - Backends might join barrier in disjoint groups with some time in
+ * between. That means that relying only on the shared dynamic barrier is
+ * not enough -- it will only synchronize resize procedure withing those
+ * groups. That's why we wait first for all participants of ProcSignal
+ * mechanism who received the message.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ pg_atomic_init_u64(&ShmemCtrl->Generation,
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
+
+ /* To order everything after setting Generation value */
+ pg_memory_barrier();
+
+ /*
+ * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
+ * the rest of the pack has started the procedure and it can resize shared
+ * memory as well.
+ *
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier. In fact even if we would like to
+ * wait here, it wouldn't be possible -- we're in the postmaster, without
+ * any waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1271,3 +1637,50 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier()
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ /* Nothing to do if resizing is not started */
+ if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
+ return;
+
+ BarrierAttach(barrier);
+
+ /* Otherwise wait through all available phases */
+ while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
+ {
+ ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
+ BarrierPhase(barrier))));
+
+ BarrierArriveAndWait(barrier, 0);
+ }
+
+ BarrierDetach(barrier);
+}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /*
+ * The barrier is missing here, it will be initialized right before
+ * starting the resizing process as a convenient way to reset it.
+ */
+
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+
+ /* shmem_resizable should be initialized by now */
+ ShmemCtrl->Resizable = shmem_resizable;
+ }
+}
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 490f7ce3664..f0cb0098dcd 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -426,6 +426,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1694,6 +1695,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2039,6 +2043,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
@@ -3852,6 +3867,9 @@ process_pm_pmsignal(void)
request_state_update = true;
}
+ if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
+ AnonymousShmemResize();
+
/*
* Try to advance postmaster's state machine, if a child requests it.
*/
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index bd68b69ee98..ac844b114bd 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -24,7 +25,6 @@ ConditionVariableMinimallyPadded *BufferIOCVArray;
WritebackContext BackendWritebackContext;
CkptSortItem *CkptBufferIds;
-
/*
* Data Structures:
* buffers live in a freelist and a lookup data structure.
@@ -62,18 +62,28 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * postmaster, or in a standalone backend) or during shared-memory resize. Size
+ * of data structures initialized here depends on NBuffers, and to be able to
+ * change NBuffers without a restart we store each structure into a separate
+ * shared memory segment, which could be resized on demand.
+ *
+ * FirstBufferToInit tells where to start initializing buffers. For
+ * initialization it always will be zero, but when resizing shared-memory it
+ * indicates the number of already initialized buffers.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(void)
+BufferManagerShmemInit(int FirstBufferToInit)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
+ elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
+ FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -110,43 +120,44 @@ BufferManagerShmemInit(void)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory, or in all cases when resizing shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
- ClearBufferTag(&buf->tag);
+#ifndef EXEC_BACKEND
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = FirstBufferToInit; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ ClearBufferTag(&buf->tag);
- buf->buf_id = i;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- pgaio_wref_clear(&buf->io_wref);
+ buf->buf_id = i;
- /*
- * Initially link all the buffers together as unused. Subsequent
- * management of this list is done by freelist.c.
- */
- buf->freeNext = i + 1;
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
+#endif
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 9d00b80b4f8..abeb91e24fd 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -84,6 +84,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -151,6 +154,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, AioShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -298,7 +309,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit();
+ BufferManagerShmemInit(0);
/*
* Set up lock manager
@@ -310,6 +321,11 @@ CreateOrAttachShmemStructs(void)
*/
PredicateLockShmemInit();
+ /*
+ * Set up shared memory resize manager
+ */
+ ShmemControlInit();
+
/*
* Set up process table
*/
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index c6bec9be423..d7b56a18b24 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -113,6 +114,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -176,6 +181,43 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -615,6 +657,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 7e1a9b43fae..c07572d6f89 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -498,17 +498,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
+ *
+ * XXX: There is an implicit assumption this can only happen in
+ * "resizable" segments, where only one shared structure is allowed.
+ * This has to be implemented more cleanly.
*/
if (result->size != size)
{
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
+ Size delta = size - result->size;
+
+ result->size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
}
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 0d1b6466d1e..0942d2bffe2 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4309,6 +4310,15 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /* Verify the shared barrier, if it's still active: join and wait. */
+ WaitOnShmemBarrier();
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 4da68312b5f..691fa14e9e3 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
@@ -352,6 +354,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index e63521e5a2d..9f00608f508 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2366,14 +2366,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 1bef98471c3..a0c37a7749e 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int MaxAvailableMemory;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index edac9db6a12..4239ebe640b 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -317,7 +317,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_skipped);
/* in buf_init.c */
-extern void BufferManagerShmemInit(void);
+extern void BufferManagerShmemInit(int);
extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6ebda479ced..bb7ae4d33b3 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index a9681738146..558da6fdd55 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
+PG_LWLOCK(54, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index f8459a5a421..19ad2e2f788 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,25 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+ pg_atomic_uint64 Generation;
+ bool Resizable;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -108,6 +128,12 @@ extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
void *ReserveAnonymousMemory(Size reserve_size);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 428aa3fd68a..1a55bf57a70 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,6 +42,7 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 2733bbb8c5b..97033f84dce 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index a8346cda633..b026a275c38 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2745,6 +2745,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.34.1
[text/x-patch] 0007-Use-anonymous-files-to-back-shared-memory-s-20250616.patch (10.7K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/9-0007-Use-anonymous-files-to-back-shared-memory-s-20250616.patch)
download | inline diff:
From 441f537b64b6bc8f0f00fa0de7850911acff621c Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sat, 15 Mar 2025 16:39:45 +0100
Subject: [PATCH 07/17] Use anonymous files to back shared memory segments
Allow to use anonymous files for shared memory, instead of plain
anonymous memory. Such an anonymous file is created via memfd_create, it
lives in memory, behaves like a regular file and semantically equivalent
to an anonymous memory allocated via mmap with MAP_ANONYMOUS.
Advantages of using anon files are following:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps. Here is how it looks like
7f90cde00000-7f90d5126000 rw-s 00000000 00:01 5463 /memfd:main (deleted)
7f90d5126000-7f914de00000 ---p 00000000 00:00 0
7f914de00000-7f9175128000 rw-s 00000000 00:01 5466 /memfd:buffers (deleted)
7f9175128000-7f944de00000 ---p 00000000 00:00 0
7f944de00000-7f9455528000 rw-s 00000000 00:01 5469 /memfd:descriptors (deleted)
7f9455528000-7f94cde00000 ---p 00000000 00:00 0
7f94cde00000-7f94d5228000 rw-s 00000000 00:01 5472 /memfd:iocv (deleted)
7f94d5228000-7f954de00000 ---p 00000000 00:00 0
7f954de00000-7f9555266000 rw-s 00000000 00:01 5475 /memfd:checkpoint (deleted)
7f9555266000-7f958de00000 ---p 00000000 00:00 0
7f958de00000-7f95954aa000 rw-s 00000000 00:01 5478 /memfd:strategy (deleted)
7f95954aa000-7f95cde00000 ---p 00000000 00:00 0
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 73 +++++++++++++++++++++++++++++-----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/portability/mem.h | 2 +-
src/include/storage/pg_shmem.h | 3 +-
5 files changed, 68 insertions(+), 14 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a3437973784..87000a24eea 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -107,6 +107,7 @@ typedef struct AnonymousMapping
Pointer shmem; /* Pointer to the start of the mapped memory */
Pointer seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -127,7 +128,7 @@ static int next_free_segment = 0;
* 00400000-00490000 /path/bin/postgres
* ...
* 012d9000-0133e000 [heap]
- * 7f443a800000-7f470a800000 /dev/zero (deleted)
+ * 7f443a800000-7f470a800000 /memfd:main (deleted)
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
* ...
@@ -150,9 +151,9 @@ static int next_free_segment = 0;
* The result would look like this:
*
* 012d9000-0133e000 [heap]
- * 7f4426f54000-7f442e010000 /dev/zero (deleted)
+ * 7f4426f54000-7f442e010000 /memfd:main (deleted)
* 7f442e010000-7f443a800000 # reserved empty space
- * 7f443a800000-7f444196c000 /dev/zero (deleted)
+ * 7f443a800000-7f444196c000 /memfd:buffers (deleted)
* 7f444196c000-7f470a800000 # reserved empty space
* 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
* 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2
@@ -643,13 +644,14 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* *hugepagesize and *mmap_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -708,6 +710,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -718,7 +721,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -727,6 +739,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -734,6 +748,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -771,7 +787,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
- int mmap_flags = PG_MMAP_FLAGS;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
#ifndef MAP_HUGETLB
/* ReserveAnonymousMemory should have dealt with this case */
@@ -785,7 +801,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
/* Round up the request size to a suitable large value */
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
@@ -794,6 +810,29 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
}
#endif
+ /*
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ */
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment),
+ memfd_flags);
+
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified size.
+ */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(mapping->shmem_segment))));
+
elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p",
MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
@@ -807,7 +846,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
* a restart.
*/
ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | MAP_FIXED, -1, 0);
+ mmap_flags | MAP_FIXED, mapping->segment_fd, 0);
mmap_errno = errno;
if (ptr == MAP_FAILED)
@@ -817,8 +856,15 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base)
"fallback to the non-resizable allocation",
MappingName(mapping->shmem_segment), allocsize, base + reserved_offset);
+ /* Specify the segment file size using allocsize. */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(mapping->shmem_segment))));
+
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
+ PG_MMAP_FLAGS, mapping->segment_fd, 0);
mmap_errno = errno;
}
else
@@ -889,7 +935,7 @@ ReserveAnonymousMemory(Size reserve_size)
Size hugepagesize, total_size = 0;
int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
/*
* Figure out how much memory is needed for all segments, keeping in
@@ -1070,6 +1116,13 @@ AnonymousShmemResize(void)
if (m->shmem_size == new_size)
continue;
+ /* Resize the backing anon file. */
+ if(ftruncate(m->segment_fd, new_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
/* Clean up some reserved space to resize into */
if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1)
ereport(FATAL,
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index ce719f1b412..ba972106de1 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index abeb91e24fd..dc2b4becf4a 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -396,7 +396,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 19ad2e2f788..192b637cc65 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -125,7 +125,8 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size,
PGShmemHeader **shim, Pointer base);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
void *ReserveAnonymousMemory(Size reserve_size);
bool ProcessBarrierShmemResize(Barrier *barrier);
--
2.34.1
[text/x-patch] 0010-Reinitialize-StrategyControl-after-resizing-20250616.patch (16.8K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/10-0010-Reinitialize-StrategyControl-after-resizing-20250616.patch)
download | inline diff:
From 2d173c9580b125dbc1248d86fb603a8508501ede Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 10 Jun 2025 11:00:36 +0530
Subject: [PATCH 10/17] Reinitialize StrategyControl after resizing buffers
... and BgBufferSync and ClockSweepTick adjustments
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. When expanding the buffer pool add new buffers to the free list.
2. When shrinking buffers, we remove any buffers, in the area being
shrunk, from the freelist. While doing so we adjust the first and
last free buffer pointers in the StrategyControl area. Hence nothing
more needed after resizing.
3. Check the sanity of the free buffer list is added after resizing.
4. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
5. &StrategyControl->buffer_strategy_lock need not be initialized again.
6. nextVictimBuffer, completePasses and numBufferAllocs are viewed in
the context of NBuffers. Now that NBuffers itself has changed, those
three do not make sense. Reset them as if the server has restarted
again.
This commit introduces a flag delay_shmem_resize, which postgresql
backends and workers can use to signal the coordinator to delay resizing
operation. Background writer sets this flag when its scanning buffers. Background
writer is blocked when the actual resizing is in progress. But if
resizing is about to begin, it does not scan the buffers by returning
from BgBufferSync(). It stops a scan in progress when it sees that the
resizing has begun. After the resizing is finished, it adjusts the
collected statistics according to the new size of the buffer pool at the
end of barrier processing.
Once the buffer resizing is finished, before resuming the regular
operation, bgwriter resets the information saved so far. This
information is viewed in the context of NBuffers and hence does not make
sense after NBuffer has changed.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 24 ++++-
src/backend/postmaster/bgwriter.c | 2 +-
src/backend/storage/buffer/buf_init.c | 11 ++-
src/backend/storage/buffer/bufmgr.c | 74 ++++++++++----
src/backend/storage/buffer/freelist.c | 133 ++++++++++++++++++++++++++
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 3 +-
src/include/storage/ipc.h | 1 +
8 files changed, 226 insertions(+), 23 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index c03464b40e2..9144f101585 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -114,6 +114,7 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Flag telling postmaster that resize is needed */
volatile bool pending_pm_shmem_resize = false;
+volatile bool delay_shmem_resize = false;
/* Keeps track of the previous NBuffers value */
static int NBuffersOld = -1;
@@ -1207,6 +1208,12 @@ AnonymousShmemResize(void)
LWLockRelease(ShmemResizeLock);
}
+
+ /*
+ * TODO: Shouldn't we call ResizeBufferPool() here as well? Or those
+ * backend who can not lock the LWLock conditionally won't resize the
+ * buffers.
+ */
}
return true;
@@ -1225,13 +1232,17 @@ AnonymousShmemResize(void)
bool
ProcessBarrierShmemResize(Barrier *barrier)
{
- elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
- NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize, delay_shmem_resize);
/* Wait until we have seen the new NBuffers value */
if (!pending_pm_shmem_resize)
return false;
+ /* Wait till this process becomes ready to resize buffers. */
+ if (delay_shmem_resize)
+ return false;
+
/*
* First thing to do after attaching to the barrier is to wait for others.
* We can't simply use BarrierArriveAndWait, because backends might arrive
@@ -1281,6 +1292,15 @@ ProcessBarrierShmemResize(Barrier *barrier)
/* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * Before resuming regular background writer activity, adjust the
+ * statistics collected so far.
+ */
+ BgBufferSyncAdjust(NBuffersOld, NBuffers);
+ }
+
BarrierDetach(barrier);
return true;
}
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index 72f5acceec7..32b34f28ead 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -233,7 +233,7 @@ BackgroundWriterMain(const void *startup_data, size_t startup_data_len)
/*
* Do one cycle of dirty-buffer writing.
*/
- can_hibernate = BgBufferSync(&wb_context);
+ can_hibernate = BgBufferSync(&wb_context, false);
/* Report pending statistics to the cumulative stats system */
pgstat_report_bgwriter();
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index f78be4700df..7b8bc577bd5 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -163,8 +163,15 @@ BufferManagerShmemInit(int FirstBufferToInit)
if (FirstBufferToInit < NBuffers)
GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
- /* Init other shared buffer-management stuff */
- StrategyInitialize(!foundDescs);
+ /*
+ * Init other shared buffer-management stuff from scratch configuring buffer
+ * pool the first time. If we are just resizing buffer pool adjust only the
+ * required structures.
+ */
+ if (FirstBufferToInit == 0)
+ StrategyInitialize(!foundDescs);
+ else
+ StrategyReInitialize(FirstBufferToInit);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 57d78c482bb..194d5da2999 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3611,6 +3611,32 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_DONE(NBuffers, num_written, num_to_scan);
}
+/*
+ * Information saved between BgBufferSync() calls so we can determine the
+ * strategy point's advance rate and avoid scanning already-cleaned buffers. The
+ * variables are global instead of static local so that BgBufferSyncAdjust() can
+ * adjust it when resizing shared buffers.
+ */
+static bool saved_info_valid = false;
+static int prev_strategy_buf_id;
+static uint32 prev_strategy_passes;
+static int next_to_clean;
+static uint32 next_passes;
+
+/* Moving averages of allocation rate and clean-buffer density */
+static float smoothed_alloc = 0;
+static float smoothed_density = 10.0;
+
+void
+BgBufferSyncAdjust(int NBuffersOld, int NBuffersNew)
+{
+ saved_info_valid = false;
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
+ NBuffersOld, NBuffersNew);
+#endif
+}
+
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
@@ -3623,27 +3649,13 @@ BufferSync(int flags)
* bgwriter_lru_maxpages to 0.)
*/
bool
-BgBufferSync(WritebackContext *wb_context)
+BgBufferSync(WritebackContext *wb_context, bool reset)
{
/* info obtained from freelist.c */
int strategy_buf_id;
uint32 strategy_passes;
uint32 recent_alloc;
- /*
- * Information saved between calls so we can determine the strategy
- * point's advance rate and avoid scanning already-cleaned buffers.
- */
- static bool saved_info_valid = false;
- static int prev_strategy_buf_id;
- static uint32 prev_strategy_passes;
- static int next_to_clean;
- static uint32 next_passes;
-
- /* Moving averages of allocation rate and clean-buffer density */
- static float smoothed_alloc = 0;
- static float smoothed_density = 10.0;
-
/* Potentially these could be tunables, but for now, not */
float smoothing_samples = 16;
float scan_whole_pool_milliseconds = 120000.0;
@@ -3666,6 +3678,22 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ /*
+ * If buffer pool is being shrunk the buffer being written out may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing out any. Hence wait till
+ * buffer resizing finishes i.e. go into hibernation mode.
+ */
+ if (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ return true;
+
+ /*
+ * Resizing shared buffers while this function is performing an LRU scan on
+ * them may lead to wrong results. Indicate that the resizing should wait for
+ * the LRU scan to complete.
+ */
+ delay_shmem_resize = true;
+
/*
* Find out where the freelist clock sweep currently is, and how many
* buffer allocations have happened since our last call.
@@ -3842,8 +3870,17 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
- while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
+ /*
+ * Execute the LRU scan.
+ *
+ * If buffer pool is being shrunk, the buffer being written may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing any. Hence stop what we are doing. This
+ * also unblocks other processes that are waiting for buffer resizing to
+ * finish.
+ */
+ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
+ pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) == NBuffers)
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -3902,6 +3939,9 @@ BgBufferSync(WritebackContext *wb_context)
#endif
}
+ /* Let the resizing commence. */
+ delay_shmem_resize = false;
+
/* Return true if OK to hibernate */
return (bufs_to_lap == 0 && recent_alloc == 0);
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 7b9ed010e2f..41641bb3ae6 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -98,6 +98,9 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
uint32 *buf_state);
static void AddBufferToRing(BufferAccessStrategy strategy,
BufferDesc *buf);
+#ifdef USE_ASSERT_CHECKING
+static void StrategyValidateFreeList(void);
+#endif /* USE_ASSERT_CHECKING */
/*
* ClockSweepTick - Helper routine for StrategyGetBuffer()
@@ -526,6 +529,88 @@ StrategyInitialize(bool init)
Assert(!init);
}
+/*
+ * StrategyReInitialize -- re-initialize the buffer cache replacement
+ * strategy.
+ *
+ * To be called when resizing buffer manager and only from the coordinator.
+ * TODO: Assess the differences between this function and StrategyInitialize().
+ */
+void
+StrategyReInitialize(int FirstBufferIdToInit)
+{
+ bool found;
+
+ /*
+ * Resizing memory for buffer pools should not affect the address of
+ * StrategyControl.
+ */
+ if (StrategyControl != (BufferStrategyControl *)
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, STRATEGY_SHMEM_SEGMENT))
+ elog(FATAL, "something went wrong while re-initializing the buffer strategy");
+
+ Assert(found);
+
+ /* TODO: Buffer lookup table adjustment: There are two options:
+ *
+ * 1. Resize the buffer lookup table to match the new number of buffers. But
+ * this requires rehashing all the entries in the buffer lookup table with
+ * the new table size.
+ *
+ * 2. Allocate maximum size of the buffer lookup table at the beginning and
+ * never resize it. This leaves sparse buffer lookup table which is
+ * inefficient from both memory and time perspective. According to David
+ * Rowley, the sparse entries in the buffer look up table cause frequent
+ * cacheline reload which affect performance. If the impact of that
+ * inefficiency in a benchmark is significant, we will need to consider first
+ * option.
+ */
+
+ /*
+ * When shrinking buffers, we must have adjusted the first and the last free
+ * buffer when removing the buffers being shrunk from the free list. Nothing
+ * to be done here.
+ *
+ * When expanding the shared buffers, new buffers are added at the end of the
+ * freelist or they form the new free list if there are no free buffers.
+ */
+ if (FirstBufferIdToInit < NBuffers)
+ {
+ if (StrategyControl->firstFreeBuffer == FREENEXT_END_OF_LIST)
+ StrategyControl->firstFreeBuffer = FirstBufferIdToInit;
+ else
+ {
+ Assert(StrategyControl->lastFreeBuffer >= 0);
+ GetBufferDescriptor(StrategyControl->lastFreeBuffer - 1)->freeNext = FirstBufferIdToInit;
+ }
+
+ StrategyControl->lastFreeBuffer = NBuffers - 1;
+ }
+
+ /* Check free list sanity after resizing. */
+#ifdef USE_ASSERT_CHECKING
+ StrategyValidateFreeList();
+#endif /* USE_ASSERT_CHECKING */
+
+ /*
+ * The clock sweep tick pointer might have got invalidated. Reset it as if
+ * starting a fresh server.
+ */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The old statistics is viewed in the context of the number of shared
+ * buffers. It does not make sense now that the number of shared buffers
+ * itself has changed.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_init_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* No pending notification */
+ StrategyControl->bgwprocno = -1;
+}
/*
* StrategyPurgeFreeList -- remove all buffers with id higher than the number of
@@ -595,6 +680,54 @@ StrategyPurgeFreeList(int numBuffers)
*/
}
+#ifdef USE_ASSERT_CHECKING
+/*
+ * StrategyValidateFreeList-- check sanity of free buffer list.
+ */
+static void
+StrategyValidateFreeList(void)
+{
+ int nextFree = StrategyControl->firstFreeBuffer;
+ int numFreeBuffers = 0;
+ int lastFreeBuffer = FREENEXT_END_OF_LIST;
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ while (nextFree != FREENEXT_END_OF_LIST)
+ {
+ BufferDesc *buf = GetBufferDescriptor(nextFree);
+
+ /* nextFree should be id of buffer being examined. */
+ Assert(nextFree == buf->buf_id);
+ Assert(buf->buf_id < NBuffers);
+ /* The buffer should not be marked as not in the list. */
+ Assert(buf->freeNext != FREENEXT_NOT_IN_LIST);
+
+ /* Update our knowledge of last buffer in the free list. */
+ lastFreeBuffer = buf->buf_id;
+
+ numFreeBuffers++;
+
+ /* Avoid infinite recursion in case there are cycles in free list. */
+ if (numFreeBuffers > NBuffers)
+ break;
+
+ nextFree = buf->freeNext;
+ }
+
+ Assert(numFreeBuffers <= NBuffers);
+
+ /*
+ * Make sure that the StrategyControl's knowledge of last free buffer
+ * agrees with what's there in the free list.
+ */
+ if (StrategyControl->firstFreeBuffer != FREENEXT_END_OF_LIST)
+ Assert(StrategyControl->lastFreeBuffer == lastFreeBuffer);
+
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+}
+#endif /* USE_ASSERT_CHECKING */
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
* ----------------------------------------------------------------
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index add15e3723b..46949e9d90e 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -454,6 +454,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
extern void StrategyPurgeFreeList(int numBuffers);
+extern void StrategyReInitialize(int FirstBufferToInit);
extern bool have_free_buffer(void);
/* buf_table.c */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 0c554f0b130..83a75eab844 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -298,7 +298,8 @@ extern bool ConditionalLockBufferForCleanup(Buffer buffer);
extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
-extern bool BgBufferSync(struct WritebackContext *wb_context);
+extern bool BgBufferSync(struct WritebackContext *wb_context, bool reset);
+extern void BgBufferSyncAdjust(int NBuffersOld, int NBuffersNew);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index bb7ae4d33b3..7d1c64a9267 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -65,6 +65,7 @@ typedef void (*shmem_startup_hook_type) (void);
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
+extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
--
2.34.1
[text/x-patch] 0009-Support-shrinking-shared-buffers-20250616.patch (13.4K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/11-0009-Support-shrinking-shared-buffers-20250616.patch)
download | inline diff:
From 23c8b2f5d75f52f14c80015ed104e82f69eca5c6 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 9 Jun 2025 14:40:34 +0530
Subject: [PATCH 09/17] Support shrinking shared buffers
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
Buffer eviction requires a separate barrier phase for two reasons:
1. No other backend should map a new page to any of buffers being
evicted when eviction is in progress. So they wait while eviction is
in progress.
2. Since a pinned buffer has the pin recorded in the backend local
memory as well as the buffer descriptor (which is in shared memory),
eviction should not coincide with remapping the shared memory of a
backend. Otherwise we might loose consistency of local and shared
pinning records. Hence it needs to be carried out in
ProcessBarrierShmemResize() and not in AnonymousShmemResize() as
indicated by now removed comment.
If a buffer being evicted is pinned, we raise a FATAL error but this should
improve. There are multiple options 1. to wait for the pinned buffer to get
unpinned, 2. the backend is killed or it itself cancels the query or 3.
rollback the operation. Note that option 1 and 2 would require the pinning
related local and shared records to be accessed. But we need infrastructure to
do either of this right now.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 35 +++++---
src/backend/storage/buffer/buf_init.c | 8 +-
src/backend/storage/buffer/bufmgr.c | 89 +++++++++++++++++++
src/backend/storage/buffer/freelist.c | 68 ++++++++++++++
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 1 +
src/include/storage/pg_shmem.h | 1 +
8 files changed, 192 insertions(+), 12 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index f0b53ce1d7c..c03464b40e2 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1096,14 +1096,6 @@ AnonymousShmemResize(void)
*/
pending_pm_shmem_resize = false;
- /*
- * XXX: Currently only increasing of shared_buffers is supported. For
- * decreasing something similar has to be done, but buffer blocks with
- * data have to be drained first.
- */
- if(NBuffersOld > NBuffers)
- return false;
-
for(int i = 0; i < next_free_segment; i++)
{
/* Note that CalculateShmemSize indirectly depends on NBuffers */
@@ -1201,11 +1193,14 @@ AnonymousShmemResize(void)
* all the pointers are still valid, and we only need to update
* structures size in the ShmemIndex once -- any other backend
* will pick up this shared structure from the index.
- *
- * XXX: This is the right place for buffer eviction as well.
*/
BufferManagerShmemInit(NBuffersOld);
+ /*
+ * Wipe out the evictor PID so that it can be used for the next
+ * buffer resizing operation.
+ */
+ ShmemCtrl->evictor_pid = 0;
/* If all fine, broadcast the new value */
pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
}
@@ -1262,6 +1257,25 @@ ProcessBarrierShmemResize(Barrier *barrier)
SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
}
+ /*
+ * Evict extra buffers when shrinking shared buffers. We need to do this
+ * while the memory for extra buffers is still mapped i.e. before remapping
+ * the shared memory segments to a smaller memory area.
+ */
+ if (NBuffersOld > NBuffersPending)
+ {
+ /*
+ * TODO: If the buffer eviction fails for any reason, we should
+ * gracefully rollback the shared buffer resizing and try again. But the
+ * infrastructure to do so is not available right now. Hence just raise
+ * a FATAL so that the system restarts.
+ */
+ if (!EvictExtraBuffers(NBuffersPending, NBuffersOld))
+ elog(FATAL, "buffer eviction failed");
+
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_EVICT);
+ }
+
AnonymousShmemResize();
/* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
@@ -1763,5 +1777,6 @@ ShmemControlInit(void)
/* shmem_resizable should be initialized by now */
ShmemCtrl->Resizable = shmem_resizable;
+ ShmemCtrl->evictor_pid = 0;
}
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ac844b114bd..f78be4700df 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -156,8 +156,12 @@ BufferManagerShmemInit(int FirstBufferToInit)
}
#endif
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ /*
+ * Correct last entry of linked list, when initializing the buffers or when
+ * expanding the buffers.
+ */
+ if (FirstBufferToInit < NBuffers)
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 667aa0c0c78..57d78c482bb 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/read_stream.h"
#include "storage/smgr.h"
@@ -7453,3 +7454,91 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int newBufSize, int oldBufSize)
+{
+ bool result = true;
+
+ /*
+ * If the buffer being evicated is locked, this function will need to wait.
+ * This function should not be called from a Postmaster since it can not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * Let only one backend perform eviction. We could split the work across all
+ * the backends but that doesn't seem necessary.
+ *
+ * The first backend to acquire ShmemResizeLock, sets its own PID as the
+ * evictor PID for other backends to know that the eviction is in progress or
+ * has already been performed. The evictor backend releases the lock when it
+ * finishes eviction. While the eviction is in progress, backends other than
+ * evictor backend won't be able to take the lock. They won't perform
+ * eviction. A backend may acquire the lock after eviction has completed, but
+ * it will not perform eviction since the evictor PID is already set. Evictor
+ * PID is reset only when the buffer resizing finishes. Thus only one backend
+ * will perform eviction in a given instance of shared buffers resizing.
+ *
+ * Any backend which acquires this lock will release it before the eviction
+ * phase finishes, hence the same lock can be reused for the next phase of
+ * resizing buffers.
+ */
+ if (LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ if (ShmemCtrl->evictor_pid == 0)
+ {
+ ShmemCtrl->evictor_pid = MyProcPid;
+
+ StrategyPurgeFreeList(newBufSize);
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = newBufSize + 1; buf <= oldBufSize; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u32(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is
+ * going one hence unlocked precheck should be safe and saves
+ * some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+ }
+ LWLockRelease(ShmemResizeLock);
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index bd390f2709d..7b9ed010e2f 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -527,6 +527,74 @@ StrategyInitialize(bool init)
}
+/*
+ * StrategyPurgeFreeList -- remove all buffers with id higher than the number of
+ * buffers in the buffer pool.
+ *
+ * This is called before evicting buffers while shrinking shared buffers, so that
+ * the free list does not reference a buffer that will be removed.
+ *
+ * The function is called after resizing has started and thus nobody should be
+ * traversing the free list and also not touching the buffers.
+ */
+void
+StrategyPurgeFreeList(int numBuffers)
+{
+ int firstBuffer = FREENEXT_END_OF_LIST;
+ int nextFree = StrategyControl->firstFreeBuffer;
+ BufferDesc *prevValidBuf = NULL;
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ while (nextFree != FREENEXT_END_OF_LIST)
+ {
+ BufferDesc *buf = GetBufferDescriptor(nextFree);
+
+ /* nextFree should be id of buffer being examined. */
+ Assert(nextFree == buf->buf_id);
+ /* The buffer should not be marked as not in the list. */
+ Assert(buf->freeNext != FREENEXT_NOT_IN_LIST);
+
+ /*
+ * If the buffer is within the new size of pool, keep it in the free list
+ * otherwise discard it.
+ */
+ if (buf->buf_id < numBuffers)
+ {
+ if (prevValidBuf != NULL)
+ prevValidBuf->freeNext = buf->buf_id;
+ prevValidBuf = buf;
+
+ /* Save the first free buffer in the list if not already known. */
+ if (firstBuffer == FREENEXT_NOT_IN_LIST)
+ firstBuffer = nextFree;
+ }
+ /* Examine the next buffer in the free list. */
+ nextFree = buf->freeNext;
+ }
+
+ /* Update the last valid free buffer, if there's any. */
+ if (prevValidBuf != NULL)
+ {
+ StrategyControl->lastFreeBuffer = prevValidBuf->buf_id;
+ prevValidBuf->freeNext = FREENEXT_END_OF_LIST;
+ }
+ else
+ StrategyControl->lastFreeBuffer = FREENEXT_END_OF_LIST;
+
+ /* Update first valid free buffer, if there's any. */
+ StrategyControl->firstFreeBuffer = firstBuffer;
+
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+
+ /*
+ * TODO: following was suggested by AI. Check whether it is required.
+ * If we removed all buffers from the freelist, reset the clock sweep
+ * pointer to zero. This is not strictly necessary, but it seems like a
+ * good idea to avoid confusion.
+ */
+}
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
* ----------------------------------------------------------------
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 691fa14e9e3..0c588b69a90 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -156,6 +156,7 @@ REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it c
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_EVICT "Waiting for other backends to finish buffer evication phase."
SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 0dec7d93b3b..add15e3723b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -453,6 +453,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyPurgeFreeList(int numBuffers);
extern bool have_free_buffer(void);
/* buf_table.c */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 4239ebe640b..0c554f0b130 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -315,6 +315,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_evicted,
int32 *buffers_flushed,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int fromBuf, int toBuf);
/* in buf_init.c */
extern void BufferManagerShmemInit(int);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 192b637cc65..23998f5469d 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -64,6 +64,7 @@ extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
typedef struct
{
pg_atomic_uint32 NSharedBuffers;
+ pid_t evictor_pid;
Barrier Barrier;
pg_atomic_uint64 Generation;
bool Resizable;
--
2.34.1
[text/x-patch] 0012-Fix-compilation-failure-in-pg_get_shmem_all-20250616.patch (1.4K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/12-0012-Fix-compilation-failure-in-pg_get_shmem_all-20250616.patch)
download | inline diff:
From aa1ef8f2a108764585f42073e2b770689e3e8b3b Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 5 Jun 2025 11:34:12 +0530
Subject: [PATCH 12/17] Fix compilation failure in
pg_get_shmem_allocations_numa()
The compilation failure is caused by
5cefa489760e34d947dbe67b4a922468b2e43668. Ideal fix should be compute
the total page count across all the shared memory segments. This commit
just fixes the compilation failure. # Please enter the commit message
for your changes. Lines starting
Ashutosh Bapat
---
src/backend/storage/ipc/shmem.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index c07572d6f89..b411fbce37e 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -696,7 +696,12 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ /*
+ * TODO: We should loop through all the Shm segments, instead of just the
+ * main segment, to find the total page count.
+ */
+ shm_total_page_count = (Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize
+ / os_page_size) + 1;
page_ptrs = palloc0(sizeof(void *) * shm_total_page_count);
pages_status = palloc(sizeof(int) * shm_total_page_count);
--
2.34.1
[text/x-patch] 0011-Additional-validation-for-buffer-in-the-rin-20250616.patch (2.1K, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/13-0011-Additional-validation-for-buffer-in-the-rin-20250616.patch)
download | inline diff:
From ac07b39c724fe5a0d4350b7b716c7690d22d4e82 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 11 Jun 2025 18:15:06 +0530
Subject: [PATCH 11/17] Additional validation for buffer in the ring
If the buffer pool has been shrunk, the buffers in the buffer list may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition. That may not be
expensive enough to affect query performance. The alternative to that is
more complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure thus accessible when processing buffer
resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work.
Ashutosh Bapat
---
src/backend/storage/buffer/freelist.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 41641bb3ae6..74d070733a4 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -948,12 +948,13 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot has been
+ * invalidated when buffer pool was shrunk, tell the caller to allocate a new
+ * buffer with the normal allocation strategy. He will then fill this slot
+ * by calling AddBufferToRing with the new buffer.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer || bufnum > NBuffers)
return NULL;
/*
--
2.34.1
[text/x-patch] 0013-Fix-compilation-failure-in-pg_get_shmem_pag-20250616.patch (959B, ../../CAExHW5sYg_d4O7oGRqbomnVODeqR3YNAeYAa526n1dsWCM=+Fg@mail.gmail.com/14-0013-Fix-compilation-failure-in-pg_get_shmem_pag-20250616.patch)
download | inline diff:
From cc706b4db87b3bd69ee2f4acc223e9306fb10674 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 5 Jun 2025 14:42:53 +0530
Subject: [PATCH 13/17] Fix compilation failure in pg_get_shmem_pagesize()
Fix compilation failure in pg_get_shmem_pagesize() due to incorrect call to
GetHugePageSize(). This is a temporary fix to allow compilation to proceed.
Ashutosh Bapat
---
src/backend/storage/ipc/shmem.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index b411fbce37e..4c2bddfe6ca 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -826,7 +826,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-06-20 10:19 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-06-20 10:22 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 2 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-06-20 10:19 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi,
> On Mon, Apr 21, 2025 at 7:47 PM Thomas Munro <thomas.munro@gmail.com>
> wrote:
>
> One thing I'm still wondering about is whether you really need all
> this multi-phase barrier stuff, or even need to stop other backends
> from running at all while doing the resize. I guess that's related to
> your remapping scheme, but supposing you find the simple
> ftruncate()-only approach to be good, my next question is: why isn't
> it enough to wait for all backends to agree to stop allocating new
> buffers in the range to be truncated, and then left them continue to
> run as normal? As far as they would be concerned, the in-progress
> downsize has already happened, though it could be reverted later if
> the eviction phase fails. Then the coordinator could start evicting
> buffers and truncating the shared memory object, which are
> phases/steps, sure, but it's not clear to me why they need other
> backends' help.
My intention behind keeping all backends waiting was to have a simple way of
not only preventing them from allocating new buffers from the truncated range,
but also eliminating any chance of them accessing those to-be-truncated
buffers. In the end it's just easier (at least for me) to reason about
correctness of the implementation this way.
> On Tue, Jun 10, 2025 at 04:39:58PM +0530, Ashutosh Bapat wrote:
>
> Here's patchset rebased on f85f6ab051b7cf6950247e5fa6072c4130613555
Thanks! I've reworked the series to implement approach suggested by
Thomas, and applied your patches to support buffers shrinking on top. I
had to restructure the patch set, here is how it looks like right now:
1. Preparation patches
Changes, that are needed to support resizing functionality, but not
strictly related to it.
* Process config reload in AIO workers. Corrects omission discussed on [1].
* Introduce pending flag for GUC assign hooks. Allowing to decouple a GUC value
change from actually applying it, sort of "pending" change. The idea is to let
a custom logic be triggered on an assign hook, and then take responsibility for
what happens later and how it's going to be applied. Doesn't do GUC reporting
yet.
* Introduce pss_barrierReceivedGeneration. Allows to distinguish situations
when a signal was processed everywhere, and when a signal was received
everywhere.
2. Resizing implementation
* Allow to use multiple shared memory mappings. A preparation patch, extending
the existing interface to support multiple shared memory segments.
* Address space reservation for shared memory. Implement the new way of
handling shared memory segments, now each segment can visually be represented
as following:
/ Address space \
+---------------<+>--------------------------+
| Actual content | Address space reservation |
| (memfd) | (mmap, PROT_NONE) |
+---------------<+>--------------------------+
The actual segment size is managed via ftruncate and mprotect. One interesting
side effect I haven't fully understood yet, is that Linux doesn't seem to
extend the existing mapping when doing mprotect on huge pages, it creates
another mapping instead. E.g. when using normal page size and resizing shared
memory we get:
7f4808600000-7f4817e00000 rw-s /memfd:buffers (deleted)
7f4817e00000-7f48a2000000 ---s /memfd:buffers (deleted)
Doing the same with huge pages ends up looking like this:
7f4808600000-7f4817e00000 rw-s /memfd:buffers (deleted)
7f4817e00000-7f4830000000 rw-s /memfd:buffers (deleted)
7f4830000000-7f48a2000000 ---s /memfd:buffers (deleted)
I'm still investigating whether it's a mistake on my side or a genuine Linux
behavior. At the same time I don't see it as a large issue, the same situation
could happen with the previous implementation as well.
* Introduce multiple shmem segments for shared buffers. Modifies necessary bits
to use new functionality.
* Allow to resize shared memory without restart. Utilizes infrastructure
introduced so far to implement stop-the-world resizing approach, where all the
active backend (and potentially new one spawning) are waiting until everyone
gets the same shared memory size.
When testing I've noticed that there seems to be concurrency issues with
interrupts, where aio workers and checkpointer sometimes do not receive
the resize signal correctly. I assume it has something to do with the
significant behavior change -- config reload processing can now fire
signals on its own. Letting those backends to always process config
reload first seems to be resolved (or at least hide) the issue, but I
still need to understand what's going on there.
3. Shared memory shrinking
So far only shared memory increase was implemented. These patches from Ashutosh
support shrinking as well, which is tricky due to the need for buffer eviction.
* Support shrinking shared buffers
* Reinitialize StrategyControl after resizing buffers
* Additional validation for buffer in the ring
> 0009 adds support to shrink shared buffers. It has two changes: a.
> evict the buffers outside the new buffer size b. remove buffers with
> buffer id outside the new buffer size from the free list. If a buffer
> being evicted is pinned, the operation is aborted and a FATAL error is
> raised. I think we need to change this behaviour to be less severe
> like rolling back the operation or waiting for the pinned buffer to be
> unpinned etc. Better even if we could let users control the behaviour.
> But we need better infrastructure to do such things. That's one TODO
> left in the patch.
I haven't reviewed those, just tested a bit to finally include into the series.
Note that I had to tweak two things:
* The way it was originally implemented was sending resize signal to postmaster
before doing eviction, which could result in sigbus when accessing LSN of a
dirty buffer to be evicted. I've reshaped it a bit to make sure eviction always
happens first.
* It seems the CurrentResource owner could be missing sometimes, so I've added
a band-aid checking its presence.
One side note, during my testing I've noticed assert failures on
pgstat_tracks_io_op inside a wal writer a few times. I couldn't reproduce it
after the fixes above, but still it may indicate that something is off. E.g.
it's somehow not expected that the wal writer will do buffer eviction IO (from
what I understand, the current shrinking implementation allows that).
> Buffer lookup table resizing
> ------------------------------------
> The size of the buffer lookup table depends upon (number of shared
> buffers + number of partitions in the shared buffer lookup table). If we
> shrink the buffer pool, the buffer lookup table will become sparse but
> still useful. If we expand the buffers we need to expand the buffer lookup
> table too. That's not implemented in the current patchset.
Just FYI, buffer lookup table has its own STRATEGY_SHMEM_SEGMENT shared memory
segment and is resized in the same way as others. There could be lots of
details missing, but at least the corresponding resizable segment is already
there.
[1]: https://www.postgresql.org/message-id/flat/sh5uqe4a4aqo5zkkpfy5fobe2rg2zzouctdjz7kou4t74c66ql%40yzpk...
From 5d4f46dc13bacf9c1df79233f8e4017a9dcd6919 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 15:14:33 +0200
Subject: [PATCH v5 01/10] Process config reload in AIO workers
Currenly AIO workers process interrupts only via CHECK_FOR_INTERRUPTS,
which does not include ConfigReloadPending. Thus we need to check for it
explicitly.
---
src/backend/storage/aio/method_worker.c | 25 +++++++++++++++++++++++++
1 file changed, 25 insertions(+)
diff --git a/src/backend/storage/aio/method_worker.c b/src/backend/storage/aio/method_worker.c
index 36be179678d..b4d5c46fb94 100644
--- a/src/backend/storage/aio/method_worker.c
+++ b/src/backend/storage/aio/method_worker.c
@@ -80,6 +80,7 @@ static void pgaio_worker_shmem_init(bool first_time);
static bool pgaio_worker_needs_synchronous_execution(PgAioHandle *ioh);
static int pgaio_worker_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+static void pgaio_worker_process_interrupts(void);
const IoMethodOps pgaio_worker_ops = {
.shmem_size = pgaio_worker_shmem_size,
@@ -461,6 +462,8 @@ IoWorkerMain(const void *startup_data, size_t startup_data_len)
int nwakeups = 0;
int worker;
+ pgaio_worker_process_interrupts();
+
/*
* Try to get a job to do.
*
@@ -584,3 +587,25 @@ pgaio_workers_enabled(void)
{
return io_method == IOMETHOD_WORKER;
}
+
+/*
+ * Process any new interrupts.
+ */
+static void
+pgaio_worker_process_interrupts(void)
+{
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
+ if (ConfigReloadPending)
+ {
+ ConfigReloadPending = false;
+ ProcessConfigFile(PGC_SIGHUP);
+ }
+
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+}
--
2.49.0
From fd29084b3221e1901a1f07656d4e7abb31335caf Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:40:32 +0200
Subject: [PATCH v5 02/10] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 61 insertions(+), 40 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 47ffc0a2307..d1be780683b 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2321,7 +2321,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 608f10d9412..e40dae2ddf2 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,12 +1161,12 @@ assign_maintenance_io_concurrency(int newval, void *extra)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra)
+assign_io_max_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(newval, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra)
+assign_io_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(io_max_combine_limit, newval);
}
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index e5171467de1..2a6a587ef76 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1952,7 +1952,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1985,7 +1985,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2008,7 +2008,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2031,7 +2031,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 2f8c3d5f918..0d1b6466d1e 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3591,7 +3591,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 667df448732..bb681f5bc60 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f619100467d..8802ad8a3cb 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 799fa7ace68..c8300cffa8e 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,14 +81,14 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
-extern void assign_io_max_combine_limit(int newval, void *extra);
-extern void assign_io_combine_limit(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
+extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -143,13 +143,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -165,7 +165,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.49.0
From efbe93b30e0174c4fba42047b14208b3fd5c0f43 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 4 Apr 2025 21:46:14 +0200
Subject: [PATCH v5 03/10] Introduce pss_barrierReceivedGeneration
Currently WaitForProcSignalBarrier allows to make sure the message sent
via EmitProcSignalBarrier was processed by all ProcSignal mechanism
participants.
Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
which will be updated when a process has received the message, but not
processed it yet. This makes it possible to support a new mode of
waiting, when ProcSignal participants want to synchronize message
processing. To do that, a participant can wait via
WaitForProcSignalBarrierReceived when processing a message, effectively
making sure that all processes are going to start processing
ProcSignalBarrier simultaneously.
---
src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------
src/include/storage/procsignal.h | 1 +
2 files changed, 54 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index a9bb540b55a..c6bec9be423 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -58,7 +58,10 @@
* of it. For such use cases, we set a bit in pss_barrierCheckMask and then
* increment the current "barrier generation"; when the new barrier generation
* (or greater) appears in the pss_barrierGeneration flag of every process,
- * we know that the message has been received everywhere.
+ * we know that the message has been received and processed everywhere. In case
+ * if we only need to know only that the message was received everywhere (e.g.
+ * receiving processes need to handle the message in a coordinated fashion)
+ * use pss_barrierReceivedGeneration in the same way.
*/
typedef struct
{
@@ -70,6 +73,7 @@ typedef struct
/* Barrier-related fields (not protected by pss_mutex) */
pg_atomic_uint64 pss_barrierGeneration;
+ pg_atomic_uint64 pss_barrierReceivedGeneration;
pg_atomic_uint32 pss_barrierCheckMask;
ConditionVariable pss_barrierCV;
} ProcSignalSlot;
@@ -152,6 +156,8 @@ ProcSignalShmemInit(void)
slot->pss_cancel_key_len = 0;
MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration,
+ PG_UINT64_MAX);
pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
ConditionVariableInit(&slot->pss_barrierCV);
}
@@ -199,6 +205,8 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
barrier_generation =
pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration);
pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration,
+ barrier_generation);
if (cancel_key_len > 0)
memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len);
@@ -263,6 +271,7 @@ CleanupProcSignalState(int status, Datum arg)
* no barrier waits block on it.
*/
pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX);
SpinLockRelease(&slot->pss_mutex);
@@ -416,12 +425,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type)
return generation;
}
-/*
- * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
- * requested by a specific call to EmitProcSignalBarrier() have taken effect.
- */
-void
-WaitForProcSignalBarrier(uint64 generation)
+static void
+WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly)
{
Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration));
@@ -436,12 +441,17 @@ WaitForProcSignalBarrier(uint64 generation)
uint64 oldval;
/*
- * It's important that we check only pss_barrierGeneration here and
- * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared
- * before the barrier is actually absorbed, but pss_barrierGeneration
+ * It's important that we check only pss_barrierGeneration &
+ * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in
+ * pss_barrierCheckMask get cleared before the barrier is actually
+ * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration
* is updated only afterward.
*/
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
while (oldval < generation)
{
if (ConditionVariableTimedSleep(&slot->pss_barrierCV,
@@ -450,7 +460,11 @@ WaitForProcSignalBarrier(uint64 generation)
ereport(LOG,
(errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier",
(int) pg_atomic_read_u32(&slot->pss_pid))));
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
}
ConditionVariableCancelSleep();
}
@@ -464,12 +478,33 @@ WaitForProcSignalBarrier(uint64 generation)
* The caller is probably calling this function because it wants to read
* the shared state or perform further writes to shared state once all
* backends are known to have absorbed the barrier. However, the read of
- * pss_barrierGeneration was performed unlocked; insert a memory barrier
- * to separate it from whatever follows.
+ * pss_barrierGeneration & pss_barrierReceivedGeneration was performed
+ * unlocked; insert a memory barrier to separate it from whatever follows.
*/
pg_memory_barrier();
}
+/*
+ * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
+ * requested by a specific call to EmitProcSignalBarrier() have taken effect.
+ */
+void
+WaitForProcSignalBarrier(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, false);
+}
+
+/*
+ * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all
+ * backends have observed the message sent by a specific call to
+ * EmitProcSignalBarrier().
+ */
+void
+WaitForProcSignalBarrierReceived(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, true);
+}
+
/*
* Handle receipt of an interrupt indicating a global barrier event.
*
@@ -523,6 +558,10 @@ ProcessProcSignalBarrier(void)
if (local_gen == shared_gen)
return;
+ /* The message is observed, record that */
+ pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
+ shared_gen);
+
/*
* Get and clear the flags that are set for this backend. Note that
* pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index afeeb1ca019..2733bbb8c5b 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, const uint8 *cancel_key, int cance
extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type);
extern void WaitForProcSignalBarrier(uint64 generation);
+extern void WaitForProcSignalBarrierReceived(uint64 generation);
extern void ProcessProcSignalBarrier(void);
extern void procsignal_sigusr1_handler(SIGNAL_ARGS);
--
2.49.0
From a8e77ba00c05765f6d7ed05c8239a6bfecdbce4c Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH v5 04/10] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 +++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 148 +++++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 13 ++-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
12 files changed, 284 insertions(+), 126 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 423b2b4f9d6..4ce2cfb662b 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -307,7 +307,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -328,7 +328,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..56af0231d24 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 567739b5be9..5b55bec8d9d 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2fa045e6b0f..8b38e985327 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -363,7 +368,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index c9ae3b45b76..72255a1c5ca 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -76,19 +76,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -101,9 +101,17 @@ Datum pg_numa_available(PG_FUNCTION_ARGS);
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -114,7 +122,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -123,9 +137,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -150,11 +164,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -184,6 +204,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -203,22 +229,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -236,6 +262,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -246,19 +278,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -273,7 +305,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -334,6 +372,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -350,9 +400,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -385,6 +435,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -393,7 +450,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -416,7 +473,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -433,8 +490,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -459,7 +516,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -478,14 +535,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -542,10 +598,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -557,15 +614,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
@@ -630,7 +687,12 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ PGShmemHeader *shmhdr = Segments[segment].ShmemSegHdr;
+ shm_total_page_count += (shmhdr->totalsize / os_page_size) + 1;
+ }
+
page_ptrs = palloc0(sizeof(void *) * shm_total_page_count);
pages_status = palloc(sizeof(int) * shm_total_page_count);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 46f44bc4511..a36b08895c8 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -618,10 +620,15 @@ LWLockNewTrancheId(void)
int *LWLockCounter;
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
- /* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ /*
+ * We use the ShmemLock spinlock to protect LWLockCounter.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
+ */
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..6ebda479ced 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 5f7d4b83a60..2348c59b5a0 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -91,4 +106,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index c1f668ded95..69663d412c3 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
--
2.49.0
From 6238657ddb8c9e63d28a1a96712278f548d3292c Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 11:47:04 +0200
Subject: [PATCH v5 05/10] Address space reservation for shared memory
Currently the shared memory layout is designed to pack everything tight
together, leaving no space between mappings for resizing. Here is how it
looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
...
Make the layout more dynamic via splitting every shared memory segment
into two parts:
* An anonymous file, which actually contains shared memory content. Such
an anonymous file is created via memfd_create, it lives in memory,
behaves like a regular file and semantically equivalent to an
anonymous memory allocated via mmap with MAP_ANONYMOUS.
* A reservation mapping, which size is much larger than required shared
segment size. This mapping is created with flags PROT_NONE (which
makes sure the reserved space is not used), and MAP_NORESERVE (to not
count the reserved space against memory limits). The anonymous file is
mapped into this reservation mapping.
The resulting layout looks like this:
00400000-00490000 /path/bin/postgres
...
3f526000-3f590000 rw-p [heap]
7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
To resize a shared memory segment in this layout it's possible to use ftruncate
on the anonymous file, adjusting access permissions on the reserved space as
needed.
This approach also do not impact the actual memory usage as reported by
the kernel. Here is the output of /proc/$PID/status for the master
version with shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages
// mapped in mm_struct. It corresponds to the mapped reserved space
// and is the only number that grows with it.
VmPeak: 2043192 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 22908 kB
// Size of resident anonymous memory
RssAnon: 768 kB
// Size of resident file mappings
RssFile: 10364 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 11776 kB
Here is the same for the patch when reserving 20GB of space:
VmPeak: 21255824 kB
VmRSS: 25020 kB
RssAnon: 768 kB
RssFile: 10812 kB
RssShmem: 13440 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
20770816 (~19.8 MB)
To control the amount of space reserved a new GUC max_available_memory
is introduced. Ideally it should be based on the maximum available
memory, hense the name.
There are also few unrelated advantages of using anon files:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps.
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 290 ++++++++++++++++++++++------
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 5 +-
src/backend/storage/ipc/shmem.c | 2 +-
src/backend/utils/init/globals.c | 1 +
src/backend/utils/misc/guc_tables.c | 14 ++
src/include/miscadmin.h | 1 +
src/include/portability/mem.h | 2 +-
src/include/storage/pg_shmem.h | 5 +-
9 files changed, 262 insertions(+), 60 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 56af0231d24..363ddfd1fca 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -97,10 +97,12 @@ void *UsedShmemSegAddr = NULL;
typedef struct AnonymousMapping
{
int shmem_segment;
- Size shmem_size; /* Size of the mapping */
+ Size shmem_size; /* Size of the actually used memory */
+ Size shmem_reserved; /* Size of the reserved mapping */
Pointer shmem; /* Pointer to the start of the mapped memory */
Pointer seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -108,6 +110,49 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping layout we use looks like this:
+ *
+ * 00400000-00c2a000 r-xp /bin/postgres
+ * ...
+ * 3f526000-3f590000 rw-p [heap]
+ * 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted)
+ * 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted)
+ * 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
+ * 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
+ * ...
+ *
+ * We need to place shared memory mappings in such a way, that there will be
+ * gaps between them in the address space. Those gaps have to be large enough
+ * to resize the mapping up to certain size, without counting towards the total
+ * memory consumption.
+ *
+ * To achieve this, for each shared memory segment we first create an anonymous
+ * file of specified size using memfd_create, which will accomodate actual
+ * shared memory mapping content. It is represented by the first /memfd:main
+ * with rw permissions. Then we create a mapping for this file using mmap, with
+ * size much larger than required and flags PROT_NONE (allows to make sure the
+ * reserved space will not be used) and MAP_NORESERVE (prevents the space from
+ * being counted against memory limits). The mapping serves as an address space
+ * reservation, into which shared memory segment can be extended and is
+ * represented by the second /memfd:main with no permissions.
+ *
+ * The reserved space for each segment is calculated as a fraction of the total
+ * reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
+ * array.
+ */
+static double SHMEM_RESIZE_RATIO[1] = {
+ 1.0, /* MAIN_SHMEM_SLOT */
+};
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -503,19 +548,20 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* hugepage sizes, we might want to think about more invasive strategies,
* such as increasing shared_buffers to absorb the extra space.
*
- * Returns the (real, assumed or config provided) page size into
- * *hugepagesize, and the hugepage-related mmap flags to use into
- * *mmap_flags if requested by the caller. If huge pages are not supported,
- * *hugepagesize and *mmap_flags are set to 0.
+ * Returns the (real, assumed or config provided) page size into *hugepagesize,
+ * the hugepage-related mmap and memfd flags to use into *mmap_flags and
+ * *memfd_flags if requested by the caller. If huge pages are not supported,
+ * *hugepagesize, *mmap_flags and *memfd_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -574,6 +620,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -584,7 +631,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -593,6 +649,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -600,6 +658,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -625,72 +685,90 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
* Creates an anonymous mmap()ed shared memory segment.
*
* This function will modify mapping size to the actual size of the allocation,
- * if it ends up allocating a segment that is larger than requested.
+ * if it ends up allocating a segment that is larger than requested. If needed,
+ * it also rounds up the mapping reserved size to be a multiple of huge page
+ * size.
+ *
+ * Note that we do not fallback from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
*/
static void
CreateAnonymousSegment(AnonymousMapping *mapping)
{
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
- int mmap_errno = 0;
+ int save_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
+
+ elog(DEBUG1, "segment[%s]: size %zu, reserved %zu",
+ MappingName(mapping->shmem_segment), mapping->shmem_size,
+ mapping->shmem_reserved);
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the request size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- {
- DebugMappings();
- elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- MappingName(mapping->shmem_segment), allocsize);
- }
+ /*
+ * The reserved space is multiple of BLCKSZ. We know the huge page
+ * size, round up the reserved space to it.
+ */
+ mapping->shmem_reserved = mapping->shmem_reserved + hugepagesize -
+ (mapping->shmem_reserved % hugepagesize);
+
+ /* Verify that the new size is withing the reserved boundaries */
+ if (mapping->shmem_reserved < mapping->shmem_size)
+ ereport(ERROR,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved"),
+ errhint("You may need to increase \"max_available_memory\".")));
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
}
#endif
/*
- * Report whether huge pages are in use. This needs to be tracked before
- * the second mmap() call if attempting to use huge pages failed
- * previously.
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
*/
- SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
- PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment),
+ memfd_flags);
- if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified value.
+ */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
{
- /*
- * Use the original size, not the rounded-up value, when falling back
- * to non-huge pages.
- */
- allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
- }
+ save_errno = errno;
- if (ptr == MAP_FAILED)
- {
- errno = mmap_errno;
DebugMappings();
+ close(mapping->segment_fd);
+
+ errno = save_errno;
ereport(FATAL,
- (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ (errmsg("segment[%s]: could not truncate anonymous file: %m",
MappingName(mapping->shmem_segment)),
- (mmap_errno == ENOMEM) ?
+ (save_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
"swap space, or huge pages. To reduce the request size "
@@ -700,10 +778,112 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
allocsize) : 0));
}
+ elog(DEBUG1, "segment[%s]: mmap(%zu)",
+ MappingName(mapping->shmem_segment), allocsize);
+
+ /*
+ * Create a reservation mapping.
+ */
+ ptr = mmap(NULL, mapping->shmem_reserved, PROT_NONE,
+ mmap_flags | MAP_NORESERVE, mapping->segment_fd, 0);
+ save_errno = errno;
+
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+
+ errno = save_errno;
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment))));
+ }
+
+ /* Make the memory accessible */
+ if(mprotect(ptr, allocsize, PROT_READ | PROT_WRITE) == -1)
+ {
+ save_errno = errno;
+ DebugMappings();
+
+ errno = save_errno;
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not mprotect anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment))));
+ }
+
mapping->shmem = ptr;
mapping->shmem_size = allocsize;
}
+/*
+ * PrepareHugePages
+ *
+ * Figure out if there are enough huge pages to allocate all shared memory
+ * segments, and report that information via huge_pages_status and
+ * huge_pages_on. It needs to be called before creating shared memory segments.
+ *
+ * It is necessary to maintain the same semantic (simple on/off) for
+ * huge_pages_status, even if there are multiple shared memory segments: all
+ * segments either use huge pages or not, there is no mix of segments with
+ * different page size. The latter might be actually beneficial, in particular
+ * because only some segments may require large amount of memory, but for now
+ * we go with a simple solution.
+ */
+void
+PrepareHugePages()
+{
+ void *ptr = MAP_FAILED;
+
+ /* Reset to handle reinitialization */
+ next_free_segment = 0;
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ */
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ int numSemas;
+ Size segment_size = CalculateShmemSize(&numSemas, segment);
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
+ }
+#endif
+
+ /*
+ * Report whether huge pages are in use. This needs to be tracked before
+ * creating shared memory segments.
+ */
+ SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
+}
+
/*
* AnonymousShmemDetach --- detach from an anonymous mmap'd block
* (called as an on_shmem_exit callback, hence funny argument list)
@@ -746,7 +926,7 @@ PGSharedMemoryCreate(Size size,
void *memAddress;
PGShmemHeader *hdr;
struct stat statbuf;
- Size sysvsize;
+ Size sysvsize, total_reserved;
AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
@@ -760,14 +940,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -776,8 +948,16 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+
+ /* Prepare the mapping information */
mapping->shmem_size = size;
mapping->shmem_segment = next_free_segment;
+ total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+ mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+
+ /* Round up to be a multiple of BLCKSZ */
+ mapping->shmem_reserved = mapping->shmem_reserved + BLCKSZ -
+ (mapping->shmem_reserved % BLCKSZ);
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..732fedee87e 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8b38e985327..b60f7ef9ce2 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -206,6 +206,9 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
+ /* Decide if we use huge pages or regular size pages */
+ PrepareHugePages();
+
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
/* Compute the size of the shared-memory block */
@@ -377,7 +380,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 72255a1c5ca..8d025f0e907 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -817,7 +817,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index d31cb45a058..90d3feb547c 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffers = 16384;
+int MaxAvailableMemory = 524288;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index f04bfedb2fd..a221e446d6a 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2376,6 +2376,20 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"max_available_memory", PGC_SIGHUP, RESOURCES_MEM,
+ gettext_noop("Sets the upper limit for the shared_buffers value."),
+ gettext_noop("Shared memory could be resized at runtime, this "
+ "parameters sets the upper limit for it, beyond which "
+ "resizing would not be supported. Normally this value "
+ "would be the same as the total available memory."),
+ GUC_UNIT_BLOCKS
+ },
+ &MaxAvailableMemory,
+ 524288, 16, INT_MAX / 2,
+ NULL, NULL, NULL
+ },
+
{
{"vacuum_buffer_usage_limit", PGC_USERSET, RESOURCES_MEM,
gettext_noop("Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum."),
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 1bef98471c3..a0c37a7749e 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int MaxAvailableMemory;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 2348c59b5a0..79b0b1ef9eb 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -61,6 +61,7 @@ extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT int MaxAvailableMemory;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -104,7 +105,9 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+void PrepareHugePages(void);
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
--
2.49.0
From f23d42ef1ccdb28b751a8f12c7737002def4e674 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 11:22:02 +0200
Subject: [PATCH v5 06/10] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 24 +++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 105 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 363ddfd1fca..dac011b766b 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -139,10 +139,18 @@ static int next_free_segment = 0;
*
* The reserved space for each segment is calculated as a fraction of the total
* reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
- * array.
+ * array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take up to 60% of the whole
+ * space when resizing, based on the fact that it most likely will be the main
+ * consumer of this memory. Those numbers are pulled out of thin air for now,
+ * makes sense to evaluate them more precise.
*/
-static double SHMEM_RESIZE_RATIO[1] = {
- 1.0, /* MAIN_SHMEM_SLOT */
+static double SHMEM_RESIZE_RATIO[6] = {
+ 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.6, /* BUFFERS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -167,6 +175,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1dc488a42..bd68b69ee98 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +77,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -156,33 +160,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..a9952b36eba 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -22,6 +22,7 @@
#include "postgres.h"
#include "storage/buf_internals.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -59,10 +60,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 01909be0272..bd390f2709d 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -491,9 +492,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index b60f7ef9ce2..2dbd81afc87 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 41fdc1e7693..edac9db6a12 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -318,7 +318,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 79b0b1ef9eb..a7b275b4db9 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -109,7 +109,29 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.49.0
From c5717c5b68439abf4488fc3e1c2e47a94c6cd9f3 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 14:16:55 +0200
Subject: [PATCH v5 07/10] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers, adjusts its size using ftruncate and adjust reservation
permissions with mprotect. One elected process signals the postmaster
to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f87909fc000-7f8798248000 rw-s /memfd:strategy (deleted)
7f8798248000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a4e84000 rw-s /memfd:checkpoint (deleted)
7f87a4e84000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1b42000 rw-s /memfd:iocv (deleted)
7f87b1b42000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cb59c000 rw-s /memfd:descriptors (deleted)
7f87cb59c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f87ece38000 rw-s /memfd:buffers (deleted)
^ buffers content, ~247 MB
7f87ece38000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~2210 MB
7f8877066000-7f887e7d0000 rw-s /memfd:main (deleted)
7f887e7d0000-7f8890a00000 ---s /memfd:main (deleted)
-- 512 MB
7f87909fc000-7f879866a000 rw-s /memfd:strategy (deleted)
7f879866a000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a50f4000 rw-s /memfd:checkpoint (deleted)
7f87a50f4000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1d82000 rw-s /memfd:iocv (deleted)
7f87b1d82000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cba1c000 rw-s /memfd:descriptors (deleted)
7f87cba1c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f8804fb8000 rw-s /memfd:buffers (deleted)
^ buffers content, ~632 MB
7f8804fb8000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~1824 MB
7f8877066000-7f887e950000 rw-s /memfd:main (deleted)
7f887e950000-7f8890a00000 ---s /memfd:main (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it receives something.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 446 ++++++++++++++++++
src/backend/postmaster/checkpointer.c | 12 +-
src/backend/postmaster/postmaster.c | 18 +
src/backend/storage/buffer/buf_init.c | 74 +--
src/backend/storage/ipc/ipci.c | 18 +-
src/backend/storage/ipc/procsignal.c | 46 ++
src/backend/storage/ipc/shmem.c | 23 +-
src/backend/tcop/postgres.c | 10 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 26 +
src/include/storage/pmsignal.h | 1 +
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
17 files changed, 644 insertions(+), 45 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index dac011b766b..b3c90d15d52 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -107,6 +113,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -161,6 +174,49 @@ static double SHMEM_RESIZE_RATIO[6] = {
*/
static bool huge_pages_on = false;
+/*
+ * Flag telling that we have prepared the memory layout to be resizable. If
+ * false after all shared memory segments creation, it means we failed to setup
+ * needed layout and falled back to the regular non-resizable approach.
+ */
+static bool shmem_resizable = false;
+
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -924,6 +980,349 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * ftruncate, which will fail if is not sufficient space to expand the anon
+ * file. When finished, based on the new and old values initialize new buffer
+ * blocks if any.
+ *
+ * If reinitializing took place, as the last step this function does buffers
+ * reinitialization as well and broadcasts the new value of NSharedBuffers. All
+ * of that needs to be done only by one backend, the first one that managed to
+ * grab the ShmemResizeLock.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ int mmap_flags = PG_MMAP_FLAGS;
+ Size hugepagesize;
+
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+#ifndef MAP_HUGETLB
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the new size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+ }
+#endif
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+#ifdef MAP_HUGETLB
+ if (huge_pages_on && (new_size % hugepagesize != 0))
+ new_size += hugepagesize - (new_size % hugepagesize);
+#endif
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ if (m->shmem_reserved < new_size)
+ ereport(ERROR,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved"),
+ errhint("You may need to increase \"max_available_memory\".")));
+
+ elog(DEBUG1, "segment[%s]: resize from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ /* Resize the backing anon file. */
+ if(ftruncate(m->segment_fd, new_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ /* Adjust memory accessibility */
+ if(mprotect(m->shmem, new_size, PROT_READ | PROT_WRITE) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not mprotect anonymous shared memory for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ /* If shrinking, make reserved space unavailable again */
+ if(new_size < m->shmem_size &&
+ mprotect(m->shmem + new_size, m->shmem_size - new_size, PROT_NONE) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not mprotect reserved shared memory for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * Note that it's enough when only one backend will do that,
+ * even the ShmemInitStruct part. The reason is that resized
+ * shared memory will maintain the same addresses, meaning that
+ * all the pointers are still valid, and we only need to update
+ * structures size in the ShmemIndex once -- any other backend
+ * will pick up this shared structure from the index.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ BufferManagerShmemInit(NBuffersOld);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Wait for all ProcSignal participants
+ * to join the barrier, then do the resize and wait on the barrier until all
+ * participating finish resizing as well -- otherwise we face danger of
+ * inconsistency between backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ Assert(IsUnderPostmaster);
+
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * First thing to do after attaching to the barrier is to wait for others.
+ * We can't simply use BarrierArriveAndWait, because backends might arrive
+ * here in disjoint groups, e.g. first two backends, pause, then second two
+ * backends. If the resize is quick enough that can lead to a situation
+ * when the first group is already finished before the second has appeared,
+ * and the barrier will only synchonize withing those groups.
+ */
+ if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
+ WaitForProcSignalBarrierReceived(
+ pg_atomic_read_u64(&ShmemCtrl->Generation));
+
+ /*
+ * Now start the procedure, and elect one backend to ping postmaster to do
+ * the same.
+ *
+ * XXX: If we need to be able to abort resizing, this has to be done later,
+ * after the SHMEM_RESIZE_DONE.
+ */
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ {
+ Assert(IsUnderPostmaster);
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ }
+
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ BarrierDetach(barrier);
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Start resizing procedure, making sure all existing processes will have
+ * consistent view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * We use dynamic barrier to help dealing with backends that were spawned
+ * during the resize.
+ */
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ *
+ * Those phases are ensured by joining the shared barrier associated with
+ * the procedure. Since resizing takes time, we need to take into account
+ * that during that time:
+ *
+ * - New backends can be spawned. They will check status of the barrier
+ * early during the bootstrap, and wait until everything is over to work
+ * with the new NBuffers value.
+ *
+ * - Old backends can exit before attempting to resize. Synchronization
+ * used between backends relies on ProcSignalBarrier and waits for all
+ * participants received the message at the beginning to gather all
+ * existing backends.
+ *
+ * - Some backends might be blocked and not responsing either before or
+ * after receiving the message. In the first case such backend still
+ * have ProcSignalSlot and should be waited for, in the second case
+ * shared barrier will make sure we still waiting for those backends. In
+ * any case there is an unbounded wait.
+ *
+ * - Backends might join barrier in disjoint groups with some time in
+ * between. That means that relying only on the shared dynamic barrier is
+ * not enough -- it will only synchronize resize procedure withing those
+ * groups. That's why we wait first for all participants of ProcSignal
+ * mechanism who received the message.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ pg_atomic_init_u64(&ShmemCtrl->Generation,
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
+
+ /* To order everything after setting Generation value */
+ pg_memory_barrier();
+
+ /*
+ * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
+ * the rest of the pack has started the procedure and it can resize shared
+ * memory as well.
+ *
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier. In fact even if we would like to
+ * wait here, it wouldn't be possible -- we're in the postmaster, without
+ * any waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1239,3 +1638,50 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier()
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ /* Nothing to do if resizing is not started */
+ if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
+ return;
+
+ BarrierAttach(barrier);
+
+ /* Otherwise wait through all available phases */
+ while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
+ {
+ ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
+ BarrierPhase(barrier))));
+
+ BarrierArriveAndWait(barrier, 0);
+ }
+
+ BarrierDetach(barrier);
+}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /*
+ * The barrier is missing here, it will be initialized right before
+ * starting the resizing process as a convenient way to reset it.
+ */
+
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+
+ /* shmem_resizable should be initialized by now */
+ ShmemCtrl->Resizable = shmem_resizable;
+ }
+}
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index fda91ffd1ce..ab08e1a182b 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -638,9 +638,12 @@ CheckpointerMain(const void *startup_data, size_t startup_data_len)
static void
ProcessCheckpointerInterrupts(void)
{
- if (ProcSignalBarrierPending)
- ProcessProcSignalBarrier();
-
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
if (ConfigReloadPending)
{
ConfigReloadPending = false;
@@ -660,6 +663,9 @@ ProcessCheckpointerInterrupts(void)
UpdateSharedMemoryConfig();
}
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+
/* Perform logging of memory contexts of this process */
if (LogMemoryContextPending)
ProcessLogMemoryContextInterrupt();
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 490f7ce3664..f0cb0098dcd 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -426,6 +426,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1694,6 +1695,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2039,6 +2043,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
@@ -3852,6 +3867,9 @@ process_pm_pmsignal(void)
request_state_update = true;
}
+ if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
+ AnonymousShmemResize();
+
/*
* Try to advance postmaster's state machine, if a child requests it.
*/
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index bd68b69ee98..8c1ea623392 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -62,18 +63,28 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * postmaster, or in a standalone backend) or during shared-memory resize. Size
+ * of data structures initialized here depends on NBuffers, and to be able to
+ * change NBuffers without a restart we store each structure into a separate
+ * shared memory segment, which could be resized on demand.
+ *
+ * FirstBufferToInit tells where to start initializing buffers. For
+ * initialization it always will be zero, but when resizing shared-memory it
+ * indicates the number of already initialized buffers.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(void)
+BufferManagerShmemInit(int FirstBufferToInit)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
+ elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
+ FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -110,43 +121,44 @@ BufferManagerShmemInit(void)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory, or in all cases when resizing shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
- ClearBufferTag(&buf->tag);
+#ifndef EXEC_BACKEND
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = FirstBufferToInit; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ ClearBufferTag(&buf->tag);
- buf->buf_id = i;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- pgaio_wref_clear(&buf->io_wref);
+ buf->buf_id = i;
- /*
- * Initially link all the buffers together as unused. Subsequent
- * management of this list is done by freelist.c.
- */
- buf->freeNext = i + 1;
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
+#endif
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2dbd81afc87..c5725f55120 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -84,6 +84,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -151,6 +154,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, AioShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -298,7 +309,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit();
+ BufferManagerShmemInit(0);
/*
* Set up lock manager
@@ -310,6 +321,11 @@ CreateOrAttachShmemStructs(void)
*/
PredicateLockShmemInit();
+ /*
+ * Set up shared memory resize manager
+ */
+ ShmemControlInit();
+
/*
* Set up process table
*/
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index c6bec9be423..d7b56a18b24 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -113,6 +114,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -176,6 +181,43 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -615,6 +657,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 8d025f0e907..9fa277b91d7 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -498,17 +498,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
+ *
+ * XXX: There is an implicit assumption this can only happen in
+ * "resizable" segments, where only one shared structure is allowed.
+ * This has to be implemented more cleanly.
*/
if (result->size != size)
{
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
+ Size delta = size - result->size;
+
+ result->size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
}
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 0d1b6466d1e..0942d2bffe2 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4309,6 +4310,15 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /* Verify the shared barrier, if it's still active: join and wait. */
+ WaitOnShmemBarrier();
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 4da68312b5f..691fa14e9e3 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
@@ -352,6 +354,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index a221e446d6a..d8acec0f911 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2366,14 +2366,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index edac9db6a12..4239ebe640b 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -317,7 +317,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_skipped);
/* in buf_init.c */
-extern void BufferManagerShmemInit(void);
+extern void BufferManagerShmemInit(int);
extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6ebda479ced..bb7ae4d33b3 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index a9681738146..558da6fdd55 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
+PG_LWLOCK(54, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index a7b275b4db9..bccdd45b1f7 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,25 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+ pg_atomic_uint64 Generation;
+ bool Resizable;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -109,6 +129,12 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 428aa3fd68a..1a55bf57a70 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,6 +42,7 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 2733bbb8c5b..97033f84dce 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 32d6e718adc..36e25ea8cd6 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2751,6 +2751,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.49.0
From 5e190dc84ce7078b6c680d18736e56be90cecdab Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 19 Jun 2025 17:38:29 +0200
Subject: [PATCH v5 08/10] Support shrinking shared buffers
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
Buffer eviction requires a separate barrier phase for two reasons:
1. No other backend should map a new page to any of buffers being
evicted when eviction is in progress. So they wait while eviction is
in progress.
2. Since a pinned buffer has the pin recorded in the backend local
memory as well as the buffer descriptor (which is in shared memory),
eviction should not coincide with remapping the shared memory of a
backend. Otherwise we might loose consistency of local and shared
pinning records. Hence it needs to be carried out in
ProcessBarrierShmemResize() and not in AnonymousShmemResize() as
indicated by now removed comment.
If a buffer being evicted is pinned, we raise a FATAL error but this should
improve. There are multiple options 1. to wait for the pinned buffer to get
unpinned, 2. the backend is killed or it itself cancels the query or 3.
rollback the operation. Note that option 1 and 2 would require the pinning
related local and shared records to be accessed. But we need infrastructure to
do either of this right now.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 42 +++++---
src/backend/storage/buffer/buf_init.c | 8 +-
src/backend/storage/buffer/bufmgr.c | 95 +++++++++++++++++++
src/backend/storage/buffer/freelist.c | 68 +++++++++++++
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 1 +
src/include/storage/pg_shmem.h | 1 +
8 files changed, 202 insertions(+), 15 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index b3c90d15d52..e612a83c9f0 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1011,14 +1011,6 @@ AnonymousShmemResize(void)
*/
pending_pm_shmem_resize = false;
- /*
- * XXX: Currently only increasing of shared_buffers is supported. For
- * decreasing something similar has to be done, but buffer blocks with
- * data have to be drained first.
- */
- if(NBuffersOld > NBuffers)
- return false;
-
#ifndef MAP_HUGETLB
/* PrepareHugePages should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
@@ -1112,11 +1104,14 @@ AnonymousShmemResize(void)
* all the pointers are still valid, and we only need to update
* structures size in the ShmemIndex once -- any other backend
* will pick up this shared structure from the index.
- *
- * XXX: This is the right place for buffer eviction as well.
*/
BufferManagerShmemInit(NBuffersOld);
+ /*
+ * Wipe out the evictor PID so that it can be used for the next
+ * buffer resizing operation.
+ */
+ ShmemCtrl->evictor_pid = 0;
/* If all fine, broadcast the new value */
pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
}
@@ -1169,11 +1164,31 @@ ProcessBarrierShmemResize(Barrier *barrier)
* XXX: If we need to be able to abort resizing, this has to be done later,
* after the SHMEM_RESIZE_DONE.
*/
- if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+
+ /*
+ * Evict extra buffers when shrinking shared buffers. We need to do this
+ * while the memory for extra buffers is still mapped i.e. before remapping
+ * the shared memory segments to a smaller memory area.
+ */
+ if (NBuffersOld > NBuffersPending)
{
- Assert(IsUnderPostmaster);
- SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
+
+ /*
+ * TODO: If the buffer eviction fails for any reason, we should
+ * gracefully rollback the shared buffer resizing and try again. But the
+ * infrastructure to do so is not available right now. Hence just raise
+ * a FATAL so that the system restarts.
+ */
+ if (!EvictExtraBuffers(NBuffersPending, NBuffersOld))
+ elog(FATAL, "buffer eviction failed");
+
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_EVICT))
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
}
+ else
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
AnonymousShmemResize();
@@ -1683,5 +1698,6 @@ ShmemControlInit(void)
/* shmem_resizable should be initialized by now */
ShmemCtrl->Resizable = shmem_resizable;
+ ShmemCtrl->evictor_pid = 0;
}
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 8c1ea623392..6f5743036a2 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -157,8 +157,12 @@ BufferManagerShmemInit(int FirstBufferToInit)
}
#endif
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ /*
+ * Correct last entry of linked list, when initializing the buffers or when
+ * expanding the buffers.
+ */
+ if (FirstBufferToInit < NBuffers)
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 667aa0c0c78..169a44dd9fc 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/read_stream.h"
#include "storage/smgr.h"
@@ -7453,3 +7454,97 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int newBufSize, int oldBufSize)
+{
+ bool result = true;
+
+ /*
+ * If the buffer being evicated is locked, this function will need to wait.
+ * This function should not be called from a Postmaster since it can not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * Let only one backend perform eviction. We could split the work across all
+ * the backends but that doesn't seem necessary.
+ *
+ * The first backend to acquire ShmemResizeLock, sets its own PID as the
+ * evictor PID for other backends to know that the eviction is in progress or
+ * has already been performed. The evictor backend releases the lock when it
+ * finishes eviction. While the eviction is in progress, backends other than
+ * evictor backend won't be able to take the lock. They won't perform
+ * eviction. A backend may acquire the lock after eviction has completed, but
+ * it will not perform eviction since the evictor PID is already set. Evictor
+ * PID is reset only when the buffer resizing finishes. Thus only one backend
+ * will perform eviction in a given instance of shared buffers resizing.
+ *
+ * Any backend which acquires this lock will release it before the eviction
+ * phase finishes, hence the same lock can be reused for the next phase of
+ * resizing buffers.
+ */
+ if (LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ if (ShmemCtrl->evictor_pid == 0)
+ {
+ ShmemCtrl->evictor_pid = MyProcPid;
+
+ StrategyPurgeFreeList(newBufSize);
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = newBufSize + 1; buf <= oldBufSize; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u32(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is
+ * going one hence unlocked precheck should be safe and saves
+ * some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ /*
+ * XXX: Looks like CurrentResourceOwner can be NULL here, find
+ * another one in that case?
+ * */
+ if (CurrentResourceOwner)
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+ }
+ LWLockRelease(ShmemResizeLock);
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index bd390f2709d..7b9ed010e2f 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -527,6 +527,74 @@ StrategyInitialize(bool init)
}
+/*
+ * StrategyPurgeFreeList -- remove all buffers with id higher than the number of
+ * buffers in the buffer pool.
+ *
+ * This is called before evicting buffers while shrinking shared buffers, so that
+ * the free list does not reference a buffer that will be removed.
+ *
+ * The function is called after resizing has started and thus nobody should be
+ * traversing the free list and also not touching the buffers.
+ */
+void
+StrategyPurgeFreeList(int numBuffers)
+{
+ int firstBuffer = FREENEXT_END_OF_LIST;
+ int nextFree = StrategyControl->firstFreeBuffer;
+ BufferDesc *prevValidBuf = NULL;
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ while (nextFree != FREENEXT_END_OF_LIST)
+ {
+ BufferDesc *buf = GetBufferDescriptor(nextFree);
+
+ /* nextFree should be id of buffer being examined. */
+ Assert(nextFree == buf->buf_id);
+ /* The buffer should not be marked as not in the list. */
+ Assert(buf->freeNext != FREENEXT_NOT_IN_LIST);
+
+ /*
+ * If the buffer is within the new size of pool, keep it in the free list
+ * otherwise discard it.
+ */
+ if (buf->buf_id < numBuffers)
+ {
+ if (prevValidBuf != NULL)
+ prevValidBuf->freeNext = buf->buf_id;
+ prevValidBuf = buf;
+
+ /* Save the first free buffer in the list if not already known. */
+ if (firstBuffer == FREENEXT_NOT_IN_LIST)
+ firstBuffer = nextFree;
+ }
+ /* Examine the next buffer in the free list. */
+ nextFree = buf->freeNext;
+ }
+
+ /* Update the last valid free buffer, if there's any. */
+ if (prevValidBuf != NULL)
+ {
+ StrategyControl->lastFreeBuffer = prevValidBuf->buf_id;
+ prevValidBuf->freeNext = FREENEXT_END_OF_LIST;
+ }
+ else
+ StrategyControl->lastFreeBuffer = FREENEXT_END_OF_LIST;
+
+ /* Update first valid free buffer, if there's any. */
+ StrategyControl->firstFreeBuffer = firstBuffer;
+
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+
+ /*
+ * TODO: following was suggested by AI. Check whether it is required.
+ * If we removed all buffers from the freelist, reset the clock sweep
+ * pointer to zero. This is not strictly necessary, but it seems like a
+ * good idea to avoid confusion.
+ */
+}
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
* ----------------------------------------------------------------
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 691fa14e9e3..0c588b69a90 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -156,6 +156,7 @@ REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it c
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_EVICT "Waiting for other backends to finish buffer evication phase."
SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 0dec7d93b3b..add15e3723b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -453,6 +453,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyPurgeFreeList(int numBuffers);
extern bool have_free_buffer(void);
/* buf_table.c */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 4239ebe640b..0c554f0b130 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -315,6 +315,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_evicted,
int32 *buffers_flushed,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int fromBuf, int toBuf);
/* in buf_init.c */
extern void BufferManagerShmemInit(int);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index bccdd45b1f7..9a3e68f76d8 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -64,6 +64,7 @@ extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
typedef struct
{
pg_atomic_uint32 NSharedBuffers;
+ pid_t evictor_pid;
Barrier Barrier;
pg_atomic_uint64 Generation;
bool Resizable;
--
2.49.0
From 475fffc314b4d5975faf94d43504f2dcc0f1dc8d Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 19 Jun 2025 17:38:51 +0200
Subject: [PATCH v5 09/10] Reinitialize StrategyControl after resizing buffers
... and BgBufferSync and ClockSweepTick adjustments
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. When expanding the buffer pool add new buffers to the free list.
2. When shrinking buffers, we remove any buffers, in the area being
shrunk, from the freelist. While doing so we adjust the first and
last free buffer pointers in the StrategyControl area. Hence nothing
more needed after resizing.
3. Check the sanity of the free buffer list is added after resizing.
4. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
5. &StrategyControl->buffer_strategy_lock need not be initialized again.
6. nextVictimBuffer, completePasses and numBufferAllocs are viewed in
the context of NBuffers. Now that NBuffers itself has changed, those
three do not make sense. Reset them as if the server has restarted
again.
This commit introduces a flag delay_shmem_resize, which postgresql
backends and workers can use to signal the coordinator to delay resizing
operation. Background writer sets this flag when its scanning buffers. Background
writer is blocked when the actual resizing is in progress. But if
resizing is about to begin, it does not scan the buffers by returning
from BgBufferSync(). It stops a scan in progress when it sees that the
resizing has begun. After the resizing is finished, it adjusts the
collected statistics according to the new size of the buffer pool at the
end of barrier processing.
Once the buffer resizing is finished, before resuming the regular
operation, bgwriter resets the information saved so far. This
information is viewed in the context of NBuffers and hence does not make
sense after NBuffer has changed.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 24 ++++-
src/backend/postmaster/bgwriter.c | 2 +-
src/backend/storage/buffer/buf_init.c | 11 ++-
src/backend/storage/buffer/bufmgr.c | 74 ++++++++++----
src/backend/storage/buffer/freelist.c | 133 ++++++++++++++++++++++++++
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 3 +-
src/include/storage/ipc.h | 1 +
8 files changed, 226 insertions(+), 23 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index e612a83c9f0..e61a557966b 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -115,6 +115,7 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Flag telling postmaster that resize is needed */
volatile bool pending_pm_shmem_resize = false;
+volatile bool delay_shmem_resize = false;
/* Keeps track of the previous NBuffers value */
static int NBuffersOld = -1;
@@ -1118,6 +1119,12 @@ AnonymousShmemResize(void)
LWLockRelease(ShmemResizeLock);
}
+
+ /*
+ * TODO: Shouldn't we call ResizeBufferPool() here as well? Or those
+ * backend who can not lock the LWLock conditionally won't resize the
+ * buffers.
+ */
}
return true;
@@ -1138,13 +1145,17 @@ ProcessBarrierShmemResize(Barrier *barrier)
{
Assert(IsUnderPostmaster);
- elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
- NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize, delay_shmem_resize);
/* Wait until we have seen the new NBuffers value */
if (!pending_pm_shmem_resize)
return false;
+ /* Wait till this process becomes ready to resize buffers. */
+ if (delay_shmem_resize)
+ return false;
+
/*
* First thing to do after attaching to the barrier is to wait for others.
* We can't simply use BarrierArriveAndWait, because backends might arrive
@@ -1195,6 +1206,15 @@ ProcessBarrierShmemResize(Barrier *barrier)
/* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * Before resuming regular background writer activity, adjust the
+ * statistics collected so far.
+ */
+ BgBufferSyncAdjust(NBuffersOld, NBuffers);
+ }
+
BarrierDetach(barrier);
return true;
}
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index 72f5acceec7..32b34f28ead 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -233,7 +233,7 @@ BackgroundWriterMain(const void *startup_data, size_t startup_data_len)
/*
* Do one cycle of dirty-buffer writing.
*/
- can_hibernate = BgBufferSync(&wb_context);
+ can_hibernate = BgBufferSync(&wb_context, false);
/* Report pending statistics to the cumulative stats system */
pgstat_report_bgwriter();
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6f5743036a2..62c1672fb70 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -164,8 +164,15 @@ BufferManagerShmemInit(int FirstBufferToInit)
if (FirstBufferToInit < NBuffers)
GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
- /* Init other shared buffer-management stuff */
- StrategyInitialize(!foundDescs);
+ /*
+ * Init other shared buffer-management stuff from scratch configuring buffer
+ * pool the first time. If we are just resizing buffer pool adjust only the
+ * required structures.
+ */
+ if (FirstBufferToInit == 0)
+ StrategyInitialize(!foundDescs);
+ else
+ StrategyReInitialize(FirstBufferToInit);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 169a44dd9fc..590fc737da8 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3611,6 +3611,32 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_DONE(NBuffers, num_written, num_to_scan);
}
+/*
+ * Information saved between BgBufferSync() calls so we can determine the
+ * strategy point's advance rate and avoid scanning already-cleaned buffers. The
+ * variables are global instead of static local so that BgBufferSyncAdjust() can
+ * adjust it when resizing shared buffers.
+ */
+static bool saved_info_valid = false;
+static int prev_strategy_buf_id;
+static uint32 prev_strategy_passes;
+static int next_to_clean;
+static uint32 next_passes;
+
+/* Moving averages of allocation rate and clean-buffer density */
+static float smoothed_alloc = 0;
+static float smoothed_density = 10.0;
+
+void
+BgBufferSyncAdjust(int NBuffersOld, int NBuffersNew)
+{
+ saved_info_valid = false;
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
+ NBuffersOld, NBuffersNew);
+#endif
+}
+
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
@@ -3623,27 +3649,13 @@ BufferSync(int flags)
* bgwriter_lru_maxpages to 0.)
*/
bool
-BgBufferSync(WritebackContext *wb_context)
+BgBufferSync(WritebackContext *wb_context, bool reset)
{
/* info obtained from freelist.c */
int strategy_buf_id;
uint32 strategy_passes;
uint32 recent_alloc;
- /*
- * Information saved between calls so we can determine the strategy
- * point's advance rate and avoid scanning already-cleaned buffers.
- */
- static bool saved_info_valid = false;
- static int prev_strategy_buf_id;
- static uint32 prev_strategy_passes;
- static int next_to_clean;
- static uint32 next_passes;
-
- /* Moving averages of allocation rate and clean-buffer density */
- static float smoothed_alloc = 0;
- static float smoothed_density = 10.0;
-
/* Potentially these could be tunables, but for now, not */
float smoothing_samples = 16;
float scan_whole_pool_milliseconds = 120000.0;
@@ -3666,6 +3678,22 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ /*
+ * If buffer pool is being shrunk the buffer being written out may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing out any. Hence wait till
+ * buffer resizing finishes i.e. go into hibernation mode.
+ */
+ if (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ return true;
+
+ /*
+ * Resizing shared buffers while this function is performing an LRU scan on
+ * them may lead to wrong results. Indicate that the resizing should wait for
+ * the LRU scan to complete.
+ */
+ delay_shmem_resize = true;
+
/*
* Find out where the freelist clock sweep currently is, and how many
* buffer allocations have happened since our last call.
@@ -3842,8 +3870,17 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
- while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
+ /*
+ * Execute the LRU scan.
+ *
+ * If buffer pool is being shrunk, the buffer being written may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing any. Hence stop what we are doing. This
+ * also unblocks other processes that are waiting for buffer resizing to
+ * finish.
+ */
+ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
+ pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) == NBuffers)
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -3902,6 +3939,9 @@ BgBufferSync(WritebackContext *wb_context)
#endif
}
+ /* Let the resizing commence. */
+ delay_shmem_resize = false;
+
/* Return true if OK to hibernate */
return (bufs_to_lap == 0 && recent_alloc == 0);
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 7b9ed010e2f..41641bb3ae6 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -98,6 +98,9 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
uint32 *buf_state);
static void AddBufferToRing(BufferAccessStrategy strategy,
BufferDesc *buf);
+#ifdef USE_ASSERT_CHECKING
+static void StrategyValidateFreeList(void);
+#endif /* USE_ASSERT_CHECKING */
/*
* ClockSweepTick - Helper routine for StrategyGetBuffer()
@@ -526,6 +529,88 @@ StrategyInitialize(bool init)
Assert(!init);
}
+/*
+ * StrategyReInitialize -- re-initialize the buffer cache replacement
+ * strategy.
+ *
+ * To be called when resizing buffer manager and only from the coordinator.
+ * TODO: Assess the differences between this function and StrategyInitialize().
+ */
+void
+StrategyReInitialize(int FirstBufferIdToInit)
+{
+ bool found;
+
+ /*
+ * Resizing memory for buffer pools should not affect the address of
+ * StrategyControl.
+ */
+ if (StrategyControl != (BufferStrategyControl *)
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, STRATEGY_SHMEM_SEGMENT))
+ elog(FATAL, "something went wrong while re-initializing the buffer strategy");
+
+ Assert(found);
+
+ /* TODO: Buffer lookup table adjustment: There are two options:
+ *
+ * 1. Resize the buffer lookup table to match the new number of buffers. But
+ * this requires rehashing all the entries in the buffer lookup table with
+ * the new table size.
+ *
+ * 2. Allocate maximum size of the buffer lookup table at the beginning and
+ * never resize it. This leaves sparse buffer lookup table which is
+ * inefficient from both memory and time perspective. According to David
+ * Rowley, the sparse entries in the buffer look up table cause frequent
+ * cacheline reload which affect performance. If the impact of that
+ * inefficiency in a benchmark is significant, we will need to consider first
+ * option.
+ */
+
+ /*
+ * When shrinking buffers, we must have adjusted the first and the last free
+ * buffer when removing the buffers being shrunk from the free list. Nothing
+ * to be done here.
+ *
+ * When expanding the shared buffers, new buffers are added at the end of the
+ * freelist or they form the new free list if there are no free buffers.
+ */
+ if (FirstBufferIdToInit < NBuffers)
+ {
+ if (StrategyControl->firstFreeBuffer == FREENEXT_END_OF_LIST)
+ StrategyControl->firstFreeBuffer = FirstBufferIdToInit;
+ else
+ {
+ Assert(StrategyControl->lastFreeBuffer >= 0);
+ GetBufferDescriptor(StrategyControl->lastFreeBuffer - 1)->freeNext = FirstBufferIdToInit;
+ }
+
+ StrategyControl->lastFreeBuffer = NBuffers - 1;
+ }
+
+ /* Check free list sanity after resizing. */
+#ifdef USE_ASSERT_CHECKING
+ StrategyValidateFreeList();
+#endif /* USE_ASSERT_CHECKING */
+
+ /*
+ * The clock sweep tick pointer might have got invalidated. Reset it as if
+ * starting a fresh server.
+ */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The old statistics is viewed in the context of the number of shared
+ * buffers. It does not make sense now that the number of shared buffers
+ * itself has changed.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_init_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* No pending notification */
+ StrategyControl->bgwprocno = -1;
+}
/*
* StrategyPurgeFreeList -- remove all buffers with id higher than the number of
@@ -595,6 +680,54 @@ StrategyPurgeFreeList(int numBuffers)
*/
}
+#ifdef USE_ASSERT_CHECKING
+/*
+ * StrategyValidateFreeList-- check sanity of free buffer list.
+ */
+static void
+StrategyValidateFreeList(void)
+{
+ int nextFree = StrategyControl->firstFreeBuffer;
+ int numFreeBuffers = 0;
+ int lastFreeBuffer = FREENEXT_END_OF_LIST;
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ while (nextFree != FREENEXT_END_OF_LIST)
+ {
+ BufferDesc *buf = GetBufferDescriptor(nextFree);
+
+ /* nextFree should be id of buffer being examined. */
+ Assert(nextFree == buf->buf_id);
+ Assert(buf->buf_id < NBuffers);
+ /* The buffer should not be marked as not in the list. */
+ Assert(buf->freeNext != FREENEXT_NOT_IN_LIST);
+
+ /* Update our knowledge of last buffer in the free list. */
+ lastFreeBuffer = buf->buf_id;
+
+ numFreeBuffers++;
+
+ /* Avoid infinite recursion in case there are cycles in free list. */
+ if (numFreeBuffers > NBuffers)
+ break;
+
+ nextFree = buf->freeNext;
+ }
+
+ Assert(numFreeBuffers <= NBuffers);
+
+ /*
+ * Make sure that the StrategyControl's knowledge of last free buffer
+ * agrees with what's there in the free list.
+ */
+ if (StrategyControl->firstFreeBuffer != FREENEXT_END_OF_LIST)
+ Assert(StrategyControl->lastFreeBuffer == lastFreeBuffer);
+
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+}
+#endif /* USE_ASSERT_CHECKING */
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
* ----------------------------------------------------------------
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index add15e3723b..46949e9d90e 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -454,6 +454,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
extern void StrategyPurgeFreeList(int numBuffers);
+extern void StrategyReInitialize(int FirstBufferToInit);
extern bool have_free_buffer(void);
/* buf_table.c */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 0c554f0b130..83a75eab844 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -298,7 +298,8 @@ extern bool ConditionalLockBufferForCleanup(Buffer buffer);
extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
-extern bool BgBufferSync(struct WritebackContext *wb_context);
+extern bool BgBufferSync(struct WritebackContext *wb_context, bool reset);
+extern void BgBufferSyncAdjust(int NBuffersOld, int NBuffersNew);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index bb7ae4d33b3..7d1c64a9267 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -65,6 +65,7 @@ typedef void (*shmem_startup_hook_type) (void);
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
+extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
--
2.49.0
From 25816411c7a77a87b8527a901633865f1c1056b7 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 11 Jun 2025 18:15:06 +0530
Subject: [PATCH v5 10/10] Additional validation for buffer in the ring
If the buffer pool has been shrunk, the buffers in the buffer list may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition. That may not be
expensive enough to affect query performance. The alternative to that is
more complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure thus accessible when processing buffer
resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work.
Ashutosh Bapat
---
src/backend/storage/buffer/freelist.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 41641bb3ae6..74d070733a4 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -948,12 +948,13 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot has been
+ * invalidated when buffer pool was shrunk, tell the caller to allocate a new
+ * buffer with the normal allocation strategy. He will then fill this slot
+ * by calling AddBufferToRing with the new buffer.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer || bufnum > NBuffers)
return NULL;
/*
--
2.49.0
Attachments:
[text/plain] v5-0001-Process-config-reload-in-AIO-workers.patch (1.8K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/2-v5-0001-Process-config-reload-in-AIO-workers.patch)
download | inline diff:
From 5d4f46dc13bacf9c1df79233f8e4017a9dcd6919 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 15:14:33 +0200
Subject: [PATCH v5 01/10] Process config reload in AIO workers
Currenly AIO workers process interrupts only via CHECK_FOR_INTERRUPTS,
which does not include ConfigReloadPending. Thus we need to check for it
explicitly.
---
src/backend/storage/aio/method_worker.c | 25 +++++++++++++++++++++++++
1 file changed, 25 insertions(+)
diff --git a/src/backend/storage/aio/method_worker.c b/src/backend/storage/aio/method_worker.c
index 36be179678d..b4d5c46fb94 100644
--- a/src/backend/storage/aio/method_worker.c
+++ b/src/backend/storage/aio/method_worker.c
@@ -80,6 +80,7 @@ static void pgaio_worker_shmem_init(bool first_time);
static bool pgaio_worker_needs_synchronous_execution(PgAioHandle *ioh);
static int pgaio_worker_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+static void pgaio_worker_process_interrupts(void);
const IoMethodOps pgaio_worker_ops = {
.shmem_size = pgaio_worker_shmem_size,
@@ -461,6 +462,8 @@ IoWorkerMain(const void *startup_data, size_t startup_data_len)
int nwakeups = 0;
int worker;
+ pgaio_worker_process_interrupts();
+
/*
* Try to get a job to do.
*
@@ -584,3 +587,25 @@ pgaio_workers_enabled(void)
{
return io_method == IOMETHOD_WORKER;
}
+
+/*
+ * Process any new interrupts.
+ */
+static void
+pgaio_worker_process_interrupts(void)
+{
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
+ if (ConfigReloadPending)
+ {
+ ConfigReloadPending = false;
+ ProcessConfigFile(PGC_SIGHUP);
+ }
+
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+}
--
2.49.0
[text/plain] v5-0002-Introduce-pending-flag-for-GUC-assign-hooks.patch (12.8K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/3-v5-0002-Introduce-pending-flag-for-GUC-assign-hooks.patch)
download | inline diff:
From fd29084b3221e1901a1f07656d4e7abb31335caf Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:40:32 +0200
Subject: [PATCH v5 02/10] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 61 insertions(+), 40 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 47ffc0a2307..d1be780683b 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2321,7 +2321,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 608f10d9412..e40dae2ddf2 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,12 +1161,12 @@ assign_maintenance_io_concurrency(int newval, void *extra)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra)
+assign_io_max_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(newval, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra)
+assign_io_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(io_max_combine_limit, newval);
}
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index e5171467de1..2a6a587ef76 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1952,7 +1952,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1985,7 +1985,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2008,7 +2008,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2031,7 +2031,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 2f8c3d5f918..0d1b6466d1e 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3591,7 +3591,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 667df448732..bb681f5bc60 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2041,13 +2046,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f619100467d..8802ad8a3cb 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 799fa7ace68..c8300cffa8e 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,14 +81,14 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
-extern void assign_io_max_combine_limit(int newval, void *extra);
-extern void assign_io_combine_limit(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
+extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
extern bool check_max_slot_wal_keep_size(int *newval, void **extra,
GucSource source);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -143,13 +143,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -165,7 +165,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.49.0
[text/plain] v5-0003-Introduce-pss_barrierReceivedGeneration.patch (7.3K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/4-v5-0003-Introduce-pss_barrierReceivedGeneration.patch)
download | inline diff:
From efbe93b30e0174c4fba42047b14208b3fd5c0f43 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 4 Apr 2025 21:46:14 +0200
Subject: [PATCH v5 03/10] Introduce pss_barrierReceivedGeneration
Currently WaitForProcSignalBarrier allows to make sure the message sent
via EmitProcSignalBarrier was processed by all ProcSignal mechanism
participants.
Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
which will be updated when a process has received the message, but not
processed it yet. This makes it possible to support a new mode of
waiting, when ProcSignal participants want to synchronize message
processing. To do that, a participant can wait via
WaitForProcSignalBarrierReceived when processing a message, effectively
making sure that all processes are going to start processing
ProcSignalBarrier simultaneously.
---
src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------
src/include/storage/procsignal.h | 1 +
2 files changed, 54 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index a9bb540b55a..c6bec9be423 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -58,7 +58,10 @@
* of it. For such use cases, we set a bit in pss_barrierCheckMask and then
* increment the current "barrier generation"; when the new barrier generation
* (or greater) appears in the pss_barrierGeneration flag of every process,
- * we know that the message has been received everywhere.
+ * we know that the message has been received and processed everywhere. In case
+ * if we only need to know only that the message was received everywhere (e.g.
+ * receiving processes need to handle the message in a coordinated fashion)
+ * use pss_barrierReceivedGeneration in the same way.
*/
typedef struct
{
@@ -70,6 +73,7 @@ typedef struct
/* Barrier-related fields (not protected by pss_mutex) */
pg_atomic_uint64 pss_barrierGeneration;
+ pg_atomic_uint64 pss_barrierReceivedGeneration;
pg_atomic_uint32 pss_barrierCheckMask;
ConditionVariable pss_barrierCV;
} ProcSignalSlot;
@@ -152,6 +156,8 @@ ProcSignalShmemInit(void)
slot->pss_cancel_key_len = 0;
MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration,
+ PG_UINT64_MAX);
pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
ConditionVariableInit(&slot->pss_barrierCV);
}
@@ -199,6 +205,8 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
barrier_generation =
pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration);
pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration,
+ barrier_generation);
if (cancel_key_len > 0)
memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len);
@@ -263,6 +271,7 @@ CleanupProcSignalState(int status, Datum arg)
* no barrier waits block on it.
*/
pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX);
SpinLockRelease(&slot->pss_mutex);
@@ -416,12 +425,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type)
return generation;
}
-/*
- * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
- * requested by a specific call to EmitProcSignalBarrier() have taken effect.
- */
-void
-WaitForProcSignalBarrier(uint64 generation)
+static void
+WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly)
{
Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration));
@@ -436,12 +441,17 @@ WaitForProcSignalBarrier(uint64 generation)
uint64 oldval;
/*
- * It's important that we check only pss_barrierGeneration here and
- * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared
- * before the barrier is actually absorbed, but pss_barrierGeneration
+ * It's important that we check only pss_barrierGeneration &
+ * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in
+ * pss_barrierCheckMask get cleared before the barrier is actually
+ * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration
* is updated only afterward.
*/
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
while (oldval < generation)
{
if (ConditionVariableTimedSleep(&slot->pss_barrierCV,
@@ -450,7 +460,11 @@ WaitForProcSignalBarrier(uint64 generation)
ereport(LOG,
(errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier",
(int) pg_atomic_read_u32(&slot->pss_pid))));
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
}
ConditionVariableCancelSleep();
}
@@ -464,12 +478,33 @@ WaitForProcSignalBarrier(uint64 generation)
* The caller is probably calling this function because it wants to read
* the shared state or perform further writes to shared state once all
* backends are known to have absorbed the barrier. However, the read of
- * pss_barrierGeneration was performed unlocked; insert a memory barrier
- * to separate it from whatever follows.
+ * pss_barrierGeneration & pss_barrierReceivedGeneration was performed
+ * unlocked; insert a memory barrier to separate it from whatever follows.
*/
pg_memory_barrier();
}
+/*
+ * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
+ * requested by a specific call to EmitProcSignalBarrier() have taken effect.
+ */
+void
+WaitForProcSignalBarrier(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, false);
+}
+
+/*
+ * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all
+ * backends have observed the message sent by a specific call to
+ * EmitProcSignalBarrier().
+ */
+void
+WaitForProcSignalBarrierReceived(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, true);
+}
+
/*
* Handle receipt of an interrupt indicating a global barrier event.
*
@@ -523,6 +558,10 @@ ProcessProcSignalBarrier(void)
if (local_gen == shared_gen)
return;
+ /* The message is observed, record that */
+ pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
+ shared_gen);
+
/*
* Get and clear the flags that are set for this backend. Note that
* pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index afeeb1ca019..2733bbb8c5b 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, const uint8 *cancel_key, int cance
extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type);
extern void WaitForProcSignalBarrier(uint64 generation);
+extern void WaitForProcSignalBarrierReceived(uint64 generation);
extern void ProcessProcSignalBarrier(void);
extern void procsignal_sigusr1_handler(SIGNAL_ARGS);
--
2.49.0
[text/plain] v5-0004-Allow-to-use-multiple-shared-memory-mappings.patch (30.5K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/5-v5-0004-Allow-to-use-multiple-shared-memory-mappings.patch)
download | inline diff:
From a8e77ba00c05765f6d7ed05c8239a6bfecdbce4c Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH v5 04/10] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 +++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 148 +++++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 13 ++-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 12 +++
12 files changed, 284 insertions(+), 126 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 423b2b4f9d6..4ce2cfb662b 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -307,7 +307,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -328,7 +328,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..56af0231d24 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 567739b5be9..5b55bec8d9d 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2fa045e6b0f..8b38e985327 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -363,7 +368,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index c9ae3b45b76..72255a1c5ca 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -76,19 +76,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -101,9 +101,17 @@ Datum pg_numa_available(PG_FUNCTION_ARGS);
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -114,7 +122,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -123,9 +137,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -150,11 +164,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -184,6 +204,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -203,22 +229,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -236,6 +262,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -246,19 +278,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -273,7 +305,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -334,6 +372,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
long max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -350,9 +400,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -385,6 +435,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -393,7 +450,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -416,7 +473,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -433,8 +490,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -459,7 +516,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -478,14 +535,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -542,10 +598,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -557,15 +614,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
@@ -630,7 +687,12 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ PGShmemHeader *shmhdr = Segments[segment].ShmemSegHdr;
+ shm_total_page_count += (shmhdr->totalsize / os_page_size) + 1;
+ }
+
page_ptrs = palloc0(sizeof(void *) * shm_total_page_count);
pages_status = palloc(sizeof(int) * shm_total_page_count);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 46f44bc4511..a36b08895c8 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -618,10 +620,15 @@ LWLockNewTrancheId(void)
int *LWLockCounter;
LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int));
- /* We use the ShmemLock spinlock to protect LWLockCounter */
- SpinLockAcquire(ShmemLock);
+ /*
+ * We use the ShmemLock spinlock to protect LWLockCounter.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
+ */
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
result = (*LWLockCounter)++;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..6ebda479ced 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 5f7d4b83a60..2348c59b5a0 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -91,4 +106,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index c1f668ded95..69663d412c3 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,15 +29,27 @@
extern PGDLLIMPORT slock_t *ShmemLock;
struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */
extern void InitShmemAccess(struct PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
+extern void InitVariableShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
--
2.49.0
[text/plain] v5-0005-Address-space-reservation-for-shared-memory.patch (24.2K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/6-v5-0005-Address-space-reservation-for-shared-memory.patch)
download | inline diff:
From 6238657ddb8c9e63d28a1a96712278f548d3292c Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 11:47:04 +0200
Subject: [PATCH v5 05/10] Address space reservation for shared memory
Currently the shared memory layout is designed to pack everything tight
together, leaving no space between mappings for resizing. Here is how it
looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
...
Make the layout more dynamic via splitting every shared memory segment
into two parts:
* An anonymous file, which actually contains shared memory content. Such
an anonymous file is created via memfd_create, it lives in memory,
behaves like a regular file and semantically equivalent to an
anonymous memory allocated via mmap with MAP_ANONYMOUS.
* A reservation mapping, which size is much larger than required shared
segment size. This mapping is created with flags PROT_NONE (which
makes sure the reserved space is not used), and MAP_NORESERVE (to not
count the reserved space against memory limits). The anonymous file is
mapped into this reservation mapping.
The resulting layout looks like this:
00400000-00490000 /path/bin/postgres
...
3f526000-3f590000 rw-p [heap]
7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
To resize a shared memory segment in this layout it's possible to use ftruncate
on the anonymous file, adjusting access permissions on the reserved space as
needed.
This approach also do not impact the actual memory usage as reported by
the kernel. Here is the output of /proc/$PID/status for the master
version with shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages
// mapped in mm_struct. It corresponds to the mapped reserved space
// and is the only number that grows with it.
VmPeak: 2043192 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 22908 kB
// Size of resident anonymous memory
RssAnon: 768 kB
// Size of resident file mappings
RssFile: 10364 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 11776 kB
Here is the same for the patch when reserving 20GB of space:
VmPeak: 21255824 kB
VmRSS: 25020 kB
RssAnon: 768 kB
RssFile: 10812 kB
RssShmem: 13440 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
20770816 (~19.8 MB)
To control the amount of space reserved a new GUC max_available_memory
is introduced. Ideally it should be based on the maximum available
memory, hense the name.
There are also few unrelated advantages of using anon files:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps.
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 290 ++++++++++++++++++++++------
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 5 +-
src/backend/storage/ipc/shmem.c | 2 +-
src/backend/utils/init/globals.c | 1 +
src/backend/utils/misc/guc_tables.c | 14 ++
src/include/miscadmin.h | 1 +
src/include/portability/mem.h | 2 +-
src/include/storage/pg_shmem.h | 5 +-
9 files changed, 262 insertions(+), 60 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 56af0231d24..363ddfd1fca 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -97,10 +97,12 @@ void *UsedShmemSegAddr = NULL;
typedef struct AnonymousMapping
{
int shmem_segment;
- Size shmem_size; /* Size of the mapping */
+ Size shmem_size; /* Size of the actually used memory */
+ Size shmem_reserved; /* Size of the reserved mapping */
Pointer shmem; /* Pointer to the start of the mapped memory */
Pointer seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -108,6 +110,49 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping layout we use looks like this:
+ *
+ * 00400000-00c2a000 r-xp /bin/postgres
+ * ...
+ * 3f526000-3f590000 rw-p [heap]
+ * 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted)
+ * 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted)
+ * 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
+ * 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
+ * ...
+ *
+ * We need to place shared memory mappings in such a way, that there will be
+ * gaps between them in the address space. Those gaps have to be large enough
+ * to resize the mapping up to certain size, without counting towards the total
+ * memory consumption.
+ *
+ * To achieve this, for each shared memory segment we first create an anonymous
+ * file of specified size using memfd_create, which will accomodate actual
+ * shared memory mapping content. It is represented by the first /memfd:main
+ * with rw permissions. Then we create a mapping for this file using mmap, with
+ * size much larger than required and flags PROT_NONE (allows to make sure the
+ * reserved space will not be used) and MAP_NORESERVE (prevents the space from
+ * being counted against memory limits). The mapping serves as an address space
+ * reservation, into which shared memory segment can be extended and is
+ * represented by the second /memfd:main with no permissions.
+ *
+ * The reserved space for each segment is calculated as a fraction of the total
+ * reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
+ * array.
+ */
+static double SHMEM_RESIZE_RATIO[1] = {
+ 1.0, /* MAIN_SHMEM_SLOT */
+};
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -503,19 +548,20 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* hugepage sizes, we might want to think about more invasive strategies,
* such as increasing shared_buffers to absorb the extra space.
*
- * Returns the (real, assumed or config provided) page size into
- * *hugepagesize, and the hugepage-related mmap flags to use into
- * *mmap_flags if requested by the caller. If huge pages are not supported,
- * *hugepagesize and *mmap_flags are set to 0.
+ * Returns the (real, assumed or config provided) page size into *hugepagesize,
+ * the hugepage-related mmap and memfd flags to use into *mmap_flags and
+ * *memfd_flags if requested by the caller. If huge pages are not supported,
+ * *hugepagesize, *mmap_flags and *memfd_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -574,6 +620,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -584,7 +631,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -593,6 +649,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -600,6 +658,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -625,72 +685,90 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
* Creates an anonymous mmap()ed shared memory segment.
*
* This function will modify mapping size to the actual size of the allocation,
- * if it ends up allocating a segment that is larger than requested.
+ * if it ends up allocating a segment that is larger than requested. If needed,
+ * it also rounds up the mapping reserved size to be a multiple of huge page
+ * size.
+ *
+ * Note that we do not fallback from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
*/
static void
CreateAnonymousSegment(AnonymousMapping *mapping)
{
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
- int mmap_errno = 0;
+ int save_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
+
+ elog(DEBUG1, "segment[%s]: size %zu, reserved %zu",
+ MappingName(mapping->shmem_segment), mapping->shmem_size,
+ mapping->shmem_reserved);
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the request size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- {
- DebugMappings();
- elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- MappingName(mapping->shmem_segment), allocsize);
- }
+ /*
+ * The reserved space is multiple of BLCKSZ. We know the huge page
+ * size, round up the reserved space to it.
+ */
+ mapping->shmem_reserved = mapping->shmem_reserved + hugepagesize -
+ (mapping->shmem_reserved % hugepagesize);
+
+ /* Verify that the new size is withing the reserved boundaries */
+ if (mapping->shmem_reserved < mapping->shmem_size)
+ ereport(ERROR,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved"),
+ errhint("You may need to increase \"max_available_memory\".")));
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
}
#endif
/*
- * Report whether huge pages are in use. This needs to be tracked before
- * the second mmap() call if attempting to use huge pages failed
- * previously.
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
*/
- SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
- PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment),
+ memfd_flags);
- if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified value.
+ */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
{
- /*
- * Use the original size, not the rounded-up value, when falling back
- * to non-huge pages.
- */
- allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
- }
+ save_errno = errno;
- if (ptr == MAP_FAILED)
- {
- errno = mmap_errno;
DebugMappings();
+ close(mapping->segment_fd);
+
+ errno = save_errno;
ereport(FATAL,
- (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ (errmsg("segment[%s]: could not truncate anonymous file: %m",
MappingName(mapping->shmem_segment)),
- (mmap_errno == ENOMEM) ?
+ (save_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
"swap space, or huge pages. To reduce the request size "
@@ -700,10 +778,112 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
allocsize) : 0));
}
+ elog(DEBUG1, "segment[%s]: mmap(%zu)",
+ MappingName(mapping->shmem_segment), allocsize);
+
+ /*
+ * Create a reservation mapping.
+ */
+ ptr = mmap(NULL, mapping->shmem_reserved, PROT_NONE,
+ mmap_flags | MAP_NORESERVE, mapping->segment_fd, 0);
+ save_errno = errno;
+
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+
+ errno = save_errno;
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment))));
+ }
+
+ /* Make the memory accessible */
+ if(mprotect(ptr, allocsize, PROT_READ | PROT_WRITE) == -1)
+ {
+ save_errno = errno;
+ DebugMappings();
+
+ errno = save_errno;
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not mprotect anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment))));
+ }
+
mapping->shmem = ptr;
mapping->shmem_size = allocsize;
}
+/*
+ * PrepareHugePages
+ *
+ * Figure out if there are enough huge pages to allocate all shared memory
+ * segments, and report that information via huge_pages_status and
+ * huge_pages_on. It needs to be called before creating shared memory segments.
+ *
+ * It is necessary to maintain the same semantic (simple on/off) for
+ * huge_pages_status, even if there are multiple shared memory segments: all
+ * segments either use huge pages or not, there is no mix of segments with
+ * different page size. The latter might be actually beneficial, in particular
+ * because only some segments may require large amount of memory, but for now
+ * we go with a simple solution.
+ */
+void
+PrepareHugePages()
+{
+ void *ptr = MAP_FAILED;
+
+ /* Reset to handle reinitialization */
+ next_free_segment = 0;
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ */
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ int numSemas;
+ Size segment_size = CalculateShmemSize(&numSemas, segment);
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
+ }
+#endif
+
+ /*
+ * Report whether huge pages are in use. This needs to be tracked before
+ * creating shared memory segments.
+ */
+ SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
+}
+
/*
* AnonymousShmemDetach --- detach from an anonymous mmap'd block
* (called as an on_shmem_exit callback, hence funny argument list)
@@ -746,7 +926,7 @@ PGSharedMemoryCreate(Size size,
void *memAddress;
PGShmemHeader *hdr;
struct stat statbuf;
- Size sysvsize;
+ Size sysvsize, total_reserved;
AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
@@ -760,14 +940,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -776,8 +948,16 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+
+ /* Prepare the mapping information */
mapping->shmem_size = size;
mapping->shmem_segment = next_free_segment;
+ total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+ mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+
+ /* Round up to be a multiple of BLCKSZ */
+ mapping->shmem_reserved = mapping->shmem_reserved + BLCKSZ -
+ (mapping->shmem_reserved % BLCKSZ);
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..732fedee87e 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8b38e985327..b60f7ef9ce2 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -206,6 +206,9 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
+ /* Decide if we use huge pages or regular size pages */
+ PrepareHugePages();
+
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
/* Compute the size of the shared-memory block */
@@ -377,7 +380,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 72255a1c5ca..8d025f0e907 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -817,7 +817,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index d31cb45a058..90d3feb547c 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffers = 16384;
+int MaxAvailableMemory = 524288;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index f04bfedb2fd..a221e446d6a 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2376,6 +2376,20 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"max_available_memory", PGC_SIGHUP, RESOURCES_MEM,
+ gettext_noop("Sets the upper limit for the shared_buffers value."),
+ gettext_noop("Shared memory could be resized at runtime, this "
+ "parameters sets the upper limit for it, beyond which "
+ "resizing would not be supported. Normally this value "
+ "would be the same as the total available memory."),
+ GUC_UNIT_BLOCKS
+ },
+ &MaxAvailableMemory,
+ 524288, 16, INT_MAX / 2,
+ NULL, NULL, NULL
+ },
+
{
{"vacuum_buffer_usage_limit", PGC_USERSET, RESOURCES_MEM,
gettext_noop("Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum."),
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 1bef98471c3..a0c37a7749e 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int MaxAvailableMemory;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 2348c59b5a0..79b0b1ef9eb 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -61,6 +61,7 @@ extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT int MaxAvailableMemory;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -104,7 +105,9 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+void PrepareHugePages(void);
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
--
2.49.0
[text/plain] v5-0006-Introduce-multiple-shmem-segments-for-shared-buff.patch (11.6K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/7-v5-0006-Introduce-multiple-shmem-segments-for-shared-buff.patch)
download | inline diff:
From f23d42ef1ccdb28b751a8f12c7737002def4e674 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 11:22:02 +0200
Subject: [PATCH v5 06/10] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 24 +++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 105 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 363ddfd1fca..dac011b766b 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -139,10 +139,18 @@ static int next_free_segment = 0;
*
* The reserved space for each segment is calculated as a fraction of the total
* reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
- * array.
+ * array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take up to 60% of the whole
+ * space when resizing, based on the fact that it most likely will be the main
+ * consumer of this memory. Those numbers are pulled out of thin air for now,
+ * makes sense to evaluate them more precise.
*/
-static double SHMEM_RESIZE_RATIO[1] = {
- 1.0, /* MAIN_SHMEM_SLOT */
+static double SHMEM_RESIZE_RATIO[6] = {
+ 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.6, /* BUFFERS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -167,6 +175,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1dc488a42..bd68b69ee98 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +77,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -156,33 +160,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a50955d5286..a9952b36eba 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -22,6 +22,7 @@
#include "postgres.h"
#include "storage/buf_internals.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -59,10 +60,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 01909be0272..bd390f2709d 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -491,9 +492,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index b60f7ef9ce2..2dbd81afc87 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 41fdc1e7693..edac9db6a12 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -318,7 +318,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 79b0b1ef9eb..a7b275b4db9 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -109,7 +109,29 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.49.0
[text/plain] v5-0007-Allow-to-resize-shared-memory-without-restart.patch (41.3K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/8-v5-0007-Allow-to-resize-shared-memory-without-restart.patch)
download | inline diff:
From c5717c5b68439abf4488fc3e1c2e47a94c6cd9f3 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 14:16:55 +0200
Subject: [PATCH v5 07/10] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers, adjusts its size using ftruncate and adjust reservation
permissions with mprotect. One elected process signals the postmaster
to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f87909fc000-7f8798248000 rw-s /memfd:strategy (deleted)
7f8798248000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a4e84000 rw-s /memfd:checkpoint (deleted)
7f87a4e84000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1b42000 rw-s /memfd:iocv (deleted)
7f87b1b42000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cb59c000 rw-s /memfd:descriptors (deleted)
7f87cb59c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f87ece38000 rw-s /memfd:buffers (deleted)
^ buffers content, ~247 MB
7f87ece38000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~2210 MB
7f8877066000-7f887e7d0000 rw-s /memfd:main (deleted)
7f887e7d0000-7f8890a00000 ---s /memfd:main (deleted)
-- 512 MB
7f87909fc000-7f879866a000 rw-s /memfd:strategy (deleted)
7f879866a000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a50f4000 rw-s /memfd:checkpoint (deleted)
7f87a50f4000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1d82000 rw-s /memfd:iocv (deleted)
7f87b1d82000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cba1c000 rw-s /memfd:descriptors (deleted)
7f87cba1c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f8804fb8000 rw-s /memfd:buffers (deleted)
^ buffers content, ~632 MB
7f8804fb8000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~1824 MB
7f8877066000-7f887e950000 rw-s /memfd:main (deleted)
7f887e950000-7f8890a00000 ---s /memfd:main (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it receives something.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 446 ++++++++++++++++++
src/backend/postmaster/checkpointer.c | 12 +-
src/backend/postmaster/postmaster.c | 18 +
src/backend/storage/buffer/buf_init.c | 74 +--
src/backend/storage/ipc/ipci.c | 18 +-
src/backend/storage/ipc/procsignal.c | 46 ++
src/backend/storage/ipc/shmem.c | 23 +-
src/backend/tcop/postgres.c | 10 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_tables.c | 4 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 26 +
src/include/storage/pmsignal.h | 1 +
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
17 files changed, 644 insertions(+), 45 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index dac011b766b..b3c90d15d52 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -107,6 +113,13 @@ typedef struct AnonymousMapping
static AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
@@ -161,6 +174,49 @@ static double SHMEM_RESIZE_RATIO[6] = {
*/
static bool huge_pages_on = false;
+/*
+ * Flag telling that we have prepared the memory layout to be resizable. If
+ * false after all shared memory segments creation, it means we failed to setup
+ * needed layout and falled back to the regular non-resizable approach.
+ */
+static bool shmem_resizable = false;
+
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -924,6 +980,349 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * ftruncate, which will fail if is not sufficient space to expand the anon
+ * file. When finished, based on the new and old values initialize new buffer
+ * blocks if any.
+ *
+ * If reinitializing took place, as the last step this function does buffers
+ * reinitialization as well and broadcasts the new value of NSharedBuffers. All
+ * of that needs to be done only by one backend, the first one that managed to
+ * grab the ShmemResizeLock.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ int mmap_flags = PG_MMAP_FLAGS;
+ Size hugepagesize;
+
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+#ifndef MAP_HUGETLB
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the new size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+ }
+#endif
+
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ Size new_size = CalculateShmemSize(&numSemas, i);
+ AnonymousMapping *m = &Mappings[i];
+
+#ifdef MAP_HUGETLB
+ if (huge_pages_on && (new_size % hugepagesize != 0))
+ new_size += hugepagesize - (new_size % hugepagesize);
+#endif
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == new_size)
+ continue;
+
+ if (m->shmem_reserved < new_size)
+ ereport(ERROR,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved"),
+ errhint("You may need to increase \"max_available_memory\".")));
+
+ elog(DEBUG1, "segment[%s]: resize from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ new_size, m->shmem);
+
+ /* Resize the backing anon file. */
+ if(ftruncate(m->segment_fd, new_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ /* Adjust memory accessibility */
+ if(mprotect(m->shmem, new_size, PROT_READ | PROT_WRITE) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not mprotect anonymous shared memory for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ /* If shrinking, make reserved space unavailable again */
+ if(new_size < m->shmem_size &&
+ mprotect(m->shmem + new_size, m->shmem_size - new_size, PROT_NONE) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not mprotect reserved shared memory for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ reinit = true;
+ m->shmem_size = new_size;
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * Note that it's enough when only one backend will do that,
+ * even the ShmemInitStruct part. The reason is that resized
+ * shared memory will maintain the same addresses, meaning that
+ * all the pointers are still valid, and we only need to update
+ * structures size in the ShmemIndex once -- any other backend
+ * will pick up this shared structure from the index.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ BufferManagerShmemInit(NBuffersOld);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Wait for all ProcSignal participants
+ * to join the barrier, then do the resize and wait on the barrier until all
+ * participating finish resizing as well -- otherwise we face danger of
+ * inconsistency between backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ Assert(IsUnderPostmaster);
+
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * First thing to do after attaching to the barrier is to wait for others.
+ * We can't simply use BarrierArriveAndWait, because backends might arrive
+ * here in disjoint groups, e.g. first two backends, pause, then second two
+ * backends. If the resize is quick enough that can lead to a situation
+ * when the first group is already finished before the second has appeared,
+ * and the barrier will only synchonize withing those groups.
+ */
+ if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
+ WaitForProcSignalBarrierReceived(
+ pg_atomic_read_u64(&ShmemCtrl->Generation));
+
+ /*
+ * Now start the procedure, and elect one backend to ping postmaster to do
+ * the same.
+ *
+ * XXX: If we need to be able to abort resizing, this has to be done later,
+ * after the SHMEM_RESIZE_DONE.
+ */
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ {
+ Assert(IsUnderPostmaster);
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ }
+
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ BarrierDetach(barrier);
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ /* Request shared memory resize only when it was initialized */
+ if (next_free_segment != 0)
+ {
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ }
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Start resizing procedure, making sure all existing processes will have
+ * consistent view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * We use dynamic barrier to help dealing with backends that were spawned
+ * during the resize.
+ */
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld || next_free_segment == 0)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d",
+ NBuffers, NBuffersOld, next_free_segment);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ *
+ * Those phases are ensured by joining the shared barrier associated with
+ * the procedure. Since resizing takes time, we need to take into account
+ * that during that time:
+ *
+ * - New backends can be spawned. They will check status of the barrier
+ * early during the bootstrap, and wait until everything is over to work
+ * with the new NBuffers value.
+ *
+ * - Old backends can exit before attempting to resize. Synchronization
+ * used between backends relies on ProcSignalBarrier and waits for all
+ * participants received the message at the beginning to gather all
+ * existing backends.
+ *
+ * - Some backends might be blocked and not responsing either before or
+ * after receiving the message. In the first case such backend still
+ * have ProcSignalSlot and should be waited for, in the second case
+ * shared barrier will make sure we still waiting for those backends. In
+ * any case there is an unbounded wait.
+ *
+ * - Backends might join barrier in disjoint groups with some time in
+ * between. That means that relying only on the shared dynamic barrier is
+ * not enough -- it will only synchronize resize procedure withing those
+ * groups. That's why we wait first for all participants of ProcSignal
+ * mechanism who received the message.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ pg_atomic_init_u64(&ShmemCtrl->Generation,
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
+
+ /* To order everything after setting Generation value */
+ pg_memory_barrier();
+
+ /*
+ * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
+ * the rest of the pack has started the procedure and it can resize shared
+ * memory as well.
+ *
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier. In fact even if we would like to
+ * wait here, it wouldn't be possible -- we're in the postmaster, without
+ * any waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1239,3 +1638,50 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier()
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ /* Nothing to do if resizing is not started */
+ if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
+ return;
+
+ BarrierAttach(barrier);
+
+ /* Otherwise wait through all available phases */
+ while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
+ {
+ ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
+ BarrierPhase(barrier))));
+
+ BarrierArriveAndWait(barrier, 0);
+ }
+
+ BarrierDetach(barrier);
+}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /*
+ * The barrier is missing here, it will be initialized right before
+ * starting the resizing process as a convenient way to reset it.
+ */
+
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+
+ /* shmem_resizable should be initialized by now */
+ ShmemCtrl->Resizable = shmem_resizable;
+ }
+}
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index fda91ffd1ce..ab08e1a182b 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -638,9 +638,12 @@ CheckpointerMain(const void *startup_data, size_t startup_data_len)
static void
ProcessCheckpointerInterrupts(void)
{
- if (ProcSignalBarrierPending)
- ProcessProcSignalBarrier();
-
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
if (ConfigReloadPending)
{
ConfigReloadPending = false;
@@ -660,6 +663,9 @@ ProcessCheckpointerInterrupts(void)
UpdateSharedMemoryConfig();
}
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+
/* Perform logging of memory contexts of this process */
if (LogMemoryContextPending)
ProcessLogMemoryContextInterrupt();
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 490f7ce3664..f0cb0098dcd 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -426,6 +426,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1694,6 +1695,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2039,6 +2043,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
@@ -3852,6 +3867,9 @@ process_pm_pmsignal(void)
request_state_update = true;
}
+ if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
+ AnonymousShmemResize();
+
/*
* Try to advance postmaster's state machine, if a child requests it.
*/
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index bd68b69ee98..8c1ea623392 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -62,18 +63,28 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * postmaster, or in a standalone backend) or during shared-memory resize. Size
+ * of data structures initialized here depends on NBuffers, and to be able to
+ * change NBuffers without a restart we store each structure into a separate
+ * shared memory segment, which could be resized on demand.
+ *
+ * FirstBufferToInit tells where to start initializing buffers. For
+ * initialization it always will be zero, but when resizing shared-memory it
+ * indicates the number of already initialized buffers.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(void)
+BufferManagerShmemInit(int FirstBufferToInit)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
+ elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
+ FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -110,43 +121,44 @@ BufferManagerShmemInit(void)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory, or in all cases when resizing shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
- ClearBufferTag(&buf->tag);
+#ifndef EXEC_BACKEND
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = FirstBufferToInit; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ ClearBufferTag(&buf->tag);
- buf->buf_id = i;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- pgaio_wref_clear(&buf->io_wref);
+ buf->buf_id = i;
- /*
- * Initially link all the buffers together as unused. Subsequent
- * management of this list is done by freelist.c.
- */
- buf->freeNext = i + 1;
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ /*
+ * Initially link all the buffers together as unused. Subsequent
+ * management of this list is done by freelist.c.
+ */
+ buf->freeNext = i + 1;
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
+#endif
+
+ /* Correct last entry of linked list */
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2dbd81afc87..c5725f55120 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -84,6 +84,9 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * XXX: Calculation for non main shared memory segments are incorrect, it
+ * includes more than needed for buffers only.
*/
Size
CalculateShmemSize(int *num_semaphores, int shmem_segment)
@@ -151,6 +154,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, AioShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -298,7 +309,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit();
+ BufferManagerShmemInit(0);
/*
* Set up lock manager
@@ -310,6 +321,11 @@ CreateOrAttachShmemStructs(void)
*/
PredicateLockShmemInit();
+ /*
+ * Set up shared memory resize manager
+ */
+ ShmemControlInit();
+
/*
* Set up process table
*/
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index c6bec9be423..d7b56a18b24 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -113,6 +114,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -176,6 +181,43 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -615,6 +657,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 8d025f0e907..9fa277b91d7 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -498,17 +498,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
+ *
+ * XXX: There is an implicit assumption this can only happen in
+ * "resizable" segments, where only one shared structure is allowed.
+ * This has to be implemented more cleanly.
*/
if (result->size != size)
{
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
+ Size delta = size - result->size;
+
+ result->size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
}
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 0d1b6466d1e..0942d2bffe2 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4309,6 +4310,15 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /* Verify the shared barrier, if it's still active: join and wait. */
+ WaitOnShmemBarrier();
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 4da68312b5f..691fa14e9e3 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
@@ -352,6 +354,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index a221e446d6a..d8acec0f911 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -2366,14 +2366,14 @@ struct config_int ConfigureNamesInt[] =
* checking for overflow, so we mustn't allow more than INT_MAX / 2.
*/
{
- {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM,
+ {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM,
gettext_noop("Sets the number of shared memory buffers used by the server."),
NULL,
GUC_UNIT_BLOCKS
},
&NBuffers,
16384, 16, INT_MAX / 2,
- NULL, NULL, NULL
+ NULL, assign_shared_buffers, NULL
},
{
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index edac9db6a12..4239ebe640b 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -317,7 +317,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_skipped);
/* in buf_init.c */
-extern void BufferManagerShmemInit(void);
+extern void BufferManagerShmemInit(int);
extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6ebda479ced..bb7ae4d33b3 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index a9681738146..558da6fdd55 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
+PG_LWLOCK(54, ShmemResize)
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index a7b275b4db9..bccdd45b1f7 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -56,6 +57,25 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+ pg_atomic_uint64 Generation;
+ bool Resizable;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -109,6 +129,12 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 428aa3fd68a..1a55bf57a70 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,6 +42,7 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 2733bbb8c5b..97033f84dce 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 32d6e718adc..36e25ea8cd6 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2751,6 +2751,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.49.0
[text/plain] v5-0008-Support-shrinking-shared-buffers.patch (14.0K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/9-v5-0008-Support-shrinking-shared-buffers.patch)
download | inline diff:
From 5e190dc84ce7078b6c680d18736e56be90cecdab Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 19 Jun 2025 17:38:29 +0200
Subject: [PATCH v5 08/10] Support shrinking shared buffers
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
Buffer eviction requires a separate barrier phase for two reasons:
1. No other backend should map a new page to any of buffers being
evicted when eviction is in progress. So they wait while eviction is
in progress.
2. Since a pinned buffer has the pin recorded in the backend local
memory as well as the buffer descriptor (which is in shared memory),
eviction should not coincide with remapping the shared memory of a
backend. Otherwise we might loose consistency of local and shared
pinning records. Hence it needs to be carried out in
ProcessBarrierShmemResize() and not in AnonymousShmemResize() as
indicated by now removed comment.
If a buffer being evicted is pinned, we raise a FATAL error but this should
improve. There are multiple options 1. to wait for the pinned buffer to get
unpinned, 2. the backend is killed or it itself cancels the query or 3.
rollback the operation. Note that option 1 and 2 would require the pinning
related local and shared records to be accessed. But we need infrastructure to
do either of this right now.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 42 +++++---
src/backend/storage/buffer/buf_init.c | 8 +-
src/backend/storage/buffer/bufmgr.c | 95 +++++++++++++++++++
src/backend/storage/buffer/freelist.c | 68 +++++++++++++
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 1 +
src/include/storage/pg_shmem.h | 1 +
8 files changed, 202 insertions(+), 15 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index b3c90d15d52..e612a83c9f0 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1011,14 +1011,6 @@ AnonymousShmemResize(void)
*/
pending_pm_shmem_resize = false;
- /*
- * XXX: Currently only increasing of shared_buffers is supported. For
- * decreasing something similar has to be done, but buffer blocks with
- * data have to be drained first.
- */
- if(NBuffersOld > NBuffers)
- return false;
-
#ifndef MAP_HUGETLB
/* PrepareHugePages should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
@@ -1112,11 +1104,14 @@ AnonymousShmemResize(void)
* all the pointers are still valid, and we only need to update
* structures size in the ShmemIndex once -- any other backend
* will pick up this shared structure from the index.
- *
- * XXX: This is the right place for buffer eviction as well.
*/
BufferManagerShmemInit(NBuffersOld);
+ /*
+ * Wipe out the evictor PID so that it can be used for the next
+ * buffer resizing operation.
+ */
+ ShmemCtrl->evictor_pid = 0;
/* If all fine, broadcast the new value */
pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
}
@@ -1169,11 +1164,31 @@ ProcessBarrierShmemResize(Barrier *barrier)
* XXX: If we need to be able to abort resizing, this has to be done later,
* after the SHMEM_RESIZE_DONE.
*/
- if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+
+ /*
+ * Evict extra buffers when shrinking shared buffers. We need to do this
+ * while the memory for extra buffers is still mapped i.e. before remapping
+ * the shared memory segments to a smaller memory area.
+ */
+ if (NBuffersOld > NBuffersPending)
{
- Assert(IsUnderPostmaster);
- SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
+
+ /*
+ * TODO: If the buffer eviction fails for any reason, we should
+ * gracefully rollback the shared buffer resizing and try again. But the
+ * infrastructure to do so is not available right now. Hence just raise
+ * a FATAL so that the system restarts.
+ */
+ if (!EvictExtraBuffers(NBuffersPending, NBuffersOld))
+ elog(FATAL, "buffer eviction failed");
+
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_EVICT))
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
}
+ else
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
AnonymousShmemResize();
@@ -1683,5 +1698,6 @@ ShmemControlInit(void)
/* shmem_resizable should be initialized by now */
ShmemCtrl->Resizable = shmem_resizable;
+ ShmemCtrl->evictor_pid = 0;
}
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 8c1ea623392..6f5743036a2 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -157,8 +157,12 @@ BufferManagerShmemInit(int FirstBufferToInit)
}
#endif
- /* Correct last entry of linked list */
- GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
+ /*
+ * Correct last entry of linked list, when initializing the buffers or when
+ * expanding the buffers.
+ */
+ if (FirstBufferToInit < NBuffers)
+ GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 667aa0c0c78..169a44dd9fc 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/read_stream.h"
#include "storage/smgr.h"
@@ -7453,3 +7454,97 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int newBufSize, int oldBufSize)
+{
+ bool result = true;
+
+ /*
+ * If the buffer being evicated is locked, this function will need to wait.
+ * This function should not be called from a Postmaster since it can not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * Let only one backend perform eviction. We could split the work across all
+ * the backends but that doesn't seem necessary.
+ *
+ * The first backend to acquire ShmemResizeLock, sets its own PID as the
+ * evictor PID for other backends to know that the eviction is in progress or
+ * has already been performed. The evictor backend releases the lock when it
+ * finishes eviction. While the eviction is in progress, backends other than
+ * evictor backend won't be able to take the lock. They won't perform
+ * eviction. A backend may acquire the lock after eviction has completed, but
+ * it will not perform eviction since the evictor PID is already set. Evictor
+ * PID is reset only when the buffer resizing finishes. Thus only one backend
+ * will perform eviction in a given instance of shared buffers resizing.
+ *
+ * Any backend which acquires this lock will release it before the eviction
+ * phase finishes, hence the same lock can be reused for the next phase of
+ * resizing buffers.
+ */
+ if (LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ if (ShmemCtrl->evictor_pid == 0)
+ {
+ ShmemCtrl->evictor_pid = MyProcPid;
+
+ StrategyPurgeFreeList(newBufSize);
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = newBufSize + 1; buf <= oldBufSize; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u32(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is
+ * going one hence unlocked precheck should be safe and saves
+ * some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ /*
+ * XXX: Looks like CurrentResourceOwner can be NULL here, find
+ * another one in that case?
+ * */
+ if (CurrentResourceOwner)
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+ }
+ LWLockRelease(ShmemResizeLock);
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index bd390f2709d..7b9ed010e2f 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -527,6 +527,74 @@ StrategyInitialize(bool init)
}
+/*
+ * StrategyPurgeFreeList -- remove all buffers with id higher than the number of
+ * buffers in the buffer pool.
+ *
+ * This is called before evicting buffers while shrinking shared buffers, so that
+ * the free list does not reference a buffer that will be removed.
+ *
+ * The function is called after resizing has started and thus nobody should be
+ * traversing the free list and also not touching the buffers.
+ */
+void
+StrategyPurgeFreeList(int numBuffers)
+{
+ int firstBuffer = FREENEXT_END_OF_LIST;
+ int nextFree = StrategyControl->firstFreeBuffer;
+ BufferDesc *prevValidBuf = NULL;
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ while (nextFree != FREENEXT_END_OF_LIST)
+ {
+ BufferDesc *buf = GetBufferDescriptor(nextFree);
+
+ /* nextFree should be id of buffer being examined. */
+ Assert(nextFree == buf->buf_id);
+ /* The buffer should not be marked as not in the list. */
+ Assert(buf->freeNext != FREENEXT_NOT_IN_LIST);
+
+ /*
+ * If the buffer is within the new size of pool, keep it in the free list
+ * otherwise discard it.
+ */
+ if (buf->buf_id < numBuffers)
+ {
+ if (prevValidBuf != NULL)
+ prevValidBuf->freeNext = buf->buf_id;
+ prevValidBuf = buf;
+
+ /* Save the first free buffer in the list if not already known. */
+ if (firstBuffer == FREENEXT_NOT_IN_LIST)
+ firstBuffer = nextFree;
+ }
+ /* Examine the next buffer in the free list. */
+ nextFree = buf->freeNext;
+ }
+
+ /* Update the last valid free buffer, if there's any. */
+ if (prevValidBuf != NULL)
+ {
+ StrategyControl->lastFreeBuffer = prevValidBuf->buf_id;
+ prevValidBuf->freeNext = FREENEXT_END_OF_LIST;
+ }
+ else
+ StrategyControl->lastFreeBuffer = FREENEXT_END_OF_LIST;
+
+ /* Update first valid free buffer, if there's any. */
+ StrategyControl->firstFreeBuffer = firstBuffer;
+
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+
+ /*
+ * TODO: following was suggested by AI. Check whether it is required.
+ * If we removed all buffers from the freelist, reset the clock sweep
+ * pointer to zero. This is not strictly necessary, but it seems like a
+ * good idea to avoid confusion.
+ */
+}
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
* ----------------------------------------------------------------
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 691fa14e9e3..0c588b69a90 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -156,6 +156,7 @@ REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it c
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_EVICT "Waiting for other backends to finish buffer evication phase."
SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized."
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 0dec7d93b3b..add15e3723b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -453,6 +453,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyPurgeFreeList(int numBuffers);
extern bool have_free_buffer(void);
/* buf_table.c */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 4239ebe640b..0c554f0b130 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -315,6 +315,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_evicted,
int32 *buffers_flushed,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int fromBuf, int toBuf);
/* in buf_init.c */
extern void BufferManagerShmemInit(int);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index bccdd45b1f7..9a3e68f76d8 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -64,6 +64,7 @@ extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
typedef struct
{
pg_atomic_uint32 NSharedBuffers;
+ pid_t evictor_pid;
Barrier Barrier;
pg_atomic_uint64 Generation;
bool Resizable;
--
2.49.0
[text/plain] v5-0009-Reinitialize-StrategyControl-after-resizing-buffe.patch (16.8K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/10-v5-0009-Reinitialize-StrategyControl-after-resizing-buffe.patch)
download | inline diff:
From 475fffc314b4d5975faf94d43504f2dcc0f1dc8d Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 19 Jun 2025 17:38:51 +0200
Subject: [PATCH v5 09/10] Reinitialize StrategyControl after resizing buffers
... and BgBufferSync and ClockSweepTick adjustments
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. When expanding the buffer pool add new buffers to the free list.
2. When shrinking buffers, we remove any buffers, in the area being
shrunk, from the freelist. While doing so we adjust the first and
last free buffer pointers in the StrategyControl area. Hence nothing
more needed after resizing.
3. Check the sanity of the free buffer list is added after resizing.
4. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
5. &StrategyControl->buffer_strategy_lock need not be initialized again.
6. nextVictimBuffer, completePasses and numBufferAllocs are viewed in
the context of NBuffers. Now that NBuffers itself has changed, those
three do not make sense. Reset them as if the server has restarted
again.
This commit introduces a flag delay_shmem_resize, which postgresql
backends and workers can use to signal the coordinator to delay resizing
operation. Background writer sets this flag when its scanning buffers. Background
writer is blocked when the actual resizing is in progress. But if
resizing is about to begin, it does not scan the buffers by returning
from BgBufferSync(). It stops a scan in progress when it sees that the
resizing has begun. After the resizing is finished, it adjusts the
collected statistics according to the new size of the buffer pool at the
end of barrier processing.
Once the buffer resizing is finished, before resuming the regular
operation, bgwriter resets the information saved so far. This
information is viewed in the context of NBuffers and hence does not make
sense after NBuffer has changed.
Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 24 ++++-
src/backend/postmaster/bgwriter.c | 2 +-
src/backend/storage/buffer/buf_init.c | 11 ++-
src/backend/storage/buffer/bufmgr.c | 74 ++++++++++----
src/backend/storage/buffer/freelist.c | 133 ++++++++++++++++++++++++++
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 3 +-
src/include/storage/ipc.h | 1 +
8 files changed, 226 insertions(+), 23 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index e612a83c9f0..e61a557966b 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -115,6 +115,7 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Flag telling postmaster that resize is needed */
volatile bool pending_pm_shmem_resize = false;
+volatile bool delay_shmem_resize = false;
/* Keeps track of the previous NBuffers value */
static int NBuffersOld = -1;
@@ -1118,6 +1119,12 @@ AnonymousShmemResize(void)
LWLockRelease(ShmemResizeLock);
}
+
+ /*
+ * TODO: Shouldn't we call ResizeBufferPool() here as well? Or those
+ * backend who can not lock the LWLock conditionally won't resize the
+ * buffers.
+ */
}
return true;
@@ -1138,13 +1145,17 @@ ProcessBarrierShmemResize(Barrier *barrier)
{
Assert(IsUnderPostmaster);
- elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
- NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize, delay_shmem_resize);
/* Wait until we have seen the new NBuffers value */
if (!pending_pm_shmem_resize)
return false;
+ /* Wait till this process becomes ready to resize buffers. */
+ if (delay_shmem_resize)
+ return false;
+
/*
* First thing to do after attaching to the barrier is to wait for others.
* We can't simply use BarrierArriveAndWait, because backends might arrive
@@ -1195,6 +1206,15 @@ ProcessBarrierShmemResize(Barrier *barrier)
/* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * Before resuming regular background writer activity, adjust the
+ * statistics collected so far.
+ */
+ BgBufferSyncAdjust(NBuffersOld, NBuffers);
+ }
+
BarrierDetach(barrier);
return true;
}
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index 72f5acceec7..32b34f28ead 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -233,7 +233,7 @@ BackgroundWriterMain(const void *startup_data, size_t startup_data_len)
/*
* Do one cycle of dirty-buffer writing.
*/
- can_hibernate = BgBufferSync(&wb_context);
+ can_hibernate = BgBufferSync(&wb_context, false);
/* Report pending statistics to the cumulative stats system */
pgstat_report_bgwriter();
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6f5743036a2..62c1672fb70 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -164,8 +164,15 @@ BufferManagerShmemInit(int FirstBufferToInit)
if (FirstBufferToInit < NBuffers)
GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST;
- /* Init other shared buffer-management stuff */
- StrategyInitialize(!foundDescs);
+ /*
+ * Init other shared buffer-management stuff from scratch configuring buffer
+ * pool the first time. If we are just resizing buffer pool adjust only the
+ * required structures.
+ */
+ if (FirstBufferToInit == 0)
+ StrategyInitialize(!foundDescs);
+ else
+ StrategyReInitialize(FirstBufferToInit);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 169a44dd9fc..590fc737da8 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3611,6 +3611,32 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_DONE(NBuffers, num_written, num_to_scan);
}
+/*
+ * Information saved between BgBufferSync() calls so we can determine the
+ * strategy point's advance rate and avoid scanning already-cleaned buffers. The
+ * variables are global instead of static local so that BgBufferSyncAdjust() can
+ * adjust it when resizing shared buffers.
+ */
+static bool saved_info_valid = false;
+static int prev_strategy_buf_id;
+static uint32 prev_strategy_passes;
+static int next_to_clean;
+static uint32 next_passes;
+
+/* Moving averages of allocation rate and clean-buffer density */
+static float smoothed_alloc = 0;
+static float smoothed_density = 10.0;
+
+void
+BgBufferSyncAdjust(int NBuffersOld, int NBuffersNew)
+{
+ saved_info_valid = false;
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
+ NBuffersOld, NBuffersNew);
+#endif
+}
+
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
@@ -3623,27 +3649,13 @@ BufferSync(int flags)
* bgwriter_lru_maxpages to 0.)
*/
bool
-BgBufferSync(WritebackContext *wb_context)
+BgBufferSync(WritebackContext *wb_context, bool reset)
{
/* info obtained from freelist.c */
int strategy_buf_id;
uint32 strategy_passes;
uint32 recent_alloc;
- /*
- * Information saved between calls so we can determine the strategy
- * point's advance rate and avoid scanning already-cleaned buffers.
- */
- static bool saved_info_valid = false;
- static int prev_strategy_buf_id;
- static uint32 prev_strategy_passes;
- static int next_to_clean;
- static uint32 next_passes;
-
- /* Moving averages of allocation rate and clean-buffer density */
- static float smoothed_alloc = 0;
- static float smoothed_density = 10.0;
-
/* Potentially these could be tunables, but for now, not */
float smoothing_samples = 16;
float scan_whole_pool_milliseconds = 120000.0;
@@ -3666,6 +3678,22 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ /*
+ * If buffer pool is being shrunk the buffer being written out may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing out any. Hence wait till
+ * buffer resizing finishes i.e. go into hibernation mode.
+ */
+ if (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ return true;
+
+ /*
+ * Resizing shared buffers while this function is performing an LRU scan on
+ * them may lead to wrong results. Indicate that the resizing should wait for
+ * the LRU scan to complete.
+ */
+ delay_shmem_resize = true;
+
/*
* Find out where the freelist clock sweep currently is, and how many
* buffer allocations have happened since our last call.
@@ -3842,8 +3870,17 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
- while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
+ /*
+ * Execute the LRU scan.
+ *
+ * If buffer pool is being shrunk, the buffer being written may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing any. Hence stop what we are doing. This
+ * also unblocks other processes that are waiting for buffer resizing to
+ * finish.
+ */
+ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
+ pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) == NBuffers)
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -3902,6 +3939,9 @@ BgBufferSync(WritebackContext *wb_context)
#endif
}
+ /* Let the resizing commence. */
+ delay_shmem_resize = false;
+
/* Return true if OK to hibernate */
return (bufs_to_lap == 0 && recent_alloc == 0);
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 7b9ed010e2f..41641bb3ae6 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -98,6 +98,9 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
uint32 *buf_state);
static void AddBufferToRing(BufferAccessStrategy strategy,
BufferDesc *buf);
+#ifdef USE_ASSERT_CHECKING
+static void StrategyValidateFreeList(void);
+#endif /* USE_ASSERT_CHECKING */
/*
* ClockSweepTick - Helper routine for StrategyGetBuffer()
@@ -526,6 +529,88 @@ StrategyInitialize(bool init)
Assert(!init);
}
+/*
+ * StrategyReInitialize -- re-initialize the buffer cache replacement
+ * strategy.
+ *
+ * To be called when resizing buffer manager and only from the coordinator.
+ * TODO: Assess the differences between this function and StrategyInitialize().
+ */
+void
+StrategyReInitialize(int FirstBufferIdToInit)
+{
+ bool found;
+
+ /*
+ * Resizing memory for buffer pools should not affect the address of
+ * StrategyControl.
+ */
+ if (StrategyControl != (BufferStrategyControl *)
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, STRATEGY_SHMEM_SEGMENT))
+ elog(FATAL, "something went wrong while re-initializing the buffer strategy");
+
+ Assert(found);
+
+ /* TODO: Buffer lookup table adjustment: There are two options:
+ *
+ * 1. Resize the buffer lookup table to match the new number of buffers. But
+ * this requires rehashing all the entries in the buffer lookup table with
+ * the new table size.
+ *
+ * 2. Allocate maximum size of the buffer lookup table at the beginning and
+ * never resize it. This leaves sparse buffer lookup table which is
+ * inefficient from both memory and time perspective. According to David
+ * Rowley, the sparse entries in the buffer look up table cause frequent
+ * cacheline reload which affect performance. If the impact of that
+ * inefficiency in a benchmark is significant, we will need to consider first
+ * option.
+ */
+
+ /*
+ * When shrinking buffers, we must have adjusted the first and the last free
+ * buffer when removing the buffers being shrunk from the free list. Nothing
+ * to be done here.
+ *
+ * When expanding the shared buffers, new buffers are added at the end of the
+ * freelist or they form the new free list if there are no free buffers.
+ */
+ if (FirstBufferIdToInit < NBuffers)
+ {
+ if (StrategyControl->firstFreeBuffer == FREENEXT_END_OF_LIST)
+ StrategyControl->firstFreeBuffer = FirstBufferIdToInit;
+ else
+ {
+ Assert(StrategyControl->lastFreeBuffer >= 0);
+ GetBufferDescriptor(StrategyControl->lastFreeBuffer - 1)->freeNext = FirstBufferIdToInit;
+ }
+
+ StrategyControl->lastFreeBuffer = NBuffers - 1;
+ }
+
+ /* Check free list sanity after resizing. */
+#ifdef USE_ASSERT_CHECKING
+ StrategyValidateFreeList();
+#endif /* USE_ASSERT_CHECKING */
+
+ /*
+ * The clock sweep tick pointer might have got invalidated. Reset it as if
+ * starting a fresh server.
+ */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The old statistics is viewed in the context of the number of shared
+ * buffers. It does not make sense now that the number of shared buffers
+ * itself has changed.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_init_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* No pending notification */
+ StrategyControl->bgwprocno = -1;
+}
/*
* StrategyPurgeFreeList -- remove all buffers with id higher than the number of
@@ -595,6 +680,54 @@ StrategyPurgeFreeList(int numBuffers)
*/
}
+#ifdef USE_ASSERT_CHECKING
+/*
+ * StrategyValidateFreeList-- check sanity of free buffer list.
+ */
+static void
+StrategyValidateFreeList(void)
+{
+ int nextFree = StrategyControl->firstFreeBuffer;
+ int numFreeBuffers = 0;
+ int lastFreeBuffer = FREENEXT_END_OF_LIST;
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ while (nextFree != FREENEXT_END_OF_LIST)
+ {
+ BufferDesc *buf = GetBufferDescriptor(nextFree);
+
+ /* nextFree should be id of buffer being examined. */
+ Assert(nextFree == buf->buf_id);
+ Assert(buf->buf_id < NBuffers);
+ /* The buffer should not be marked as not in the list. */
+ Assert(buf->freeNext != FREENEXT_NOT_IN_LIST);
+
+ /* Update our knowledge of last buffer in the free list. */
+ lastFreeBuffer = buf->buf_id;
+
+ numFreeBuffers++;
+
+ /* Avoid infinite recursion in case there are cycles in free list. */
+ if (numFreeBuffers > NBuffers)
+ break;
+
+ nextFree = buf->freeNext;
+ }
+
+ Assert(numFreeBuffers <= NBuffers);
+
+ /*
+ * Make sure that the StrategyControl's knowledge of last free buffer
+ * agrees with what's there in the free list.
+ */
+ if (StrategyControl->firstFreeBuffer != FREENEXT_END_OF_LIST)
+ Assert(StrategyControl->lastFreeBuffer == lastFreeBuffer);
+
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+}
+#endif /* USE_ASSERT_CHECKING */
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
* ----------------------------------------------------------------
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index add15e3723b..46949e9d90e 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -454,6 +454,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
extern void StrategyPurgeFreeList(int numBuffers);
+extern void StrategyReInitialize(int FirstBufferToInit);
extern bool have_free_buffer(void);
/* buf_table.c */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 0c554f0b130..83a75eab844 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -298,7 +298,8 @@ extern bool ConditionalLockBufferForCleanup(Buffer buffer);
extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
-extern bool BgBufferSync(struct WritebackContext *wb_context);
+extern bool BgBufferSync(struct WritebackContext *wb_context, bool reset);
+extern void BgBufferSyncAdjust(int NBuffersOld, int NBuffersNew);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index bb7ae4d33b3..7d1c64a9267 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -65,6 +65,7 @@ typedef void (*shmem_startup_hook_type) (void);
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
+extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
--
2.49.0
[text/plain] v5-0010-Additional-validation-for-buffer-in-the-ring.patch (2.1K, ../../my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov2/11-v5-0010-Additional-validation-for-buffer-in-the-ring.patch)
download | inline diff:
From 25816411c7a77a87b8527a901633865f1c1056b7 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 11 Jun 2025 18:15:06 +0530
Subject: [PATCH v5 10/10] Additional validation for buffer in the ring
If the buffer pool has been shrunk, the buffers in the buffer list may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition. That may not be
expensive enough to affect query performance. The alternative to that is
more complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure thus accessible when processing buffer
resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work.
Ashutosh Bapat
---
src/backend/storage/buffer/freelist.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 41641bb3ae6..74d070733a4 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -948,12 +948,13 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot has been
+ * invalidated when buffer pool was shrunk, tell the caller to allocate a new
+ * buffer with the normal allocation strategy. He will then fill this slot
+ * by calling AddBufferToRing with the new buffer.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer || bufnum > NBuffers)
return NULL;
/*
--
2.49.0
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-06-20 10:22 ` Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-06-20 10:22 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Fri, Jun 20, 2025 at 12:19:31PM +0200, Dmitry Dolgov wrote:
> Thanks! I've reworked the series to implement approach suggested by
> Thomas, and applied your patches to support buffers shrinking on top. I
> had to restructure the patch set, here is how it looks like right now:
The base-commit was left in the cover letter which I didn't post, so for
posterity:
base-commit: 4464fddf7b50abe3dbb462f76fd925e10eedad1c
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-02 12:35 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
1 sibling, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-07-02 12:35 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi Dmitry,
Thanks for sharing the patches.
On Fri, Jun 20, 2025 at 3:49 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> 3. Shared memory shrinking
>
> So far only shared memory increase was implemented. These patches from Ashutosh
> support shrinking as well, which is tricky due to the need for buffer eviction.
>
> * Support shrinking shared buffers
> * Reinitialize StrategyControl after resizing buffers
This applies to both shrinking and expansion of shared buffers. When
expanding we need to add the new buffers to the freelist by changing
next pointer of last buffer in the free list to point to the first new
buffer.
> * Additional validation for buffer in the ring
>
> > 0009 adds support to shrink shared buffers. It has two changes: a.
> > evict the buffers outside the new buffer size b. remove buffers with
> > buffer id outside the new buffer size from the free list. If a buffer
> > being evicted is pinned, the operation is aborted and a FATAL error is
> > raised. I think we need to change this behaviour to be less severe
> > like rolling back the operation or waiting for the pinned buffer to be
> > unpinned etc. Better even if we could let users control the behaviour.
> > But we need better infrastructure to do such things. That's one TODO
> > left in the patch.
>
> I haven't reviewed those, just tested a bit to finally include into the series.
> Note that I had to tweak two things:
>
> * The way it was originally implemented was sending resize signal to postmaster
> before doing eviction, which could result in sigbus when accessing LSN of a
> dirty buffer to be evicted. I've reshaped it a bit to make sure eviction always
> happens first.
Will take a look at this.
>
> * It seems the CurrentResource owner could be missing sometimes, so I've added
> a band-aid checking its presence.
>
> One side note, during my testing I've noticed assert failures on
> pgstat_tracks_io_op inside a wal writer a few times. I couldn't reproduce it
> after the fixes above, but still it may indicate that something is off. E.g.
> it's somehow not expected that the wal writer will do buffer eviction IO (from
> what I understand, the current shrinking implementation allows that).
Yes. I think, we have to find a better way to choose a backend which
does the actual work. Eviction can be done in that backend itself.
Compiler gives warning about an uninitialized variable, which seems to
be a real bug. Fix attached.
--
Best Wishes,
Ashutosh Bapat
commit 5c54d36d28edcdd72edd6374132ac88e1693fcaa
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed Jul 2 15:47:45 2025 +0530
fixup! Allow to use multiple shared memory mappings
shm_total_page_count is used unitialized. If this variable has a random
value to start with, the final sum would be wrong.
Ashutosh Bapat
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9dd22920a79..24aec528c3c 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -658,7 +658,7 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size os_page_size;
void **page_ptrs;
int *pages_status;
- uint64 shm_total_page_count,
+ uint64 shm_total_page_count = 0,
shm_ent_page_count,
max_nodes;
Size *nodes;
Attachments:
[text/plain] unitialized_variable_warning.patch.txt (815B, ../../CAExHW5sEfnpMvCe_NRY8wRO41b1Satx+mNYzwzkLsJkmegTFwA@mail.gmail.com/2-unitialized_variable_warning.patch.txt)
download | inline diff:
commit 5c54d36d28edcdd72edd6374132ac88e1693fcaa
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed Jul 2 15:47:45 2025 +0530
fixup! Allow to use multiple shared memory mappings
shm_total_page_count is used unitialized. If this variable has a random
value to start with, the final sum would be wrong.
Ashutosh Bapat
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9dd22920a79..24aec528c3c 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -658,7 +658,7 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size os_page_size;
void **page_ptrs;
int *pages_status;
- uint64 shm_total_page_count,
+ uint64 shm_total_page_count = 0,
shm_ent_page_count,
max_nodes;
Size *nodes;
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-07-04 00:06 ` Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-18 04:47 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Tomas Vondra @ 2025-07-04 00:06 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi Ashutosh dn Dmitry,
I took a look at this patch, because it's somewhat related to the NUMA
patch series I posted a couple days ago, and I've been wondering if
it makes some of the NUMA stuff harder or simpler.
I don't think it makes a bit difference (for the NUMA stuff). My main
question was when would we adjust the "NUMA location" of parts of memory
to keep stuff balanced, but this patch series already needs to update
some of these structs (like the freelists), so those places would be
updated to be NUMA-aware. Some of the changes could be made lazily,
to minimize the amount of time when activity is stopped (like shifting
the buffers to different NUMA nodes). It'd be harder if we wanted to
resize e.g. PGPROC, but that's not the case. So I think this is fine.
I agree it'd be useful to be able to resize shared buffers, without
having to restart the instance (which is obviously very disruptive). So
if we can make this work reliably, with reasonable trade offs (both on
the backends, and also the risks/complexity introduced by the feature).
I'm far from an expert on mmap() and similar low-level stuff, but the
current appproach (reserving a big chunk of shared memory and slicing
it by mmap() into smaller segments) seems reasonable.
But I'm getting a bit lost in how exactly this interacts with things
like overcommit, system memory accounting / OOM killer and this sort of
stuff. I went through the thread and it seems to me the reserve+map
approach works OK in this regard (and the messages on linux-mm seem to
confirm this). But this information is scattered over many messages and
it's hard to say for sure, because some of this might be relevant for
an earlier approach, or a subtly different variant of it.
A similar question is portability. The comments and commit messages
seem to suggest most of this is linux-specific, and other platforms just
don't have these capabilities. But there's a bunch of messages (mostly
by Thomas Munro) that hint FreeBSD might be capable of this too, even if
to some limited extent. And possibly even Windows/EXEC_BACKEND, although
that seems much trickier.
FWIW I think it's perfectly fine to only support resizing on selected
platforms, especially considering Linux is the most widely used system
for running Postgres. We still need to be able to build/run on other
systems, of course. And maybe it'd be good to be able to disable this
even on Linux, if that eliminates some overhead and/or risks for people
who don't need the feature. Just a thought.
Anyway, my main point is that this information is important, but very
scattered over the thread. It's a bit foolish to expect everyone who
wants to do a review to read the whole thread (which will inevitably
grow longer over time), and assemble all these pieces again an again,
following all the changes in the design etc. Few people will get over
that hurdle, IMHO.
So I think it'd be very helpful to write a README, explaining the
currnent design/approach, and summarizing all these aspects in a single
place. Including things like portability, interaction with the OS
accounting, OOM killer, this kind of stuff. Some of this stuff may be
already mentioned in code comments, but you it's hard to find those.
Especially worth documenting are the states the processes need to go
through (using the barriers), and the transacitons between them (i.e.
what is allowed in each phase, what blocks can be visible, etc.).
I'll go over some higher-level items first, and then over some comments
for individual patches.
1) no user docs
There are no user .sgml docs, and maybe it's time to write some,
explaining how to use this thing - how to configure it, how to trigger
the resizing, etc. It took me a while to realize I need to do ALTER
SYSTEM + pg_reload_conf() to kick this off.
It should also document the user-visible limitations, e.g. what activity
is blocked during the resizing, etc.
2) pending GUC changes
I'm somewhat skeptical about the GUC approach. I don't think it was
designed with this kind of use case in mind, and so I think it's quite
likely it won't be able to handle it well.
For example, there's almost no validation of the values, so how do you
ensure the new value makes sense? Because if it doesn't, it can easily
crash the system (I've seen such crashes repeatedly, I'll get to that).
Sure, you may do ALTER SYSTEM to set shared_buffers to nonsense and it
won't start after restart/reboot, but crashing an instance is maybe a
little bit more annoying.
Let's say we did the ALTER SYSTEM + pg_reload_conf(), and it gets stuck
waiting on something (can't evict a buffer or something). How do you
cancel it, when the change is already written to the .auto.conf file?
Can you simply do ALTER SYSTEM + pg_reload_conf() again?
It also seems a bit strange that the "switch" gets to be be driven by a
randomly selected backend (unless I'm misunderstanding this bit). It
seems to be true for the buffer eviction during shrinking, at least.
Perhaps this should be a separate utility command, or maybe even just
a new ALTER SYSTEM variant? Or even just a function, similar to what
the "online checksums" patch did, possibly combined with a bgworder
(but probably not needed, there are no db-specific tasks to do).
3) max_available_memory
Speaking of GUCs, I dislike how max_available_memory works. It seems a
bit backwards to me. I mean, we're specifying shared_buffers (and some
other parameters), and the system calculates the amount of shared memory
needed. But the limit determines the total limit?
I think the GUC should specify the maximum shared_buffers we want to
allow, and then we'd work out the total to pre-allocate? Considering
we're only allowing to resize shared_buffers, that should be pretty
trivial. Yes, it might happen that the "total limit" happens to exceed
the available memory or something, but we already have the problem
with shared_buffers. Seems fine if we explain this in the docs, and
perhaps print the calculated memory limit on start.
In any case, we should not allow setting a value that ends up
overflowing the internal reserved space. It's true we don't have a good
way to do checks for GUcs, but it's a bit silly to crash because of
hitting some non-obvious internal limit that we necessarily know about.
Maybe this is a reason why GUC hooks are not a good way to set this.
4) SHMEM_RESIZE_RATIO
The SHMEM_RESIZE_RATIO thing seems a bit strange too. There's no way
these ratios can make sense. For example, BLCKSZ is 8192 but the buffer
descriptor is 64B. That's 128x difference, but the ratios says 0.6 and
0.1, so 6x. Sure, we'll actually allocate only the memory we need, and
the rest is only "reserved".
However, that just makes the max_available_memory a bit misleading,
because you can't ever use it. You can use the 60% for shared buffers
(which is not mentioned anywhere, and good luck not overflowing that,
as it's never checked), but those smaller regions are guaranteed to be
mostly unused. Unfortunate.
And it's not just a matter of fixing those ratios, because then someone
rebuilds with 32kB blocks and you're in the same situation.
Moreover, all of the above is for mappings sized based on NBuffers. But
if we allocate 10% for MAIN_SHMEM_SEGMENT, won't that be a problem the
moment someone increases of max_connection, max_locks_per_transaction
and possibly some other stuff?
5) no tests
I mentioned no "user docs", but the patch has 0 tests too. Which seems
a bit strange for a patch of this age.
A really serious part of the patch series seems to be the coordination
of processes when going through the phases, enforced by the barriers.
This seems like a perfect match for testing using injection points, and
I know we did something like this in the online checksums patch, which
needs to coordinate processes in a similar way.
But even just a simple TAP test that does a bunch of (random?) resizes
while running a pgbench seem better than no tests. (That's what I did
manually, and it crashed right away.)
There's a lot more stuff to test here, I think. Idle sessions with
buffers pinned by open cursors, multiple backends doing ALTER SYSTEM
+ pg_reload_conf concurrently, other kinds of failures.
6) SIGBUS failures
As mentioned, I did some simple tests with shrink/resize with a pgbench
in the background, and it almost immediately crashed for me :-( With a
SIGBUS, which I think is fairly rare on x86 (definitely much less common
than e.g. SIGSEGV).
An example backtrace attached.
7) EXEC_BACKEND, FreeBSD
We clearly need to keep this working on systems without the necessary
bits (so likely EXEC_BACKEND, FreeBSD etc.). But the builds currently
fail in both cases, it seems.
I think it's fine to not support resizing on every platform, then we'd
never get it, but it still needs to build. It would be good to not have
two very different code versions, one for resizing and one without it,
though. I wonder if we can just have the "no-resize" use the same struct
(with the segments/mapping, ...) and all that, but skipping the space
reservation.
8) monitoring
So, let's say I start a resize of shared buffers. How will I know what
it's currently doing, how much longer it might take, what it's waiting
for, etc.? I think it'd be good to have progress monitoring, through
the regular system view (e.g. pg_stat_shmem_resize_progress?).
10) what to do about stuck resize?
AFAICS the resize can get stuck for various reasons, e.g. because it
can't evict pinned buffers, possibly indefinitely. Not great, it's not
clear to me if there's a way out (canceling the resize) after a timeout,
or something like that? Not great to start an "online resize" only to
get stuck with all activity blocked for indefinite amount of time, and
get to restart anyway.
Seems related to Thomas' message [2], but AFAICS the patch does not do
anything about this yet, right? What's the plan here?
11) preparatory actions?
Even if it doesn't get stuck, some of the actions can take a while, like
evicting dirty buffers before shrinking, etc. This is similar to what
happens on restart, when the shutdown checkpoint can take a while, while
the system is (partly) unavailable.
The common mitigation is to do an explicit checkpoint right before the
restart, to make the shutdown checkpoint cheap. Could we do something
similar for the shrinking, e.g. flush buffers from the part to be
removed before actually starting the resize?
12) does this affect e.g. fork() costs?
I wonder if this affects the cost of fork() in some undesirable way?
Could it make fork() more measurably more expensive?
13) resize "state" is all over the place
For me, a big hurdle when reasoning about the resizing correctness is
that there's quite a lot of distinct pieces tracking what the current
"state" is. I mean, there's:
- ShmemCtrl->NSharedBuffers
- NBuffers
- NBuffersOld
- NBuffersPending
- ... (I'm sure I missed something)
There's no cohesive description how this fits together, it seems a bit
"ad hoc". Could be correct, but I find it hard to reason about.
14) interesting messages from the thread
While reading through the thread, I noticed a couple messages that I
think are still relevant:
- I see Peter E posted some review in 2024/11 [3], but it seems his
comments were mostly ignored. I agree with most of them.
- Robert mentioned a couple interesting failure scenarios in [4], not
sure if all of this was handled. He howerver assumes pointers would
not be stable (and that's something we should not allow, and the
current approach works OK in this regard, I think). He also outlines
how it'd happen in phases - this would be useful for the design README
I think. It also reminds me the "phases" in the checksums patch.
- Robert asked [5] if Linux might abruptly break this, but I find that
unlikely. We'd point out we rely on this, and they'd likely rethink.
This would be made safer if this was specified by POSIX - taking that
away once implemented seems way harder than for custom extensions.
It's likely they'd not take away the feature without an alternative
way to achieve the same effect, I think (yes, harder to maintain).
Tom suggests [7] this is not in POSIX.
- Matthias mentioned [6] similar flags on other operating systems. Could
some of those be used to implement the same resizing?
- Andres had an interesting comment about how overcommit interacts with
MAP_NORESERVE. AFAIK it means we need the flag to not break overcommit
accounting. There's also some comments about from linux-mm people [9].
- There seem to be some issues with releasing memory backing a mapping
with hugetlb [10]. With the fd (and truncating the file), this seems
to release the memory, but it's linux-specific? But most of this stuff
is specific to linux, it seems. So is this a problem? With this it
should be working even for hugetlb ...
- It seems FreeBSD has MFD_HUGETLB [11], so maybe we could use this and
make the hugetlb stuff work just like on Linux? Unclear. Also, I
thought the mfd stuff is linux-specific ... or am I confused?
- Andres objected to any approach without pointer stability, and I agree
with that. If we can figure out such solution, of course.
- Thomas asked [13] why we need to stop all the backends, instead of
just waiting for them to acknowledge the new (smaller) NBuffers value
and then let them continue. I also don't quite see why this should
not work, and it'd limit the disruption when we have to wait for
eviction of buffers pinned by paused cursors, etc.
Now, some comments about the individual patches (some of this may be a
bit redundant with the earlier points):
v5-0001-Process-config-reload-in-AIO-workers.patch
1) Hmmm, so which other workers may need such explicit handling? Do all
other processes participate in procsignal stuff, or does anything
need an explicit handling?
v5-0002-Introduce-pending-flag-for-GUC-assign-hooks.patch
No additional comments, see the points about resizing through a GUC
callback with pending flag vs. a separate utility command, monitoring
and so on.
v5-0003-Introduce-pss_barrierReceivedGeneration.patch
1) Do we actually need this? Isn't it enough to just have two barriers?
Or a barrier + condition variable, or something like that.
2) The comment talks about "coordinated way" when processing messages,
but it's not very clear to me. It should explain what is needed and
not possible with the current barrier code.
3) This very much reminds me what the online checksums patch needed to
do, and we managed to do it using plain barriers. So why does this
need this new thing? (No opinion on whether it's correct.)
v5-0004-Allow-to-use-multiple-shared-memory-mappings.patch
1) "int shmem_segment" - wouldn't it be better to have a separate enum
for this? I mean, we'll have a predefined list of segments, right?
2) typedef struct AnonymousMapping would deserve some comment
3) ANON_MAPPINGS - Probably should be MAX_ANON_MAPPINGS? But we'll know
how many we have, so why not to allocate exactly the right number?
Or even just an array of structs, like in similar cases?
4) static int next_free_segment = 0;
We exactly know what segments we'll create and in which order, no? So
why do we even bother with this next_free_segment thing? Can't we
simply declare an array of AnonymousMapping elements, with all the
elements, and then just walk it and calculate the sizes/pointers?
5) I'm a bit confused about the segment/mapping difference. The patch
seems to randomly mix those, or maybe I'm just confused. I mean,
we are creating just shmem segment, and the pieces are mappings,
right? So why do we index them by "shmem_segment"?
Also, consider
CreateAnonymousSegment(AnonymousMapping *mapping)
so is that creating a segment or mapping? Or what's the difference?
Or are we creating multiple segments, and I missed that? Or are there
different "segment" concepts, or what?
6) There should probably be some sort of API wrapping the mappings, so
that the various places don't need to mess with next_free_segments
directly, etc. Perhaps PGSharedMemoryCreate() shouldn't do this, and
should just pass size to CreateAnonymousSegment(), and that finding
empty slot in Mappings, etc.? Not sure that'll work, but it's a bit
error-prone if a struct is modified from multiple places like this.
7) We should remember which segments got to use huge pages and which
did not. And we should make it optional for each segment. Although,
maybe I'm just confused about the "segment" definition - if we only
have one, that's where huge pages are applied.
If we could have multiple segments for different segments (whatever
that means), not sure what we'll report for cases when some segments
get to use huge pages and others don't. Either because we don't want
to use that for some segments, or because we happen to run out of
the available huge pages.
8) It seems PGSharedMemoryDetach got some significant changes, but the
comment was not modified at all. I'd guess that means the comment is
perhaps stale, or maybe there's something we should mention.
9) I doubt the Assert on GetConfigOption needs to be repeated for all
segments (in CreateSharedMemoryAndSemaphores).
10) Why do we have the Mapping and Segments indexed in different ways?
I mean, Mappings seem to be filled in FIFO (just grab the next free
slot), while Segments are indexed by segment ID.
11) Actually, what's the difference between the contents of Mappings
and Segments? Isn't that the same thing, indexed in the same way?
Or could it be unified? Or are they conceptually different thing?
12) I believe we'll have a predefined list of segments, with fixed IDs,
so why not just have a MAX of those IDs as the capacity?
13) Would it be good to have some checks on shmem_segment values? That
it's valid with respect to defined segments, etc. An assert, maybe?
What about some asserts on the Mapping/Segment elements? To check
that the element is sensible, and that the arrays "match" (if we
need both).
14) Some of the lines got pretty long, e.g. in pg_get_shmem_allocations.
I suggest we define some macros to make this shorter, or something
like that.
15) I'd maybe rename ShmemSegment to PGShmemSegment, for consistency
with PGShmemHeader?
16) Is MAIN_SHMEM_SEGMENT something we want to expose in a public header
file? Seems very much like an internal thing, people should access
it only through APIs ...
v5-0005-Address-space-reservation-for-shared-memory.patch
1) Shouldn't reserved_offset and huge_pages_on really be in the segment
info? Or maybe even in mapping info? (again, maybe I'm confused
about what these structs store)
2) CreateSharedMemoryAndSemaphores comment is rather light on what it
does, considering it now reserves space and then carves is into
segments.
3) So ReserveAnonymousMemory is what makes decisions about huge pages,
for the whole reserved space / all segments in it. That's a bit
unfortunate with respect to the desirability of some segments
benefiting from huge pages and others not. Maybe we should have two
"reserved" areas, one with huge pages, one without?
I guess we don't want too many segments, because that might make
fork() more expensive, etc. Just guessing, though. Also, how would
this work with threading?
4) Any particular reason to define max_available_memory as
GUC_UNIT_BLOCKS and not GUC_UNIT_MB? Of course, if we change this
to have "max shared buffers limit" then it'd make sense to use
blocks, but "total limit" is not in blocks.
5) The general approach seems sound to me, but I'm not expert on this.
I wonder how portable this behavior is. I mean, will it work on other
Unix systems / Windows? Is it POSIX or Linux extension?
6) It might be a good idea to have Assert procedures to chech mappings
and segments (that it doesn't overflow reserved space, etc.). It
took me ages to realize I can change shared_buffers to >60% of the
limit, it'll happily oblige and then just crash with OOM when
calling mprotect().
v5-0006-Introduce-multiple-shmem-segments-for-shared-buff.patch
1) I suspect the SHMEM_RESIZE_RATIO is the wrong direction, because it
entirely ignores relationships between the parts. See the earlier
comment about this.
2) In fact, what happens if the user tries to resize to a value that is
too large for one of the segments? How would the system know before
starting the resize (and failing)?
3) It seems wrong to modify the BufferManagerShmemSize like this. It's
probably better to have a "...SegmentSize" function for individual
segments, and let BufferManagerShmemSize() to still return a sum of
all segments.
4) I think MaxAvailableMemory is the wrong abstraction, because that's
not what people specify. See earlier comment.
5) Let's say we change the shared memory size (ALTER SYSTEM), trigger
the config reload (pg_reload_conf). But then we find that we can't
actually shrink the buffers, for some unpredictable reason (e.g.
there's pinned buffers). How do we "undo" the change? We can't
really undo the ALTER SYSTEM, that's already written in the .conf
and we don't know the old value, IIRC. Is it reasonable to start
killing backends from the assign_hook or something? Seems weird.
v5-0007-Allow-to-resize-shared-memory-without-restart.patch
1) Why would AdjustShmemSize be needed? Isn't that a sign of a bug
somewhere in the resizing?
2) Isn't the pg_memory_barrier() in CoordinateShmemResize a bit weird?
Why is it needed, exactly? If it's to flush stuff for processes
consuming EmitProcSignalBarrier, it's that too late? What if a
process consumes the barrier between the emit and memory barrier?
3) WaitOnShmemBarrier seem a bit under-documented.
4) Is this actually adding buffers to the freelist? I see buf_init only
links the new buffers by seeting freeNext, but where are the new
buffers added to the existing freelist?
5) The issue with a new backend seeing an old NBuffers value reminds me
of the "support enabling checksums online" thread, where we ran into
similar race conditions. See message [1], the part about race #2
(the other race might be relevant too, not sure). It's been a while,
but I think our conclusion ini that thread was that the "best" fix
would be to change the order of steps in InitPostgres(), i.e. setup
the ProcSignal stuff first, and only then "copy" the NBuffers value.
And handle the possibility that we receive a "duplicate" barriers.
6) In fact, the online checksums thread seems like a possible source of
inspiration for some of the issues, because it needs to do similar
stuff (e.g. make sure all backends follow steps in a synchronized
way, etc.). And it didn't need new types of Barrier to do that.
7) Also, this seems like a perfect match for testing using injection
points. In fact, there's not a single test in the whole patch series.
Or a single line of .sgml docs, for that matter. It took me a while
to realize I'm supposed to change the size by ALTER SYSTEM + reload
the config.
v5-0008-Support-shrinking-shared-buffers.patch
1) Why is ShmemCtrl->evictor_pid reset in AnonymousShmemResize? Isn't
there a place starting it and waiting for it to complete? Why
shouldn't it do EvictExtraBuffers itself?
2) Isn't the change to BufferManagerShmemInit wrong? How do we know the
last buffer is still at the end of the freelist? Seems unlikely.
3) Seems a bit strange to do it from a random backend. Shouldn't it
be the responsibility of a process like checkpointer/bgwriter, or
maybe a dedicated dynamic bgworker? Can we even rely on a backend
to be available?
4) Unsolved issues with buffers pinned for a long time. Could be an
issue if the buffer is pinned indefinitely (e.g. cursor in idle
connection), and the resizing blocks some activity (new connections
or stuff like that).
5) Funny that "AI suggests" something, but doesn't the block fail to
reset nextVictimBuffer of the clocksweep? It may point to a buffer
we're removing, and it'll be invalid, no?
6) It's not clear to me in what situations this triggers (in the call
to BufferManagerShmemInit)
if (FirstBufferToInit < NBuffers) ...
v5-0009-Reinitialize-StrategyControl-after-resizing-buffe.patch
1) IMHO this should be included in the earlier resize/shrink patches,
I don't see a reason to keep it separate (assuming this is the
correct way, and the "init" is not).
2) Doesn't StrategyPurgeFreeList already do some of this for the case
of shrinking memory?
3) Not great adding a bunch of static variables to bufmgr.c. Why do we
need to make "everything" static global? Isn't it enough to make
only the "valid" flag global? The rest can stay local, no?
If everything needs to be global for some reason, could we at least
make it a struct, to group the fields, not just separate random
variables? And maybe at the top, not half-way throught the file?
4) Isn't the name BgBufferSyncAdjust misleading? It's not adjusting
anything, it's just invalidating the info about past runs.
5) I don't quite understand why BufferSync needs to do the dance with
delay_shmem_resize. I mean, we certainly should not run BufferSync
from the code that resizes buffers, right? Certainly not after the
eviction, from the part that actually rebuilds shmem structs etc.
So perhaps something could trigger resize while we're running the
BufferSync()? Isn't that a bit strange? If this flag is needed, it
seems more like a band-aid for some issue in the architecture.
6) Also, why should it be fine to get into situation that some of the
buffers might not be valid, during shrinking? I mean, why should
this check (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers).
It seems better to ensure we never get into "sync" in a way that
might lead some of the buffers invalid. Seems way too lowlevel to
care about whether resize is happening.
7) I don't understand the new condition for "Execute the LRU scan".
Won't this stop LRU scan even in cases when we want it to happen?
Don't we want to scan the buffers in the remaining part (after
shrinking), for example? Also, we already checked this shmem flag at
the beginning of the function - sure, it could change (if some other
process modifies it), but does that make sense? Wouldn't it cause
problems if it can change at an arbitrary point while running the
BufferSync? IMHO just another sign it may not make sense to allow
this, i.e. buffer sync should not run during the "actual" resize.
v5-0010-Additional-validation-for-buffer-in-the-ring.patch
1) So the problem is we might create a ring before shrinking shared
buffers, and then GetBufferFromRing will see bogus buffers? OK, but
we should be more careful with these checks, otherwise we'll miss
real issues when we incorrectly get an invalid buffer. Can't the
backends do this only when they for sure know we did shrink the
shared buffers? Or maybe even handle that during the barrier?
2) IMHO a sign there's the "transitions" between different NBuffers
values may not be clear enough, and we're allowing stuff to happen
in the "blurry" area. I think that's likely to cause bugs (it did
cause issues for the online checksums patch, I think).
[1]
https://www.postgresql.org/message-id/3372a09c-d1f6-4974-ad60-eec15ee0c734%40vondra.me
[2]
https://www.postgresql.org/message-id/CA%2BhUKGL5hW3i_pk5y_gcbF_C5kP-pWFjCuM8bAyCeHo3xUaH8g%40mail.g...
[3]
https://www.postgresql.org/message-id/12add41a-7625-4639-a394-a5563e349322%40eisentraut.org
[4]
https://www.postgresql.org/message-id/CA%2BTgmoZFfn0E%2BEkUAjnv_QM_00eUJPkgCJKzm3n1G4itJKMSsA%40mail...
[5]
https://www.postgresql.org/message-id/flat/cnthxg2eekacrejyeonuhiaezc7vd7o2uowlsbenxqfkjwgvwj%40qgzu...
[6]
https://www.postgresql.org/message-id/CAEze2WiMkmXUWg10y%2B_oGhJzXirZbYHB5bw0%3DVWte%2BYHwSBa%3DA%40...
[7] https://www.postgresql.org/message-id/397218.1732844567%40sss.pgh.pa.us
[8]
https://www.postgresql.org/message-id/gzhuqq3eszx7w46j5de5jehycygipsy7zmfrtdkhfbj5utl6zh%40sxyejudix...
[9]
https://lore.kernel.org/linux-mm/pr7zggtdgjqjwyrfqzusih2suofszxvlfxdptbo2smneixkp7i@nrmtbhemy3is/
[10]
https://www.postgresql.org/message-id/3qzw5fhhb3eqwl3huqabyxechbz7frxs2vk3hx3tb3h7euyvul%40pc2rmhehu...
[11]
https://www.postgresql.org/message-id/CA%2BhUKGJ-RfwSe3%3DZS2HRV9rvgrZTJJButfE8Kh5C6Ta2Eb%2BmPQ%40ma...
[12]
https://www.postgresql.org/message-id/94B56B9C-025A-463F-BC57-DF5B15B8E808%40anarazel.de
[13]
https://www.postgresql.org/message-id/CA%2BhUKGLQhsZ1dEf5Zo6JuPbs6n-qX%3DcTGy49feKf1iFA_TBP1g%40mail...
--
Tomas Vondra
Core was generated by `postgres: tomas test [local] UPDATE '.
Program terminated with signal SIGBUS, Bus error.
warning: Section `.reg-xstate/113350' in core file too small.
#0 0x000055a44eaec9b5 in PageIsNew (page=0x7f9929eea000 "") at ../../../../src/include/storage/bufpage.h:237
237 return ((const PageHeaderData *) page)->pd_upper == 0;
(gdb) bt
#0 0x000055a44eaec9b5 in PageIsNew (page=0x7f9929eea000 "") at ../../../../src/include/storage/bufpage.h:237
#1 0x000055a44eaee8d5 in _bt_checkpage (rel=0x7fa7c55ec5f8, buf=38256) at nbtpage.c:807
#2 0x000055a44eaeeeb2 in _bt_relandgetbuf (rel=0x7fa7c55ec5f8, obuf=4280, blkno=2476, access=1) at nbtpage.c:1013
#3 0x000055a44eafad3d in _bt_search (rel=0x7fa7c55ec5f8, heaprel=0x0, key=0x7ffe711ad690, bufP=0x55a464cb16c0, access=1) at nbtsearch.c:186
#4 0x000055a44eafcae5 in _bt_first (scan=0x55a464cad1d8, dir=ForwardScanDirection) at nbtsearch.c:1518
#5 0x000055a44eaf7944 in btgettuple (scan=0x55a464cad1d8, dir=ForwardScanDirection) at nbtree.c:245
#6 0x000055a44eae285e in index_getnext_tid (scan=0x55a464cad1d8, direction=ForwardScanDirection) at indexam.c:637
#7 0x000055a44eae2a78 in index_getnext_slot (scan=0x55a464cad1d8, direction=ForwardScanDirection, slot=0x55a464cac248) at indexam.c:729
#8 0x000055a44ed860aa in IndexNext (node=0x55a464cabed8) at nodeIndexscan.c:131
#9 0x000055a44ed5a425 in ExecScanFetch (node=0x55a464cabed8, epqstate=0x0, accessMtd=0x55a44ed85eda <IndexNext>, recheckMtd=0x55a44ed8665a <IndexRecheck>)
at ../../../src/include/executor/execScan.h:126
#10 0x000055a44ed5a4b9 in ExecScanExtended (node=0x55a464cabed8, accessMtd=0x55a44ed85eda <IndexNext>, recheckMtd=0x55a44ed8665a <IndexRecheck>, epqstate=0x0, qual=0x0,
projInfo=0x55a464cac548) at ../../../src/include/executor/execScan.h:187
#11 0x000055a44ed5a5f8 in ExecScan (node=0x55a464cabed8, accessMtd=0x55a44ed85eda <IndexNext>, recheckMtd=0x55a44ed8665a <IndexRecheck>) at execScan.c:59
#12 0x000055a44ed86b78 in ExecIndexScan (pstate=0x55a464cabed8) at nodeIndexscan.c:536
#13 0x000055a44ed55b86 in ExecProcNodeFirst (node=0x55a464cabed8) at execProcnode.c:469
#14 0x000055a44ed92826 in ExecProcNode (node=0x55a464cabed8) at ../../../src/include/executor/executor.h:313
#15 0x000055a44ed995aa in ExecModifyTable (pstate=0x55a464caba38) at nodeModifyTable.c:4241
#16 0x000055a44ed55b86 in ExecProcNodeFirst (node=0x55a464caba38) at execProcnode.c:469
#17 0x000055a44ed487ef in ExecProcNode (node=0x55a464caba38) at ../../../src/include/executor/executor.h:313
#18 0x000055a44ed4b64d in ExecutePlan (queryDesc=0x55a464beb6d0, operation=CMD_UPDATE, sendTuples=false, numberTuples=0, direction=ForwardScanDirection, dest=0x55a464cae600)
at execMain.c:1679
#19 0x000055a44ed48e7f in standard_ExecutorRun (queryDesc=0x55a464beb6d0, direction=ForwardScanDirection, count=0) at execMain.c:367
#20 0x000055a44ed48ce1 in ExecutorRun (queryDesc=0x55a464beb6d0, direction=ForwardScanDirection, count=0) at execMain.c:304
#21 0x000055a44f02bd94 in ProcessQuery (plan=0x55a464ca2198, sourceText=0x55a464ba90c0 "UPDATE pgbench_accounts SET abalance = abalance + 900 WHERE aid = 902417;", params=0x0,
queryEnv=0x0, dest=0x55a464cae600, qc=0x7ffe711ae740) at pquery.c:161
#22 0x000055a44f02d763 in PortalRunMulti (portal=0x55a464c31eb0, isTopLevel=true, setHoldSnapshot=false, dest=0x55a464cae600, altdest=0x55a464cae600, qc=0x7ffe711ae740)
at pquery.c:1272
#23 0x000055a44f02ccff in PortalRun (portal=0x55a464c31eb0, count=9223372036854775807, isTopLevel=true, dest=0x55a464cae600, altdest=0x55a464cae600, qc=0x7ffe711ae740)
at pquery.c:788
#24 0x000055a44f025914 in exec_simple_query (query_string=0x55a464ba90c0 "UPDATE pgbench_accounts SET abalance = abalance + 900 WHERE aid = 902417;") at postgres.c:1274
#25 0x000055a44f02ad63 in PostgresMain (dbname=0x55a464bef6c0 "test", username=0x55a464bef6a8 "tomas") at postgres.c:4776
#26 0x000055a44f021446 in BackendMain (startup_data=0x7ffe711aea50, startup_data_len=24) at backend_startup.c:124
#27 0x000055a44ef2cd7a in postmaster_child_launch (child_type=B_BACKEND, child_slot=1, startup_data=0x7ffe711aea50, startup_data_len=24, client_sock=0x7ffe711aeaa0)
at launch_backend.c:290
#28 0x000055a44ef33473 in BackendStartup (client_sock=0x7ffe711aeaa0) at postmaster.c:3595
#29 0x000055a44ef30b16 in ServerLoop () at postmaster.c:1706
#30 0x000055a44ef3046a in PostmasterMain (argc=3, argv=0x55a464ba38e0) at postmaster.c:1401
#31 0x000055a44edd663f in main (argc=3, argv=0x55a464ba38e0) at main.c:227
Attachments:
[text/plain] resize-crash.txt (4.5K, ../../a162df8e-fc90-4645-822b-93e7c9c94608@vondra.me/2-resize-crash.txt)
download | inline:
Core was generated by `postgres: tomas test [local] UPDATE '.
Program terminated with signal SIGBUS, Bus error.
warning: Section `.reg-xstate/113350' in core file too small.
#0 0x000055a44eaec9b5 in PageIsNew (page=0x7f9929eea000 "") at ../../../../src/include/storage/bufpage.h:237
237 return ((const PageHeaderData *) page)->pd_upper == 0;
(gdb) bt
#0 0x000055a44eaec9b5 in PageIsNew (page=0x7f9929eea000 "") at ../../../../src/include/storage/bufpage.h:237
#1 0x000055a44eaee8d5 in _bt_checkpage (rel=0x7fa7c55ec5f8, buf=38256) at nbtpage.c:807
#2 0x000055a44eaeeeb2 in _bt_relandgetbuf (rel=0x7fa7c55ec5f8, obuf=4280, blkno=2476, access=1) at nbtpage.c:1013
#3 0x000055a44eafad3d in _bt_search (rel=0x7fa7c55ec5f8, heaprel=0x0, key=0x7ffe711ad690, bufP=0x55a464cb16c0, access=1) at nbtsearch.c:186
#4 0x000055a44eafcae5 in _bt_first (scan=0x55a464cad1d8, dir=ForwardScanDirection) at nbtsearch.c:1518
#5 0x000055a44eaf7944 in btgettuple (scan=0x55a464cad1d8, dir=ForwardScanDirection) at nbtree.c:245
#6 0x000055a44eae285e in index_getnext_tid (scan=0x55a464cad1d8, direction=ForwardScanDirection) at indexam.c:637
#7 0x000055a44eae2a78 in index_getnext_slot (scan=0x55a464cad1d8, direction=ForwardScanDirection, slot=0x55a464cac248) at indexam.c:729
#8 0x000055a44ed860aa in IndexNext (node=0x55a464cabed8) at nodeIndexscan.c:131
#9 0x000055a44ed5a425 in ExecScanFetch (node=0x55a464cabed8, epqstate=0x0, accessMtd=0x55a44ed85eda <IndexNext>, recheckMtd=0x55a44ed8665a <IndexRecheck>)
at ../../../src/include/executor/execScan.h:126
#10 0x000055a44ed5a4b9 in ExecScanExtended (node=0x55a464cabed8, accessMtd=0x55a44ed85eda <IndexNext>, recheckMtd=0x55a44ed8665a <IndexRecheck>, epqstate=0x0, qual=0x0,
projInfo=0x55a464cac548) at ../../../src/include/executor/execScan.h:187
#11 0x000055a44ed5a5f8 in ExecScan (node=0x55a464cabed8, accessMtd=0x55a44ed85eda <IndexNext>, recheckMtd=0x55a44ed8665a <IndexRecheck>) at execScan.c:59
#12 0x000055a44ed86b78 in ExecIndexScan (pstate=0x55a464cabed8) at nodeIndexscan.c:536
#13 0x000055a44ed55b86 in ExecProcNodeFirst (node=0x55a464cabed8) at execProcnode.c:469
#14 0x000055a44ed92826 in ExecProcNode (node=0x55a464cabed8) at ../../../src/include/executor/executor.h:313
#15 0x000055a44ed995aa in ExecModifyTable (pstate=0x55a464caba38) at nodeModifyTable.c:4241
#16 0x000055a44ed55b86 in ExecProcNodeFirst (node=0x55a464caba38) at execProcnode.c:469
#17 0x000055a44ed487ef in ExecProcNode (node=0x55a464caba38) at ../../../src/include/executor/executor.h:313
#18 0x000055a44ed4b64d in ExecutePlan (queryDesc=0x55a464beb6d0, operation=CMD_UPDATE, sendTuples=false, numberTuples=0, direction=ForwardScanDirection, dest=0x55a464cae600)
at execMain.c:1679
#19 0x000055a44ed48e7f in standard_ExecutorRun (queryDesc=0x55a464beb6d0, direction=ForwardScanDirection, count=0) at execMain.c:367
#20 0x000055a44ed48ce1 in ExecutorRun (queryDesc=0x55a464beb6d0, direction=ForwardScanDirection, count=0) at execMain.c:304
#21 0x000055a44f02bd94 in ProcessQuery (plan=0x55a464ca2198, sourceText=0x55a464ba90c0 "UPDATE pgbench_accounts SET abalance = abalance + 900 WHERE aid = 902417;", params=0x0,
queryEnv=0x0, dest=0x55a464cae600, qc=0x7ffe711ae740) at pquery.c:161
#22 0x000055a44f02d763 in PortalRunMulti (portal=0x55a464c31eb0, isTopLevel=true, setHoldSnapshot=false, dest=0x55a464cae600, altdest=0x55a464cae600, qc=0x7ffe711ae740)
at pquery.c:1272
#23 0x000055a44f02ccff in PortalRun (portal=0x55a464c31eb0, count=9223372036854775807, isTopLevel=true, dest=0x55a464cae600, altdest=0x55a464cae600, qc=0x7ffe711ae740)
at pquery.c:788
#24 0x000055a44f025914 in exec_simple_query (query_string=0x55a464ba90c0 "UPDATE pgbench_accounts SET abalance = abalance + 900 WHERE aid = 902417;") at postgres.c:1274
#25 0x000055a44f02ad63 in PostgresMain (dbname=0x55a464bef6c0 "test", username=0x55a464bef6a8 "tomas") at postgres.c:4776
#26 0x000055a44f021446 in BackendMain (startup_data=0x7ffe711aea50, startup_data_len=24) at backend_startup.c:124
#27 0x000055a44ef2cd7a in postmaster_child_launch (child_type=B_BACKEND, child_slot=1, startup_data=0x7ffe711aea50, startup_data_len=24, client_sock=0x7ffe711aeaa0)
at launch_backend.c:290
#28 0x000055a44ef33473 in BackendStartup (client_sock=0x7ffe711aeaa0) at postmaster.c:3595
#29 0x000055a44ef30b16 in ServerLoop () at postmaster.c:1706
#30 0x000055a44ef3046a in PostmasterMain (argc=3, argv=0x55a464ba38e0) at postmaster.c:1401
#31 0x000055a44edd663f in main (argc=3, argv=0x55a464ba38e0) at main.c:227
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
@ 2025-07-04 14:41 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-04 15:23 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 22:55 ` Re: Changing shared_buffers without restart Jim Nasby <jnasby@upgrade.com>
1 sibling, 3 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-04 14:41 UTC (permalink / raw)
To: Tomas Vondra <tomas@vondra.me>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Fri, Jul 04, 2025 at 02:06:16AM +0200, Tomas Vondra wrote:
> I took a look at this patch, because it's somewhat related to the NUMA
> patch series I posted a couple days ago, and I've been wondering if
> it makes some of the NUMA stuff harder or simpler.
Thanks a lot for the review! It's a plenty of feedback, and I'll
probably take time to answer all of it, but I still want to address
couple of most important topics quickly.
> But I'm getting a bit lost in how exactly this interacts with things
> like overcommit, system memory accounting / OOM killer and this sort of
> stuff. I went through the thread and it seems to me the reserve+map
> approach works OK in this regard (and the messages on linux-mm seem to
> confirm this). But this information is scattered over many messages and
> it's hard to say for sure, because some of this might be relevant for
> an earlier approach, or a subtly different variant of it.
>
> A similar question is portability. The comments and commit messages
> seem to suggest most of this is linux-specific, and other platforms just
> don't have these capabilities. But there's a bunch of messages (mostly
> by Thomas Munro) that hint FreeBSD might be capable of this too, even if
> to some limited extent. And possibly even Windows/EXEC_BACKEND, although
> that seems much trickier.
>
> [...]
>
> So I think it'd be very helpful to write a README, explaining the
> currnent design/approach, and summarizing all these aspects in a single
> place. Including things like portability, interaction with the OS
> accounting, OOM killer, this kind of stuff. Some of this stuff may be
> already mentioned in code comments, but you it's hard to find those.
>
> Especially worth documenting are the states the processes need to go
> through (using the barriers), and the transacitons between them (i.e.
> what is allowed in each phase, what blocks can be visible, etc.).
Agree, I'll add some comprehensive readme in the next version. Note,
that on the topic of portability the latest version implements a new
approach suggested by Thomas Munro, which reduces problematic parts to
memfd_create only, which is mentioned as Linux specific in the
documentation, but AFAICT has FreeBSD counterparts.
> 1) no user docs
>
> There are no user .sgml docs, and maybe it's time to write some,
> explaining how to use this thing - how to configure it, how to trigger
> the resizing, etc. It took me a while to realize I need to do ALTER
> SYSTEM + pg_reload_conf() to kick this off.
>
> It should also document the user-visible limitations, e.g. what activity
> is blocked during the resizing, etc.
While the user interface is still under discussion, I agree, it makes
sense to capture this information in sgml docs.
> 2) pending GUC changes
>
> [...]
>
> It also seems a bit strange that the "switch" gets to be be driven by a
> randomly selected backend (unless I'm misunderstanding this bit). It
> seems to be true for the buffer eviction during shrinking, at least.
The resize itself is coordinated by the postmaster alone, not by a
randomly selected backend. But looks like buffer eviction indeed can
happen anywhere, which is what we were discussing in the previous
messages.
> Perhaps this should be a separate utility command, or maybe even just
> a new ALTER SYSTEM variant? Or even just a function, similar to what
> the "online checksums" patch did, possibly combined with a bgworder
> (but probably not needed, there are no db-specific tasks to do).
This is one topic we still actively discuss, but haven't had much
feedback otherwise. The pros and cons seem to be clear:
* Utilizing the existing GUC mechanism would allow treating
shared_buffers as any other configuration, meaning that potential
users of this feature don't have to do anything new to use it -- they
still can use whatever method they prefer to apply new configuration
(pg_reload_conf, pg_ctr reload, maybe even sending SIGHUP directly).
I'm also wondering if it's only shared_buffers, or some other options
could use similar approach.
* Having a separate utility command is a mighty simplification, which
helps avoiding problems you've described above.
So far we've got two against one in favour of simple utility command, so
we can as well go with that.
> 3) max_available_memory
>
> Speaking of GUCs, I dislike how max_available_memory works. It seems a
> bit backwards to me. I mean, we're specifying shared_buffers (and some
> other parameters), and the system calculates the amount of shared memory
> needed. But the limit determines the total limit?
The reason it's so backwards is that it's coming from the need to
specify how much memory we would like to reserve, and what would be the
upper boundary for increasing shared_buffers. My intention is eventually
to get rid of this GUC and figure its value at runtime as a function of
the total available memory.
> I think the GUC should specify the maximum shared_buffers we want to
> allow, and then we'd work out the total to pre-allocate? Considering
> we're only allowing to resize shared_buffers, that should be pretty
> trivial. Yes, it might happen that the "total limit" happens to exceed
> the available memory or something, but we already have the problem
> with shared_buffers. Seems fine if we explain this in the docs, and
> perhaps print the calculated memory limit on start.
Somehow I'm not following what you suggest here. You mean having the
maximum shared_buffers specified, but not as a separate GUC?
> 4) SHMEM_RESIZE_RATIO
>
> The SHMEM_RESIZE_RATIO thing seems a bit strange too. There's no way
> these ratios can make sense. For example, BLCKSZ is 8192 but the buffer
> descriptor is 64B. That's 128x difference, but the ratios says 0.6 and
> 0.1, so 6x. Sure, we'll actually allocate only the memory we need, and
> the rest is only "reserved".
SHMEM_RESIZE_RATIO is a temporary hack, waiting for more decent
solution, nothing more. I probably have to mention that in the
commentaries.
> Moreover, all of the above is for mappings sized based on NBuffers. But
> if we allocate 10% for MAIN_SHMEM_SEGMENT, won't that be a problem the
> moment someone increases of max_connection, max_locks_per_transaction
> and possibly some other stuff?
Can you elaborate, what do you mean by that? Increasing max_connection,
etc. leading to increased memory consumption in the MAIN_SHMEM_SEGMENT,
but the ratio is for memory reservation only.
> 5) no tests
>
> I mentioned no "user docs", but the patch has 0 tests too. Which seems
> a bit strange for a patch of this age.
>
> A really serious part of the patch series seems to be the coordination
> of processes when going through the phases, enforced by the barriers.
> This seems like a perfect match for testing using injection points, and
> I know we did something like this in the online checksums patch, which
> needs to coordinate processes in a similar way.
Exactly what we're talking about recently, figuring out how to use
injections points for testing. Keep in mind, that the scope of this work
turned out to be huge, and with just two people on board we're
addressing one thing at the time.
> But even just a simple TAP test that does a bunch of (random?) resizes
> while running a pgbench seem better than no tests. (That's what I did
> manually, and it crashed right away.)
This is the type of testing I was doing before posting the series. I
assume you've crashed it on buffers shrinking, singe you've got SIGBUS
which would indicate that the memory is not available anymore. Before we
go into debugging, just to be on the safe side I would like to make sure
you were testing the latest patch version (there are some signs that
it's not the case, about that later)?
> 10) what to do about stuck resize?
>
> AFAICS the resize can get stuck for various reasons, e.g. because it
> can't evict pinned buffers, possibly indefinitely. Not great, it's not
> clear to me if there's a way out (canceling the resize) after a timeout,
> or something like that? Not great to start an "online resize" only to
> get stuck with all activity blocked for indefinite amount of time, and
> get to restart anyway.
>
> Seems related to Thomas' message [2], but AFAICS the patch does not do
> anything about this yet, right? What's the plan here?
It's another open discussion right now, with an idea to eventually allow
canceling after a timeout. I think canceling when stuck on buffer
eviction should be pretty straightforward (the evition must take place
before actual shared memory resize, so we know nothing has changed yet),
but in some other failure scenarios it would be harder (e.g. if one
backend is stuck resizing, while other have succeeded -- this would
require another round of synchronization and some way to figure out what
is the current status).
> 11) preparatory actions?
>
> Even if it doesn't get stuck, some of the actions can take a while, like
> evicting dirty buffers before shrinking, etc. This is similar to what
> happens on restart, when the shutdown checkpoint can take a while, while
> the system is (partly) unavailable.
>
> The common mitigation is to do an explicit checkpoint right before the
> restart, to make the shutdown checkpoint cheap. Could we do something
> similar for the shrinking, e.g. flush buffers from the part to be
> removed before actually starting the resize?
Yeah, that's a good idea, we will try to explore it.
> 12) does this affect e.g. fork() costs?
>
> I wonder if this affects the cost of fork() in some undesirable way?
> Could it make fork() more measurably more expensive?
The number of new mappings is quite limited, so I would not expect that.
But I can measure the impact.
> 14) interesting messages from the thread
>
> While reading through the thread, I noticed a couple messages that I
> think are still relevant:
Right, I'm aware there is a lot of not yet addressed feedback, even more
than you've mentioned below. None of this feedback was ignored, we're
just solving large problems step by step. So far the focus was on how to
do memory reservation and to coordinate resize, and everybody is more
than welcome to join. But thanks for collecting the list, I probably
need to start tracking what was addressed and what was not.
> - Robert asked [5] if Linux might abruptly break this, but I find that
> unlikely. We'd point out we rely on this, and they'd likely rethink.
> This would be made safer if this was specified by POSIX - taking that
> away once implemented seems way harder than for custom extensions.
> It's likely they'd not take away the feature without an alternative
> way to achieve the same effect, I think (yes, harder to maintain).
> Tom suggests [7] this is not in POSIX.
This conversation was related to the original implementation, which was
based on mremap and slicing of mappings. As I've mentioned, the new
approach doesn't have most of those controversial points, it uses
memfd_create and regular compatible mmap -- I don't see any of those
changing their behavior any time soon.
> - Andres had an interesting comment about how overcommit interacts with
> MAP_NORESERVE. AFAIK it means we need the flag to not break overcommit
> accounting. There's also some comments about from linux-mm people [9].
The new implementation uses MAP_NORESERVE for the mapping.
> - There seem to be some issues with releasing memory backing a mapping
> with hugetlb [10]. With the fd (and truncating the file), this seems
> to release the memory, but it's linux-specific? But most of this stuff
> is specific to linux, it seems. So is this a problem? With this it
> should be working even for hugetlb ...
Again, the new implementation got rid of problematic bits here, and I
haven't found any weak points related to hugetlb in testing so far.
> - It seems FreeBSD has MFD_HUGETLB [11], so maybe we could use this and
> make the hugetlb stuff work just like on Linux? Unclear. Also, I
> thought the mfd stuff is linux-specific ... or am I confused?
Yep, probably.
> - Thomas asked [13] why we need to stop all the backends, instead of
> just waiting for them to acknowledge the new (smaller) NBuffers value
> and then let them continue. I also don't quite see why this should
> not work, and it'd limit the disruption when we have to wait for
> eviction of buffers pinned by paused cursors, etc.
I think I've replied to that one, the idea so far was to eliminate any
chance of accessing to-be-truncated buffers and make it easier to reason
about correctness of the implementation this way. I don't see any other
way how to prevent backends from accessing buffers that may disappear
without adding overhead on the read path, but if you folks have some
ideas -- please share!
> v5-0001-Process-config-reload-in-AIO-workers.patch
>
> 1) Hmmm, so which other workers may need such explicit handling? Do all
> other processes participate in procsignal stuff, or does anything
> need an explicit handling?
So far I've noticed the issue only with io_workers and the checkpointer.
> v5-0003-Introduce-pss_barrierReceivedGeneration.patch
>
> 1) Do we actually need this? Isn't it enough to just have two barriers?
> Or a barrier + condition variable, or something like that.
The issue with two barriers is that they do not prevent disjoint groups,
i.e. one backend joins the barrier, finishes the work and detaches from
the barrier, then another backends joins. I'm not familiar with how this
was solved for online checkums patch though, will take a look. Having a
barrier and a condition variable would be possible, but it's hard to
figure out for how many backends to wait. All in all, a small extention
to the ProcSignalBarrier feels to me much more elegant.
> 2) The comment talks about "coordinated way" when processing messages,
> but it's not very clear to me. It should explain what is needed and
> not possible with the current barrier code.
Yeah, I need to work on the commentaries across the patch. Here in
particular it means any coordinated way, whatever that could be. I can
add an example to clarify that part.
> v5-0004-Allow-to-use-multiple-shared-memory-mappings.patch
Most of the commentaries here and in the following patches are obviously
reasonable and I'll incorporate them into the next version.
> 5) I'm a bit confused about the segment/mapping difference. The patch
> seems to randomly mix those, or maybe I'm just confused. I mean,
> we are creating just shmem segment, and the pieces are mappings,
> right? So why do we index them by "shmem_segment"?
Indeed, the patch uses "segment" and "mapping" interchangeably, I need
to tighten it up. The relation is still one to one, thus are multiple
segments as well as mappings.
> 7) We should remember which segments got to use huge pages and which
> did not. And we should make it optional for each segment. Although,
> maybe I'm just confused about the "segment" definition - if we only
> have one, that's where huge pages are applied.
>
> If we could have multiple segments for different segments (whatever
> that means), not sure what we'll report for cases when some segments
> get to use huge pages and others don't.
Exactly to avoid solving this, I've consciously decided to postpone
implementing possibility to mix huge and regular pages so far. Any
opinions, should a single reported value be removed and this information
is instead represented as part of an informational view about shared
memory (the one you were suggesting in this thread)?
> 11) Actually, what's the difference between the contents of Mappings
> and Segments? Isn't that the same thing, indexed in the same way?
> Or could it be unified? Or are they conceptually different thing?
Unless I'm mixing something badly, the content is the same. The relation
is a segment as a structure "contains" a mapping.
> v5-0005-Address-space-reservation-for-shared-memory.patch
>
> 1) Shouldn't reserved_offset and huge_pages_on really be in the segment
> info? Or maybe even in mapping info? (again, maybe I'm confused
> about what these structs store)
I don't think there is reserved_offset variable in the latest version
anymore, can you please confirm you use it instead of ther one I've
posted in April?
> 3) So ReserveAnonymousMemory is what makes decisions about huge pages,
> for the whole reserved space / all segments in it. That's a bit
> unfortunate with respect to the desirability of some segments
> benefiting from huge pages and others not. Maybe we should have two
> "reserved" areas, one with huge pages, one without?
Again, there is no ReserveAnonymousMemory anymore, the new approach is
to reserve the memory via separate mappings.
> I guess we don't want too many segments, because that might make
> fork() more expensive, etc. Just guessing, though. Also, how would
> this work with threading?
I assume multithreading will render it unnecessary to use shared memory
favoring some other types of memory usage, but the mechanism around it
could still be the same.
> 5) The general approach seems sound to me, but I'm not expert on this.
> I wonder how portable this behavior is. I mean, will it work on other
> Unix systems / Windows? Is it POSIX or Linux extension?
Don't know yet, it's a topic for investigation.
> v5-0006-Introduce-multiple-shmem-segments-for-shared-buff.patch
>
> 2) In fact, what happens if the user tries to resize to a value that is
> too large for one of the segments? How would the system know before
> starting the resize (and failing)?
This type of situation is handled (doing hard stop) in the latest
version, because all the necessary information is present in the mapping
structure.
> v5-0007-Allow-to-resize-shared-memory-without-restart.patch
>
> 1) Why would AdjustShmemSize be needed? Isn't that a sign of a bug
> somewhere in the resizing?
When coordination with barriers kicks in, there is a cut off line after
which any newly spawned backend will not be able to take part in it
(e.g. it was too slow to init ProcSignal infrastructure).
AdjustShmemSize is used to handle this cases.
> 2) Isn't the pg_memory_barrier() in CoordinateShmemResize a bit weird?
> Why is it needed, exactly? If it's to flush stuff for processes
> consuming EmitProcSignalBarrier, it's that too late? What if a
> process consumes the barrier between the emit and memory barrier?
I think it's not needed, a leftover after code modifications.
> v5-0008-Support-shrinking-shared-buffers.patch
> v5-0009-Reinitialize-StrategyControl-after-resizing-buffe.patch
> v5-0010-Additional-validation-for-buffer-in-the-ring.patch
This reminds me I still need to review those, so Ashutosh probably can
answer those questions better than I.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-04 15:23 ` Tomas Vondra <tomas@vondra.me>
2025-07-05 10:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2 siblings, 1 reply; 167+ messages in thread
From: Tomas Vondra @ 2025-07-04 15:23 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
On 7/4/25 16:41, Dmitry Dolgov wrote:
>> On Fri, Jul 04, 2025 at 02:06:16AM +0200, Tomas Vondra wrote:
>> I took a look at this patch, because it's somewhat related to the NUMA
>> patch series I posted a couple days ago, and I've been wondering if
>> it makes some of the NUMA stuff harder or simpler.
>
> Thanks a lot for the review! It's a plenty of feedback, and I'll
> probably take time to answer all of it, but I still want to address
> couple of most important topics quickly.
>
>> But I'm getting a bit lost in how exactly this interacts with things
>> like overcommit, system memory accounting / OOM killer and this sort of
>> stuff. I went through the thread and it seems to me the reserve+map
>> approach works OK in this regard (and the messages on linux-mm seem to
>> confirm this). But this information is scattered over many messages and
>> it's hard to say for sure, because some of this might be relevant for
>> an earlier approach, or a subtly different variant of it.
>>
>> A similar question is portability. The comments and commit messages
>> seem to suggest most of this is linux-specific, and other platforms just
>> don't have these capabilities. But there's a bunch of messages (mostly
>> by Thomas Munro) that hint FreeBSD might be capable of this too, even if
>> to some limited extent. And possibly even Windows/EXEC_BACKEND, although
>> that seems much trickier.
>>
>> [...]
>>
>> So I think it'd be very helpful to write a README, explaining the
>> currnent design/approach, and summarizing all these aspects in a single
>> place. Including things like portability, interaction with the OS
>> accounting, OOM killer, this kind of stuff. Some of this stuff may be
>> already mentioned in code comments, but you it's hard to find those.
>>
>> Especially worth documenting are the states the processes need to go
>> through (using the barriers), and the transacitons between them (i.e.
>> what is allowed in each phase, what blocks can be visible, etc.).
>
> Agree, I'll add some comprehensive readme in the next version. Note,
> that on the topic of portability the latest version implements a new
> approach suggested by Thomas Munro, which reduces problematic parts to
> memfd_create only, which is mentioned as Linux specific in the
> documentation, but AFAICT has FreeBSD counterparts.
>
OK. It's not entirely clear to me if this README should be temporary, or
if it should eventually get committed. I'd probably vote to have a
proper README explaining the basic design / resizing processes etc. It
probably should not discuss portability in too much detail, that can get
stale pretty quick.
>> 1) no user docs
>>
>> There are no user .sgml docs, and maybe it's time to write some,
>> explaining how to use this thing - how to configure it, how to trigger
>> the resizing, etc. It took me a while to realize I need to do ALTER
>> SYSTEM + pg_reload_conf() to kick this off.
>>
>> It should also document the user-visible limitations, e.g. what activity
>> is blocked during the resizing, etc.
>
> While the user interface is still under discussion, I agree, it makes
> sense to capture this information in sgml docs.
>
Yeah. Spelling out the "official" way to use something is helpful.
>> 2) pending GUC changes
>>
>> [...]
>>
>> It also seems a bit strange that the "switch" gets to be be driven by a
>> randomly selected backend (unless I'm misunderstanding this bit). It
>> seems to be true for the buffer eviction during shrinking, at least.
>
> The resize itself is coordinated by the postmaster alone, not by a
> randomly selected backend. But looks like buffer eviction indeed can
> happen anywhere, which is what we were discussing in the previous
> messages.
>
>> Perhaps this should be a separate utility command, or maybe even just
>> a new ALTER SYSTEM variant? Or even just a function, similar to what
>> the "online checksums" patch did, possibly combined with a bgworder
>> (but probably not needed, there are no db-specific tasks to do).
>
> This is one topic we still actively discuss, but haven't had much
> feedback otherwise. The pros and cons seem to be clear:
>
> * Utilizing the existing GUC mechanism would allow treating
> shared_buffers as any other configuration, meaning that potential
> users of this feature don't have to do anything new to use it -- they
> still can use whatever method they prefer to apply new configuration
> (pg_reload_conf, pg_ctr reload, maybe even sending SIGHUP directly).
>
> I'm also wondering if it's only shared_buffers, or some other options
> could use similar approach.
>
I don't know. What are the "potential users" of this feature? I don't
recall any, but there may be some. How do we know the new pending flag
will work for them too?
> * Having a separate utility command is a mighty simplification, which
> helps avoiding problems you've described above.
>
> So far we've got two against one in favour of simple utility command, so
> we can as well go with that.
>
Not sure voting is a good way to make design decisions ...
>> 3) max_available_memory
>>
>> Speaking of GUCs, I dislike how max_available_memory works. It seems a
>> bit backwards to me. I mean, we're specifying shared_buffers (and some
>> other parameters), and the system calculates the amount of shared memory
>> needed. But the limit determines the total limit?
>
> The reason it's so backwards is that it's coming from the need to
> specify how much memory we would like to reserve, and what would be the
> upper boundary for increasing shared_buffers. My intention is eventually
> to get rid of this GUC and figure its value at runtime as a function of
> the total available memory.
>
I understand why it's like this. It's simple, and people do want to
limit the memory the instance will allocate. That's understandable. The
trouble is it makes it very unclear what's the implied limit on shared
buffers size. Maybe if there was a sensible way to expose that, we could
keep the max_available_memory.
But I don't think you can get rid of the GUC, at least not entirely. You
need to leave some memory aside for queries, people may start multiple
instances at once, ...
>> I think the GUC should specify the maximum shared_buffers we want to
>> allow, and then we'd work out the total to pre-allocate? Considering
>> we're only allowing to resize shared_buffers, that should be pretty
>> trivial. Yes, it might happen that the "total limit" happens to exceed
>> the available memory or something, but we already have the problem
>> with shared_buffers. Seems fine if we explain this in the docs, and
>> perhaps print the calculated memory limit on start.
>
> Somehow I'm not following what you suggest here. You mean having the
> maximum shared_buffers specified, but not as a separate GUC?
>
My suggestion was to have a guc max_shared_buffers. Based on that you
can easily calculate the size of all other segments dependent on
NBuffers, and reserve memory for that.
>> 4) SHMEM_RESIZE_RATIO
>>
>> The SHMEM_RESIZE_RATIO thing seems a bit strange too. There's no way
>> these ratios can make sense. For example, BLCKSZ is 8192 but the buffer
>> descriptor is 64B. That's 128x difference, but the ratios says 0.6 and
>> 0.1, so 6x. Sure, we'll actually allocate only the memory we need, and
>> the rest is only "reserved".
>
> SHMEM_RESIZE_RATIO is a temporary hack, waiting for more decent
> solution, nothing more. I probably have to mention that in the
> commentaries.
>
OK
>> Moreover, all of the above is for mappings sized based on NBuffers. But
>> if we allocate 10% for MAIN_SHMEM_SEGMENT, won't that be a problem the
>> moment someone increases of max_connection, max_locks_per_transaction
>> and possibly some other stuff?
>
> Can you elaborate, what do you mean by that? Increasing max_connection,
> etc. leading to increased memory consumption in the MAIN_SHMEM_SEGMENT,
> but the ratio is for memory reservation only.
>
Stuff like PGPROC, fast-path locks etc. are allocated as part of
MAIN_SHMEM_SEGMENT, right? Yet the ratio assigns 10% of the maximum
space for that. If I significantly increase GUCs like max_connections or
max_locks_per_transaction, how do you know it didn't exceed the 10%?
>> 5) no tests
>>
>> I mentioned no "user docs", but the patch has 0 tests too. Which seems
>> a bit strange for a patch of this age.
>>
>> A really serious part of the patch series seems to be the coordination
>> of processes when going through the phases, enforced by the barriers.
>> This seems like a perfect match for testing using injection points, and
>> I know we did something like this in the online checksums patch, which
>> needs to coordinate processes in a similar way.
>
> Exactly what we're talking about recently, figuring out how to use
> injections points for testing. Keep in mind, that the scope of this work
> turned out to be huge, and with just two people on board we're
> addressing one thing at the time.
>
Sure.
>> But even just a simple TAP test that does a bunch of (random?) resizes
>> while running a pgbench seem better than no tests. (That's what I did
>> manually, and it crashed right away.)
>
> This is the type of testing I was doing before posting the series. I
> assume you've crashed it on buffers shrinking, singe you've got SIGBUS
> which would indicate that the memory is not available anymore. Before we
> go into debugging, just to be on the safe side I would like to make sure
> you were testing the latest patch version (there are some signs that
> it's not the case, about that later)?
>
Maybe, I don't remember. But I also see crashes while expanding the
buffers, with assert failure here:
#4 0x0000556f159c43d1 in ExceptionalCondition
(conditionName=0x556f15c00e00 "node->prev != INVALID_PROC_NUMBER ||
list->head == procno", fileName=0x556f15c00ce0
"../../../../src/include/storage/proclist.h", lineNumber=163) at assert.c:66
#5 0x0000556f157a9831 in proclist_contains_offset (list=0x7f296333ce24,
procno=140, node_offset=100) at
../../../../src/include/storage/proclist.h:163
#6 0x0000556f157a9add in ConditionVariableTimedSleep
(cv=0x7f296333ce20, timeout=-1, wait_event_info=134217782) at
condition_variable.c:184
#7 0x0000556f157a99c9 in ConditionVariableSleep (cv=0x7f296333ce20,
wait_event_info=134217782) at condition_variable.c:98
#8 0x0000556f157902df in BarrierArriveAndWait (barrier=0x7f296333ce08,
wait_event_info=134217782) at barrier.c:191
#9 0x0000556f156d1226 in ProcessBarrierShmemResize
(barrier=0x7f296333ce08) at pg_shmem.c:1201
>> 10) what to do about stuck resize?
>>
>> AFAICS the resize can get stuck for various reasons, e.g. because it
>> can't evict pinned buffers, possibly indefinitely. Not great, it's not
>> clear to me if there's a way out (canceling the resize) after a timeout,
>> or something like that? Not great to start an "online resize" only to
>> get stuck with all activity blocked for indefinite amount of time, and
>> get to restart anyway.
>>
>> Seems related to Thomas' message [2], but AFAICS the patch does not do
>> anything about this yet, right? What's the plan here?
>
> It's another open discussion right now, with an idea to eventually allow
> canceling after a timeout. I think canceling when stuck on buffer
> eviction should be pretty straightforward (the evition must take place
> before actual shared memory resize, so we know nothing has changed yet),
> but in some other failure scenarios it would be harder (e.g. if one
> backend is stuck resizing, while other have succeeded -- this would
> require another round of synchronization and some way to figure out what
> is the current status).
>
I think it'll be crucial to structure it so that it can't get stuck
while resizing.
>> 11) preparatory actions?
>>
>> Even if it doesn't get stuck, some of the actions can take a while, like
>> evicting dirty buffers before shrinking, etc. This is similar to what
>> happens on restart, when the shutdown checkpoint can take a while, while
>> the system is (partly) unavailable.
>>
>> The common mitigation is to do an explicit checkpoint right before the
>> restart, to make the shutdown checkpoint cheap. Could we do something
>> similar for the shrinking, e.g. flush buffers from the part to be
>> removed before actually starting the resize?
>
> Yeah, that's a good idea, we will try to explore it.
>
>> 12) does this affect e.g. fork() costs?
>>
>> I wonder if this affects the cost of fork() in some undesirable way?
>> Could it make fork() more measurably more expensive?
>
> The number of new mappings is quite limited, so I would not expect that.
> But I can measure the impact.
>
>> 14) interesting messages from the thread
>>
>> While reading through the thread, I noticed a couple messages that I
>> think are still relevant:
>
> Right, I'm aware there is a lot of not yet addressed feedback, even more
> than you've mentioned below. None of this feedback was ignored, we're
> just solving large problems step by step. So far the focus was on how to
> do memory reservation and to coordinate resize, and everybody is more
> than welcome to join. But thanks for collecting the list, I probably
> need to start tracking what was addressed and what was not.
>
>> - Robert asked [5] if Linux might abruptly break this, but I find that
>> unlikely. We'd point out we rely on this, and they'd likely rethink.
>> This would be made safer if this was specified by POSIX - taking that
>> away once implemented seems way harder than for custom extensions.
>> It's likely they'd not take away the feature without an alternative
>> way to achieve the same effect, I think (yes, harder to maintain).
>> Tom suggests [7] this is not in POSIX.
>
> This conversation was related to the original implementation, which was
> based on mremap and slicing of mappings. As I've mentioned, the new
> approach doesn't have most of those controversial points, it uses
> memfd_create and regular compatible mmap -- I don't see any of those
> changing their behavior any time soon.
>
>> - Andres had an interesting comment about how overcommit interacts with
>> MAP_NORESERVE. AFAIK it means we need the flag to not break overcommit
>> accounting. There's also some comments about from linux-mm people [9].
>
> The new implementation uses MAP_NORESERVE for the mapping.
>
>> - There seem to be some issues with releasing memory backing a mapping
>> with hugetlb [10]. With the fd (and truncating the file), this seems
>> to release the memory, but it's linux-specific? But most of this stuff
>> is specific to linux, it seems. So is this a problem? With this it
>> should be working even for hugetlb ...
>
> Again, the new implementation got rid of problematic bits here, and I
> haven't found any weak points related to hugetlb in testing so far.
>
>> - It seems FreeBSD has MFD_HUGETLB [11], so maybe we could use this and
>> make the hugetlb stuff work just like on Linux? Unclear. Also, I
>> thought the mfd stuff is linux-specific ... or am I confused?
>
> Yep, probably.
>
>> - Thomas asked [13] why we need to stop all the backends, instead of
>> just waiting for them to acknowledge the new (smaller) NBuffers value
>> and then let them continue. I also don't quite see why this should
>> not work, and it'd limit the disruption when we have to wait for
>> eviction of buffers pinned by paused cursors, etc.
>
> I think I've replied to that one, the idea so far was to eliminate any
> chance of accessing to-be-truncated buffers and make it easier to reason
> about correctness of the implementation this way. I don't see any other
> way how to prevent backends from accessing buffers that may disappear
> without adding overhead on the read path, but if you folks have some
> ideas -- please share!
>
>> v5-0001-Process-config-reload-in-AIO-workers.patch
>>
>> 1) Hmmm, so which other workers may need such explicit handling? Do all
>> other processes participate in procsignal stuff, or does anything
>> need an explicit handling?
>
> So far I've noticed the issue only with io_workers and the checkpointer.
>
>> v5-0003-Introduce-pss_barrierReceivedGeneration.patch
>>
>> 1) Do we actually need this? Isn't it enough to just have two barriers?
>> Or a barrier + condition variable, or something like that.
>
> The issue with two barriers is that they do not prevent disjoint groups,
> i.e. one backend joins the barrier, finishes the work and detaches from
> the barrier, then another backends joins. I'm not familiar with how this
> was solved for online checkums patch though, will take a look. Having a
> barrier and a condition variable would be possible, but it's hard to
> figure out for how many backends to wait. All in all, a small extention
> to the ProcSignalBarrier feels to me much more elegant.
>
>> 2) The comment talks about "coordinated way" when processing messages,
>> but it's not very clear to me. It should explain what is needed and
>> not possible with the current barrier code.
>
> Yeah, I need to work on the commentaries across the patch. Here in
> particular it means any coordinated way, whatever that could be. I can
> add an example to clarify that part.
>
>> v5-0004-Allow-to-use-multiple-shared-memory-mappings.patch
>
> Most of the commentaries here and in the following patches are obviously
> reasonable and I'll incorporate them into the next version.
>
>> 5) I'm a bit confused about the segment/mapping difference. The patch
>> seems to randomly mix those, or maybe I'm just confused. I mean,
>> we are creating just shmem segment, and the pieces are mappings,
>> right? So why do we index them by "shmem_segment"?
>
> Indeed, the patch uses "segment" and "mapping" interchangeably, I need
> to tighten it up. The relation is still one to one, thus are multiple
> segments as well as mappings.
>
>> 7) We should remember which segments got to use huge pages and which
>> did not. And we should make it optional for each segment. Although,
>> maybe I'm just confused about the "segment" definition - if we only
>> have one, that's where huge pages are applied.
>>
>> If we could have multiple segments for different segments (whatever
>> that means), not sure what we'll report for cases when some segments
>> get to use huge pages and others don't.
>
> Exactly to avoid solving this, I've consciously decided to postpone
> implementing possibility to mix huge and regular pages so far. Any
> opinions, should a single reported value be removed and this information
> is instead represented as part of an informational view about shared
> memory (the one you were suggesting in this thread)?
>
>> 11) Actually, what's the difference between the contents of Mappings
>> and Segments? Isn't that the same thing, indexed in the same way?
>> Or could it be unified? Or are they conceptually different thing?
>
> Unless I'm mixing something badly, the content is the same. The relation
> is a segment as a structure "contains" a mapping.
>
Then, why do we need to track it in two places? Doesn't it just increase
the likelihood that someone misses updating one of them?
>> v5-0005-Address-space-reservation-for-shared-memory.patch
>>
>> 1) Shouldn't reserved_offset and huge_pages_on really be in the segment
>> info? Or maybe even in mapping info? (again, maybe I'm confused
>> about what these structs store)
>
> I don't think there is reserved_offset variable in the latest version
> anymore, can you please confirm you use it instead of ther one I've
> posted in April?
>
>> 3) So ReserveAnonymousMemory is what makes decisions about huge pages,
>> for the whole reserved space / all segments in it. That's a bit
>> unfortunate with respect to the desirability of some segments
>> benefiting from huge pages and others not. Maybe we should have two
>> "reserved" areas, one with huge pages, one without?
>
> Again, there is no ReserveAnonymousMemory anymore, the new approach is
> to reserve the memory via separate mappings.
>
Will check. These may indeed be stale comments, from looking at the
earlier version of the patch (the last one from Ashutosh).
>> I guess we don't want too many segments, because that might make
>> fork() more expensive, etc. Just guessing, though. Also, how would
>> this work with threading?
>
> I assume multithreading will render it unnecessary to use shared memory
> favoring some other types of memory usage, but the mechanism around it
> could still be the same.
>
>> 5) The general approach seems sound to me, but I'm not expert on this.
>> I wonder how portable this behavior is. I mean, will it work on other
>> Unix systems / Windows? Is it POSIX or Linux extension?
>
> Don't know yet, it's a topic for investigation.
>
>> v5-0006-Introduce-multiple-shmem-segments-for-shared-buff.patch
>>
>> 2) In fact, what happens if the user tries to resize to a value that is
>> too large for one of the segments? How would the system know before
>> starting the resize (and failing)?
>
> This type of situation is handled (doing hard stop) in the latest
> version, because all the necessary information is present in the mapping
> structure.
>
I don't know, but crashing the instance (I assume that's what you mean
by hard stop) does not seem like something we want to do. AFAIK the GUC
hook should be able to determine if the value is too large, and reject
it at that point. Not proceed and crash everything.
>> v5-0007-Allow-to-resize-shared-memory-without-restart.patch
>>
>> 1) Why would AdjustShmemSize be needed? Isn't that a sign of a bug
>> somewhere in the resizing?
>
> When coordination with barriers kicks in, there is a cut off line after
> which any newly spawned backend will not be able to take part in it
> (e.g. it was too slow to init ProcSignal infrastructure).
> AdjustShmemSize is used to handle this cases.
>
>> 2) Isn't the pg_memory_barrier() in CoordinateShmemResize a bit weird?
>> Why is it needed, exactly? If it's to flush stuff for processes
>> consuming EmitProcSignalBarrier, it's that too late? What if a
>> process consumes the barrier between the emit and memory barrier?
>
> I think it's not needed, a leftover after code modifications.
>
>> v5-0008-Support-shrinking-shared-buffers.patch
>> v5-0009-Reinitialize-StrategyControl-after-resizing-buffe.patch
>> v5-0010-Additional-validation-for-buffer-in-the-ring.patch
>
> This reminds me I still need to review those, so Ashutosh probably can
> answer those questions better than I.
--
Tomas Vondra
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-04 15:23 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
@ 2025-07-05 10:35 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-07 11:57 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-05 10:35 UTC (permalink / raw)
To: Tomas Vondra <tomas@vondra.me>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Fri, Jul 04, 2025 at 05:23:29PM +0200, Tomas Vondra wrote:
> >> 2) pending GUC changes
> >>
> >> Perhaps this should be a separate utility command, or maybe even just
> >> a new ALTER SYSTEM variant? Or even just a function, similar to what
> >> the "online checksums" patch did, possibly combined with a bgworder
> >> (but probably not needed, there are no db-specific tasks to do).
> >
> > This is one topic we still actively discuss, but haven't had much
> > feedback otherwise. The pros and cons seem to be clear:
> >
> > * Utilizing the existing GUC mechanism would allow treating
> > shared_buffers as any other configuration, meaning that potential
> > users of this feature don't have to do anything new to use it -- they
> > still can use whatever method they prefer to apply new configuration
> > (pg_reload_conf, pg_ctr reload, maybe even sending SIGHUP directly).
> >
> > I'm also wondering if it's only shared_buffers, or some other options
> > could use similar approach.
> >
>
> I don't know. What are the "potential users" of this feature? I don't
> recall any, but there may be some. How do we know the new pending flag
> will work for them too?
It could be potentialy useful for any GUC that controls a resource
shared between backend, and requires restart. To make this GUC
changeable online, every backend has to perform some action, and they
have to coordinate to make sure things are consistent -- exactly the use
case we're trying to address, shared_buffers is just happened to be one
of such resources. While I agree that the currently implemented
interface is wrong (e.g. it doesn't prevent pending GUCs from being
stored in PG_AUTOCONF_FILENAME, this has to happen only when the new
value is actually applied), it still makes sense to me to allow more
flexible lifecycle for certain GUC.
An example I could think of is shared_preload_libraries. If we ever want
to do a hot reload of libraries, this will follow the procedure above:
every backend has to do something like dlclose / dlopen and make sure
that other backends have the same version of the library. Another maybe
less far fetched example is max_worker_processes, which AFAICT is mostly
used to control number of slots in shared memory (altough it's also
stored in the control file, which makes things more complicated).
> > * Having a separate utility command is a mighty simplification, which
> > helps avoiding problems you've described above.
> >
> > So far we've got two against one in favour of simple utility command, so
> > we can as well go with that.
> >
>
> Not sure voting is a good way to make design decisions ...
I'm somewhat torn between those two options myself. The more I think
about this topic, the more I convinced that pending GUC makes sense, but
the more work I see needed to implement that. Maybe a good middle ground
is to go with a simple utility command, as Ashutosh was suggesting, and
keep pending GUC infrastructure on top of that as an optional patch.
> >> 3) max_available_memory
> >>
> >> I think the GUC should specify the maximum shared_buffers we want to
> >> allow, and then we'd work out the total to pre-allocate? Considering
> >> we're only allowing to resize shared_buffers, that should be pretty
> >> trivial. Yes, it might happen that the "total limit" happens to exceed
> >> the available memory or something, but we already have the problem
> >> with shared_buffers. Seems fine if we explain this in the docs, and
> >> perhaps print the calculated memory limit on start.
> >
> > Somehow I'm not following what you suggest here. You mean having the
> > maximum shared_buffers specified, but not as a separate GUC?
>
> My suggestion was to have a guc max_shared_buffers. Based on that you
> can easily calculate the size of all other segments dependent on
> NBuffers, and reserve memory for that.
Got it, ok.
> >> Moreover, all of the above is for mappings sized based on NBuffers. But
> >> if we allocate 10% for MAIN_SHMEM_SEGMENT, won't that be a problem the
> >> moment someone increases of max_connection, max_locks_per_transaction
> >> and possibly some other stuff?
> >
> > Can you elaborate, what do you mean by that? Increasing max_connection,
> > etc. leading to increased memory consumption in the MAIN_SHMEM_SEGMENT,
> > but the ratio is for memory reservation only.
> >
>
> Stuff like PGPROC, fast-path locks etc. are allocated as part of
> MAIN_SHMEM_SEGMENT, right? Yet the ratio assigns 10% of the maximum
> space for that. If I significantly increase GUCs like max_connections or
> max_locks_per_transaction, how do you know it didn't exceed the 10%?
Still don't see the problem. The 10% we're talking about is the reserved
space, thus it affects only shared memory resizing operation and nothing
else. The real memory allocated is less than or equal to the reserved
size, but is allocated and managed completely in the same way as without
the patch, including size calculations. If some GUCs are increased and
drive real memory usage high, it will be handled as before. Are we on
the same page about this?
> >> 11) Actually, what's the difference between the contents of Mappings
> >> and Segments? Isn't that the same thing, indexed in the same way?
> >> Or could it be unified? Or are they conceptually different thing?
> >
> > Unless I'm mixing something badly, the content is the same. The relation
> > is a segment as a structure "contains" a mapping.
> >
> Then, why do we need to track it in two places? Doesn't it just increase
> the likelihood that someone misses updating one of them?
To clarify, under "contents" I mean the shared memory content (the
actual data) behind both "segment" and the "mapping", maybe you had
something else in mind.
On the surface of it those are two different data structures that have
mostly different, but related, fields: a shared memory segment contains
stuff needed for working with memory (header, base, end, lock), mapping
has more lower level details, e.g. reserved space, fd, IPC key. The
only common fields are size and address, maybe I can factor them out to
not repeat.
> >> 2) In fact, what happens if the user tries to resize to a value that is
> >> too large for one of the segments? How would the system know before
> >> starting the resize (and failing)?
> >
> > This type of situation is handled (doing hard stop) in the latest
> > version, because all the necessary information is present in the mapping
> > structure.
> >
> I don't know, but crashing the instance (I assume that's what you mean
> by hard stop) does not seem like something we want to do. AFAIK the GUC
> hook should be able to determine if the value is too large, and reject
> it at that point. Not proceed and crash everything.
I see, you're pointing out that it would be good to have more validation
at the GUC level, right?
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-04 15:23 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-05 10:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-07 11:57 ` Tomas Vondra <tomas@vondra.me>
2025-07-07 13:06 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Tomas Vondra @ 2025-07-07 11:57 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
On 7/5/25 12:35, Dmitry Dolgov wrote:
>> On Fri, Jul 04, 2025 at 05:23:29PM +0200, Tomas Vondra wrote:
>>>> 2) pending GUC changes
>>>>
>>>> Perhaps this should be a separate utility command, or maybe even just
>>>> a new ALTER SYSTEM variant? Or even just a function, similar to what
>>>> the "online checksums" patch did, possibly combined with a bgworder
>>>> (but probably not needed, there are no db-specific tasks to do).
>>>
>>> This is one topic we still actively discuss, but haven't had much
>>> feedback otherwise. The pros and cons seem to be clear:
>>>
>>> * Utilizing the existing GUC mechanism would allow treating
>>> shared_buffers as any other configuration, meaning that potential
>>> users of this feature don't have to do anything new to use it -- they
>>> still can use whatever method they prefer to apply new configuration
>>> (pg_reload_conf, pg_ctr reload, maybe even sending SIGHUP directly).
>>>
>>> I'm also wondering if it's only shared_buffers, or some other options
>>> could use similar approach.
>>>
>>
>> I don't know. What are the "potential users" of this feature? I don't
>> recall any, but there may be some. How do we know the new pending flag
>> will work for them too?
>
> It could be potentialy useful for any GUC that controls a resource
> shared between backend, and requires restart. To make this GUC
> changeable online, every backend has to perform some action, and they
> have to coordinate to make sure things are consistent -- exactly the use
> case we're trying to address, shared_buffers is just happened to be one
> of such resources. While I agree that the currently implemented
> interface is wrong (e.g. it doesn't prevent pending GUCs from being
> stored in PG_AUTOCONF_FILENAME, this has to happen only when the new
> value is actually applied), it still makes sense to me to allow more
> flexible lifecycle for certain GUC.
>
> An example I could think of is shared_preload_libraries. If we ever want
> to do a hot reload of libraries, this will follow the procedure above:
> every backend has to do something like dlclose / dlopen and make sure
> that other backends have the same version of the library. Another maybe
> less far fetched example is max_worker_processes, which AFAICT is mostly
> used to control number of slots in shared memory (altough it's also
> stored in the control file, which makes things more complicated).
>
Not sure. My concern is the config reload / GUC assign hook was not
designed with this use case in mind, and we'll run into issues. I also
dislike the "async" nature of this, which makes it harder to e.g. abort
the change, etc.
>>> * Having a separate utility command is a mighty simplification, which
>>> helps avoiding problems you've described above.
>>>
>>> So far we've got two against one in favour of simple utility command, so
>>> we can as well go with that.
>>>
>>
>> Not sure voting is a good way to make design decisions ...
>
> I'm somewhat torn between those two options myself. The more I think
> about this topic, the more I convinced that pending GUC makes sense, but
> the more work I see needed to implement that. Maybe a good middle ground
> is to go with a simple utility command, as Ashutosh was suggesting, and
> keep pending GUC infrastructure on top of that as an optional patch.
>
What about a simple function? Probably not as clean as a proper utility
command, and it implies a transaction - not sure if that could be a
problem for some part of this.
>>>> 3) max_available_memory
>>>>
>>>> I think the GUC should specify the maximum shared_buffers we want to
>>>> allow, and then we'd work out the total to pre-allocate? Considering
>>>> we're only allowing to resize shared_buffers, that should be pretty
>>>> trivial. Yes, it might happen that the "total limit" happens to exceed
>>>> the available memory or something, but we already have the problem
>>>> with shared_buffers. Seems fine if we explain this in the docs, and
>>>> perhaps print the calculated memory limit on start.
>>>
>>> Somehow I'm not following what you suggest here. You mean having the
>>> maximum shared_buffers specified, but not as a separate GUC?
>>
>> My suggestion was to have a guc max_shared_buffers. Based on that you
>> can easily calculate the size of all other segments dependent on
>> NBuffers, and reserve memory for that.
>
> Got it, ok.
>
>>>> Moreover, all of the above is for mappings sized based on NBuffers. But
>>>> if we allocate 10% for MAIN_SHMEM_SEGMENT, won't that be a problem the
>>>> moment someone increases of max_connection, max_locks_per_transaction
>>>> and possibly some other stuff?
>>>
>>> Can you elaborate, what do you mean by that? Increasing max_connection,
>>> etc. leading to increased memory consumption in the MAIN_SHMEM_SEGMENT,
>>> but the ratio is for memory reservation only.
>>>
>>
>> Stuff like PGPROC, fast-path locks etc. are allocated as part of
>> MAIN_SHMEM_SEGMENT, right? Yet the ratio assigns 10% of the maximum
>> space for that. If I significantly increase GUCs like max_connections or
>> max_locks_per_transaction, how do you know it didn't exceed the 10%?
>
> Still don't see the problem. The 10% we're talking about is the reserved
> space, thus it affects only shared memory resizing operation and nothing
> else. The real memory allocated is less than or equal to the reserved
> size, but is allocated and managed completely in the same way as without
> the patch, including size calculations. If some GUCs are increased and
> drive real memory usage high, it will be handled as before. Are we on
> the same page about this?
>
How do you know reserving 10% is sufficient? Imagine I set
max_available_memory = '256MB'
max_connections = 1000000
max_locks_per_transaction = 10000
How do you know it's not more than 10% of the available memory?
FWIW if I add a simple assert to CreateAnonymousSegment
Assert(mapping->shmem_reserved >= allocsize);
it crashes even with just the max_available_memory=256MB
#4 0x0000000000b74fbd in ExceptionalCondition (conditionName=0xe25920
"mapping->shmem_reserved >= allocsize", fileName=0xe251e7 "pg_shmem.c",
lineNumber=878) at assert.c:66
because we happen to execute it with this:
mapping->shmem_reserved 26845184 allocsize 125042688
I think I mentioned a similar crash earlier, not sure if that's the same
issue or a different one.
>>>> 11) Actually, what's the difference between the contents of Mappings
>>>> and Segments? Isn't that the same thing, indexed in the same way?
>>>> Or could it be unified? Or are they conceptually different thing?
>>>
>>> Unless I'm mixing something badly, the content is the same. The relation
>>> is a segment as a structure "contains" a mapping.
>>>
>> Then, why do we need to track it in two places? Doesn't it just increase
>> the likelihood that someone misses updating one of them?
>
> To clarify, under "contents" I mean the shared memory content (the
> actual data) behind both "segment" and the "mapping", maybe you had
> something else in mind.
>
> On the surface of it those are two different data structures that have
> mostly different, but related, fields: a shared memory segment contains
> stuff needed for working with memory (header, base, end, lock), mapping
> has more lower level details, e.g. reserved space, fd, IPC key. The
> only common fields are size and address, maybe I can factor them out to
> not repeat.
>
OK, I think I'm just confused by the ambiguous definitions of
segment/mapping. It'd be good to document/explain this in a comment
somewhere.
>>>> 2) In fact, what happens if the user tries to resize to a value that is
>>>> too large for one of the segments? How would the system know before
>>>> starting the resize (and failing)?
>>>
>>> This type of situation is handled (doing hard stop) in the latest
>>> version, because all the necessary information is present in the mapping
>>> structure.
>>>
>> I don't know, but crashing the instance (I assume that's what you mean
>> by hard stop) does not seem like something we want to do. AFAIK the GUC
>> hook should be able to determine if the value is too large, and reject
>> it at that point. Not proceed and crash everything.
>
> I see, you're pointing out that it would be good to have more validation
> at the GUC level, right?
Well, that'd be a starting point. We definitely should not allow setting
a value that end up crashing an instance (it does not matter if it's
because of FATAL or hitting a segfault/sigbut somewhere).
cheers
--
Tomas Vondra
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-04 15:23 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-05 10:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-07 11:57 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
@ 2025-07-07 13:06 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-07 13:42 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-07 13:06 UTC (permalink / raw)
To: Tomas Vondra <tomas@vondra.me>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 07, 2025 at 01:57:42PM +0200, Tomas Vondra wrote:
> > It could be potentialy useful for any GUC that controls a resource
> > shared between backend, and requires restart. To make this GUC
> > changeable online, every backend has to perform some action, and they
> > have to coordinate to make sure things are consistent -- exactly the use
> > case we're trying to address, shared_buffers is just happened to be one
> > of such resources. While I agree that the currently implemented
> > interface is wrong (e.g. it doesn't prevent pending GUCs from being
> > stored in PG_AUTOCONF_FILENAME, this has to happen only when the new
> > value is actually applied), it still makes sense to me to allow more
> > flexible lifecycle for certain GUC.
> >
> > An example I could think of is shared_preload_libraries. If we ever want
> > to do a hot reload of libraries, this will follow the procedure above:
> > every backend has to do something like dlclose / dlopen and make sure
> > that other backends have the same version of the library. Another maybe
> > less far fetched example is max_worker_processes, which AFAICT is mostly
> > used to control number of slots in shared memory (altough it's also
> > stored in the control file, which makes things more complicated).
> >
>
> Not sure. My concern is the config reload / GUC assign hook was not
> designed with this use case in mind, and we'll run into issues. I also
> dislike the "async" nature of this, which makes it harder to e.g. abort
> the change, etc.
Yes, GUC assing hook was not designed for that. That's why the idea is
to extend the design and see if it will be good enough.
> > I'm somewhat torn between those two options myself. The more I think
> > about this topic, the more I convinced that pending GUC makes sense, but
> > the more work I see needed to implement that. Maybe a good middle ground
> > is to go with a simple utility command, as Ashutosh was suggesting, and
> > keep pending GUC infrastructure on top of that as an optional patch.
> >
>
> What about a simple function? Probably not as clean as a proper utility
> command, and it implies a transaction - not sure if that could be a
> problem for some part of this.
I'm currently inclined towards this and a new one worker to coordinate
the process, with everything else provided as an optional follow-up
step. Will try this out unless there are any objections.
> >> Stuff like PGPROC, fast-path locks etc. are allocated as part of
> >> MAIN_SHMEM_SEGMENT, right? Yet the ratio assigns 10% of the maximum
> >> space for that. If I significantly increase GUCs like max_connections or
> >> max_locks_per_transaction, how do you know it didn't exceed the 10%?
> >
> > Still don't see the problem. The 10% we're talking about is the reserved
> > space, thus it affects only shared memory resizing operation and nothing
> > else. The real memory allocated is less than or equal to the reserved
> > size, but is allocated and managed completely in the same way as without
> > the patch, including size calculations. If some GUCs are increased and
> > drive real memory usage high, it will be handled as before. Are we on
> > the same page about this?
> >
>
> How do you know reserving 10% is sufficient? Imagine I set
I see, I was convinced you're talking about changing something at
runtime, which will hit the reservation boundary. But you mean all of
that at simply the start, and yes, of course it will fail -- see the
point about SHMEM_RATIO being just a temporary hack.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-04 15:23 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-05 10:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-07 11:57 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-07 13:06 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-07 13:42 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-07 13:58 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-07-07 13:42 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
On Mon, Jul 7, 2025 at 6:36 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Mon, Jul 07, 2025 at 01:57:42PM +0200, Tomas Vondra wrote:
> > > It could be potentialy useful for any GUC that controls a resource
> > > shared between backend, and requires restart. To make this GUC
> > > changeable online, every backend has to perform some action, and they
> > > have to coordinate to make sure things are consistent -- exactly the use
> > > case we're trying to address, shared_buffers is just happened to be one
> > > of such resources. While I agree that the currently implemented
> > > interface is wrong (e.g. it doesn't prevent pending GUCs from being
> > > stored in PG_AUTOCONF_FILENAME, this has to happen only when the new
> > > value is actually applied), it still makes sense to me to allow more
> > > flexible lifecycle for certain GUC.
> > >
> > > An example I could think of is shared_preload_libraries. If we ever want
> > > to do a hot reload of libraries, this will follow the procedure above:
> > > every backend has to do something like dlclose / dlopen and make sure
> > > that other backends have the same version of the library. Another maybe
> > > less far fetched example is max_worker_processes, which AFAICT is mostly
> > > used to control number of slots in shared memory (altough it's also
> > > stored in the control file, which makes things more complicated).
> > >
> >
> > Not sure. My concern is the config reload / GUC assign hook was not
> > designed with this use case in mind, and we'll run into issues. I also
> > dislike the "async" nature of this, which makes it harder to e.g. abort
> > the change, etc.
>
> Yes, GUC assing hook was not designed for that. That's why the idea is
> to extend the design and see if it will be good enough.
>
> > > I'm somewhat torn between those two options myself. The more I think
> > > about this topic, the more I convinced that pending GUC makes sense, but
> > > the more work I see needed to implement that. Maybe a good middle ground
> > > is to go with a simple utility command, as Ashutosh was suggesting, and
> > > keep pending GUC infrastructure on top of that as an optional patch.
> > >
> >
> > What about a simple function? Probably not as clean as a proper utility
> > command, and it implies a transaction - not sure if that could be a
> > problem for some part of this.
>
> I'm currently inclined towards this and a new one worker to coordinate
> the process, with everything else provided as an optional follow-up
> step. Will try this out unless there are any objections.
I will reply to the questions but let me summarise my offlist
discussion with Andres.
I had proposed ALTER SYSTEM ... UPDATE ... approach in pgconf.dev for
any system wide GUC change such as this. However, Andres pointed out
that any UI proposal has to honour the current ability to edit
postgresql.conf and trigger the change in a running server. ALTER
SYSTEM ... UDPATE ... does not allow that. So, I think we have to
build something similar or on top of the current ALTER SYSTEM ... SET
+ pg_reload_conf().
My current proposal is ALTER SYSTEM ... SET + pg_reload_conf() with
pending mark + pg_apply_pending_conf(<name of GUC>, <more
parameters>). The third function would take a GUC name as parameter
and complete the pending application change. If the proposed change is
not valid, it will throw an error. If there are problems completing
the change it will throw an error and keep the pending mark intact.
Further the function can take GUC specific parameters which control
the application process. E.g. for example it could tell whether to
wait for a backend to unpin a buffer or cancel that query or kill the
backend or abort the application itself. If the operation takes too
long, a user may want to cancel the function execution just like
cancelling a query. Running two concurrent instances of the function,
both applying the same GUC won't be allowed.
Does that look good?
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-04 15:23 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-05 10:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-07 11:57 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-07 13:06 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-07 13:42 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-07-07 13:58 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-07 13:58 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 07, 2025 at 07:12:50PM +0530, Ashutosh Bapat wrote:
>
> My current proposal is ALTER SYSTEM ... SET + pg_reload_conf() with
> pending mark + pg_apply_pending_conf(<name of GUC>, <more
> parameters>). The third function would take a GUC name as parameter
> and complete the pending application change. If the proposed change is
> not valid, it will throw an error. If there are problems completing
> the change it will throw an error and keep the pending mark intact.
> Further the function can take GUC specific parameters which control
> the application process. E.g. for example it could tell whether to
> wait for a backend to unpin a buffer or cancel that query or kill the
> backend or abort the application itself. If the operation takes too
> long, a user may want to cancel the function execution just like
> cancelling a query. Running two concurrent instances of the function,
> both applying the same GUC won't be allowed.
Yeah, it can look like this, but it's a large chunk of work as well as
improving the current implementation. I'm still convinced that using GUC
mechanism one or another way is the right choice here, but maybe better
as a follow-up step I was mentioning above -- simply to limit the scope
and move step by step. How does it sound?
Regarding the proposal, I'm somehow uncomfortable with the fact that
between those two function call the system will be in an awkward state
for some time, and how long would it take will not be controlled by
the resizing logic anymore. But otherwise it seems to be equivalent of
what we want to achieve in many other apspects.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-06 13:01 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-06 13:01 UTC (permalink / raw)
To: Tomas Vondra <tomas@vondra.me>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Fri, Jul 04, 2025 at 04:41:51PM +0200, Dmitry Dolgov wrote:
> > v5-0003-Introduce-pss_barrierReceivedGeneration.patch
> >
> > 1) Do we actually need this? Isn't it enough to just have two barriers?
> > Or a barrier + condition variable, or something like that.
>
> The issue with two barriers is that they do not prevent disjoint groups,
> i.e. one backend joins the barrier, finishes the work and detaches from
> the barrier, then another backends joins. I'm not familiar with how this
> was solved for online checkums patch though, will take a look. Having a
> barrier and a condition variable would be possible, but it's hard to
> figure out for how many backends to wait. All in all, a small extention
> to the ProcSignalBarrier feels to me much more elegant.
After quickly checking how online checksums patch is dealing with the
coordination, I've realized my answer here about the disjoint groups is
not quite correct. You were asking about ProcSignalBarrier, I was
answering about the barrier within the resizing logic. Here is how it
looks like to me:
* We could follow the same way as the online checksums, launch a
coordinator worker (Ashutosh was suggesting that, but no
implementation has materialized yet) and fire two ProcSignalBarriers,
one to kick off resizing and another one to finish it. Maybe it could
even be three phases, one extra to tell backends to not pull in new
buffers into the pool to help buffer eviction process.
* This way any backend between the ProcSignalBarriers will be able
proceed with whatever it's doing, and there is need to make sure it
will not access buffers that will soon disappear. A suggestion so far
was to get all backends agree to not allocate any new buffers in the
to-be-truncated range, but accessing already existing buffers that
will soon go away is a problem as well. As far as I can tell there is
no rock solid method to make sure a backend doesn't have a reference
to such a buffer somewhere (this was discussed earlier in thre
thread), meaning that either a backend has to wait or buffers have to
be checked every time on access.
* Since the latter adds a performance overhead, we went with the former
(making backends wait). And here is where all the complexity comes
from, because waiting backends cannot reply on a ProcSignalBarrier and
thus require some other approach. If I've overlooked any other
alternative to backends waiting, let me know.
> It also seems a bit strange that the "switch" gets to be be driven by
> a randomly selected backend (unless I'm misunderstanding this bit). It
> seems to be true for the buffer eviction during shrinking, at least.
But looks like the eviction could be indeed improved via a new
coordinator worker. Before resizing shared memory such a worker will
first tell all the backends to not allocate new buffers via
ProcSignalBarrier, then will do buffer eviction. Since backends don't
need to be waiting after this type of ProcSignalBarrier, it should work
and establish only one worker to do the eviction. But the second
ProcSignalBarrier for resizing would still follow the current procedure
with everybody waiting.
Does it make sense to you folks?
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-06 13:21 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-06 13:21 UTC (permalink / raw)
To: Tomas Vondra <tomas@vondra.me>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Sun, Jul 06, 2025 at 03:01:34PM +0200, Dmitry Dolgov wrote:
> * This way any backend between the ProcSignalBarriers will be able
> proceed with whatever it's doing, and there is need to make sure it
> will not access buffers that will soon disappear. A suggestion so far
> was to get all backends agree to not allocate any new buffers in the
> to-be-truncated range, but accessing already existing buffers that
> will soon go away is a problem as well. As far as I can tell there is
> no rock solid method to make sure a backend doesn't have a reference
> to such a buffer somewhere (this was discussed earlier in thre
> thread), meaning that either a backend has to wait or buffers have to
> be checked every time on access.
And sure enough, after I wrote this I've realized there should be no
such references after the buffer eviction and prohibiting new buffer
allocation. I still need to check it though, because not only buffers,
but other shared memory structures (which number depends on NBuffers)
will be truncated. But if they will also be handled by the eviction,
then maybe everything is just fine.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-13 18:37 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-13 18:37 UTC (permalink / raw)
To: Tomas Vondra <tomas@vondra.me>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Sun, Jul 06, 2025 at 03:21:08PM +0200, Dmitry Dolgov wrote:
> > On Sun, Jul 06, 2025 at 03:01:34PM +0200, Dmitry Dolgov wrote:
> > * This way any backend between the ProcSignalBarriers will be able
> > proceed with whatever it's doing, and there is need to make sure it
> > will not access buffers that will soon disappear. A suggestion so far
> > was to get all backends agree to not allocate any new buffers in the
> > to-be-truncated range, but accessing already existing buffers that
> > will soon go away is a problem as well. As far as I can tell there is
> > no rock solid method to make sure a backend doesn't have a reference
> > to such a buffer somewhere (this was discussed earlier in thre
> > thread), meaning that either a backend has to wait or buffers have to
> > be checked every time on access.
>
> And sure enough, after I wrote this I've realized there should be no
> such references after the buffer eviction and prohibiting new buffer
> allocation. I still need to check it though, because not only buffers,
> but other shared memory structures (which number depends on NBuffers)
> will be truncated. But if they will also be handled by the eviction,
> then maybe everything is just fine.
Pondering more about this topic, I've realized there was one more
problematic case mentioned by Robert early in the thread, which is
relatively easy to construct:
* When increasing shared buffers from NBuffers_small to NBuffers_large
it's possible that one backend already has applied NBuffers_large,
then allocated a buffer B from (NBuffer_small, NBuffers_large] and put
it into the buffer lookup table.
* In the meantime another backend still has NBuffers_small, but got
buffer B from the lookup table.
Currently it's being addressed via every backend waiting for each other,
but I guess it could be as well managed via handling the freelist, so
that only "available" buffers will be inserted into the lookup table.
It's probably the only such case, but I can't tell that for sure (hard
to say, maybe there are more tricky cases with the latest async io). If
you folks have some other examples that may break, let me know. The
idea behind making everyone wait was to be rock solid that no similar
but unknown scenarios could damage the resize procedure.
As for other structures, BufferBlocks, BufferDescriptors and
BufferIOCVArray are all buffer indexed, so making sure shared memory
resizing works for buffers should automatically mean the same for the
rest. But CkptBufferIds is a different case, as it collects buffers to
sync and process them at later point in time -- it has to be explicitely
handled when shrinking shared memory I guess.
Long story short, in the next version of the patch I'll try to
experiment with a simplified design: a simple function to trigger
resizing, launching a coordinator worker, with backends not waiting for
each other and buffers first allocated and then marked as "available to
use".
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 04:55 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-07-14 04:55 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
On Mon, Jul 14, 2025 at 12:07 AM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Sun, Jul 06, 2025 at 03:21:08PM +0200, Dmitry Dolgov wrote:
> > > On Sun, Jul 06, 2025 at 03:01:34PM +0200, Dmitry Dolgov wrote:
> > > * This way any backend between the ProcSignalBarriers will be able
> > > proceed with whatever it's doing, and there is need to make sure it
> > > will not access buffers that will soon disappear. A suggestion so far
> > > was to get all backends agree to not allocate any new buffers in the
> > > to-be-truncated range, but accessing already existing buffers that
> > > will soon go away is a problem as well. As far as I can tell there is
> > > no rock solid method to make sure a backend doesn't have a reference
> > > to such a buffer somewhere (this was discussed earlier in thre
> > > thread), meaning that either a backend has to wait or buffers have to
> > > be checked every time on access.
> >
> > And sure enough, after I wrote this I've realized there should be no
> > such references after the buffer eviction and prohibiting new buffer
> > allocation. I still need to check it though, because not only buffers,
> > but other shared memory structures (which number depends on NBuffers)
> > will be truncated. But if they will also be handled by the eviction,
> > then maybe everything is just fine.
>
> Pondering more about this topic, I've realized there was one more
> problematic case mentioned by Robert early in the thread, which is
> relatively easy to construct:
>
> * When increasing shared buffers from NBuffers_small to NBuffers_large
> it's possible that one backend already has applied NBuffers_large,
> then allocated a buffer B from (NBuffer_small, NBuffers_large] and put
> it into the buffer lookup table.
>
> * In the meantime another backend still has NBuffers_small, but got
> buffer B from the lookup table.
>
> Currently it's being addressed via every backend waiting for each other,
> but I guess it could be as well managed via handling the freelist, so
> that only "available" buffers will be inserted into the lookup table.
I didn't get how can that be managed by freelist? Buffers are also
allocated through clocksweep, which needs to be managed as well.
> Long story short, in the next version of the patch I'll try to
> experiment with a simplified design: a simple function to trigger
> resizing, launching a coordinator worker, with backends not waiting for
> each other and buffers first allocated and then marked as "available to
> use".
Should all the backends wait between buffer allocation and them being
marked as "available"? I assume that marking them as available means
"declaring the new NBuffers". What about when shrinking the buffers?
Do you plan to make all the backends wait while the coordinator is
evicting buffers?
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-07-14 08:10 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 08:10 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 10:25:51AM +0530, Ashutosh Bapat wrote:
> > Currently it's being addressed via every backend waiting for each other,
> > but I guess it could be as well managed via handling the freelist, so
> > that only "available" buffers will be inserted into the lookup table.
>
> I didn't get how can that be managed by freelist? Buffers are also
> allocated through clocksweep, which needs to be managed as well.
The way it is implemented in the patch right now is new buffers are
added into the freelist right away, when they're initialized by the
virtue of nextFree. What I have in mind is to do this as the last step,
when all backends have confirmed shared memory signal was absorbed. This
would mean that StrategyControll will not return a buffer id from the
freshly allocated range until everything is done, and no such buffer
will be inserted into the buffer lookup table.
You're right of course, a buffer id could be returned from the
ClockSweep and from the custom strategy buffer ring. Buf from what I see
those are picking a buffer from the set of already utilized buffers,
meaning that for a buffer to land there it first has to go through
StrategyControl->firstFreeBuffer, and hence the idea above will be a
requirement for those as well.
> > Long story short, in the next version of the patch I'll try to
> > experiment with a simplified design: a simple function to trigger
> > resizing, launching a coordinator worker, with backends not waiting for
> > each other and buffers first allocated and then marked as "available to
> > use".
>
> Should all the backends wait between buffer allocation and them being
> marked as "available"? I assume that marking them as available means
> "declaring the new NBuffers".
Yep, making buffers available would be equivalent to declaring the new
NBuffers. What I think is important here is to note, that we use two
mechanisms for coordination: the shared structure ShmemControl that
shares the state of operation, and ProcSignal that tells backends to do
something (change the memory mapping). Declaring the new NBuffers could
be done via ShmemControl, atomically applying the new value, instead of
sending a ProcSignal -- this way there is no need for backends to wait,
but StrategyControl would need to use the ShmemControl instead of local
copy of NBuffers. Does it make sense to you?
> What about when shrinking the buffers? Do you plan to make all the
> backends wait while the coordinator is evicting buffers?
No, it was never planned like that, since it could easily end up with
coordinator waiting for the backend to unpin a buffer, and the backend
to wait for a signal from the coordinator.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 08:25 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-07-14 08:25 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
On Mon, Jul 14, 2025 at 1:40 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Mon, Jul 14, 2025 at 10:25:51AM +0530, Ashutosh Bapat wrote:
> > > Currently it's being addressed via every backend waiting for each other,
> > > but I guess it could be as well managed via handling the freelist, so
> > > that only "available" buffers will be inserted into the lookup table.
> >
> > I didn't get how can that be managed by freelist? Buffers are also
> > allocated through clocksweep, which needs to be managed as well.
>
> The way it is implemented in the patch right now is new buffers are
> added into the freelist right away, when they're initialized by the
> virtue of nextFree. What I have in mind is to do this as the last step,
> when all backends have confirmed shared memory signal was absorbed. This
> would mean that StrategyControll will not return a buffer id from the
> freshly allocated range until everything is done, and no such buffer
> will be inserted into the buffer lookup table.
>
> You're right of course, a buffer id could be returned from the
> ClockSweep and from the custom strategy buffer ring. Buf from what I see
> those are picking a buffer from the set of already utilized buffers,
> meaning that for a buffer to land there it first has to go through
> StrategyControl->firstFreeBuffer, and hence the idea above will be a
> requirement for those as well.
That isn't true. A buffer which was never in the free list can still
be picked up by clock sweep. But you are raising a relevant point
about StrategyControl below
>
> > > Long story short, in the next version of the patch I'll try to
> > > experiment with a simplified design: a simple function to trigger
> > > resizing, launching a coordinator worker, with backends not waiting for
> > > each other and buffers first allocated and then marked as "available to
> > > use".
> >
> > Should all the backends wait between buffer allocation and them being
> > marked as "available"? I assume that marking them as available means
> > "declaring the new NBuffers".
>
> Yep, making buffers available would be equivalent to declaring the new
> NBuffers. What I think is important here is to note, that we use two
> mechanisms for coordination: the shared structure ShmemControl that
> shares the state of operation, and ProcSignal that tells backends to do
> something (change the memory mapping). Declaring the new NBuffers could
> be done via ShmemControl, atomically applying the new value, instead of
> sending a ProcSignal -- this way there is no need for backends to wait,
> but StrategyControl would need to use the ShmemControl instead of local
> copy of NBuffers. Does it make sense to you?
When expanding buffers, letting StrategyControl continue with the old
NBuffers may work. When propagating the new buffer value we have to
reinitialize StrategyControl to use new NBuffers. But when shrinking,
the StrategyControl needs to be initialized with the new NBuffers,
lest it picks a victim from buffers being shrunk. And then if the
operation fails, we have to reinitialize the StrategyControl again
with the old NBuffers.
>
> > What about when shrinking the buffers? Do you plan to make all the
> > backends wait while the coordinator is evicting buffers?
>
> No, it was never planned like that, since it could easily end up with
> coordinator waiting for the backend to unpin a buffer, and the backend
> to wait for a signal from the coordinator.
I agree with the deadlock situation. How do we prevent the backends
from picking or continuing to work with a buffer from buffers being
shrunk then? Each backend then has to do something about their
respective pinned buffers.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-07-14 08:54 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 08:54 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 01:55:39PM +0530, Ashutosh Bapat wrote:
> > You're right of course, a buffer id could be returned from the
> > ClockSweep and from the custom strategy buffer ring. Buf from what I see
> > those are picking a buffer from the set of already utilized buffers,
> > meaning that for a buffer to land there it first has to go through
> > StrategyControl->firstFreeBuffer, and hence the idea above will be a
> > requirement for those as well.
>
> That isn't true. A buffer which was never in the free list can still
> be picked up by clock sweep.
How's that?
> > Yep, making buffers available would be equivalent to declaring the new
> > NBuffers. What I think is important here is to note, that we use two
> > mechanisms for coordination: the shared structure ShmemControl that
> > shares the state of operation, and ProcSignal that tells backends to do
> > something (change the memory mapping). Declaring the new NBuffers could
> > be done via ShmemControl, atomically applying the new value, instead of
> > sending a ProcSignal -- this way there is no need for backends to wait,
> > but StrategyControl would need to use the ShmemControl instead of local
> > copy of NBuffers. Does it make sense to you?
>
> When expanding buffers, letting StrategyControl continue with the old
> NBuffers may work. When propagating the new buffer value we have to
> reinitialize StrategyControl to use new NBuffers. But when shrinking,
> the StrategyControl needs to be initialized with the new NBuffers,
> lest it picks a victim from buffers being shrunk. And then if the
> operation fails, we have to reinitialize the StrategyControl again
> with the old NBuffers.
Right, those two cases will become more asymmetrical: for expanding
number of available buffers would have to be propagated to the backends
at the end, when they're ready; for shrinking number of available
buffers would have to be propagated at the start, so that backends will
stop allocating unavailable buffers.
> > > What about when shrinking the buffers? Do you plan to make all the
> > > backends wait while the coordinator is evicting buffers?
> >
> > No, it was never planned like that, since it could easily end up with
> > coordinator waiting for the backend to unpin a buffer, and the backend
> > to wait for a signal from the coordinator.
>
> I agree with the deadlock situation. How do we prevent the backends
> from picking or continuing to work with a buffer from buffers being
> shrunk then? Each backend then has to do something about their
> respective pinned buffers.
The idea I've got so far is stop allocating buffers from the unavailable
range and wait until backends will unpin all unavailable buffers. We
either wait unconditionally until it happens, or bail out after certain
timeout.
It's probably possible to force backends to unpin buffers they work
with, but it sounds much more problematic to me. What do you think?
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 09:24 ` Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Thom Brown @ 2025-07-14 09:24 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
On Mon, 14 Jul 2025, 09:54 Dmitry Dolgov, <9erthalion6@gmail.com> wrote:
> > On Mon, Jul 14, 2025 at 01:55:39PM +0530, Ashutosh Bapat wrote:
> > > You're right of course, a buffer id could be returned from the
> > > ClockSweep and from the custom strategy buffer ring. Buf from what I
> see
> > > those are picking a buffer from the set of already utilized buffers,
> > > meaning that for a buffer to land there it first has to go through
> > > StrategyControl->firstFreeBuffer, and hence the idea above will be a
> > > requirement for those as well.
> >
> > That isn't true. A buffer which was never in the free list can still
> > be picked up by clock sweep.
>
> How's that?
>
Isn't it its job to find usable buffers from the used buffer list when no
free ones are available? The next victim buffer can be selected (and
cleaned if dirty) and then immediately used without touching the free list.
Thom
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
@ 2025-07-14 09:32 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 09:32 UTC (permalink / raw)
To: Thom Brown <thom@linux.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 10:24:50AM +0100, Thom Brown wrote:
> On Mon, 14 Jul 2025, 09:54 Dmitry Dolgov, <9erthalion6@gmail.com> wrote:
>
> > > On Mon, Jul 14, 2025 at 01:55:39PM +0530, Ashutosh Bapat wrote:
> > > > You're right of course, a buffer id could be returned from the
> > > > ClockSweep and from the custom strategy buffer ring. Buf from what I
> > see
> > > > those are picking a buffer from the set of already utilized buffers,
> > > > meaning that for a buffer to land there it first has to go through
> > > > StrategyControl->firstFreeBuffer, and hence the idea above will be a
> > > > requirement for those as well.
> > >
> > > That isn't true. A buffer which was never in the free list can still
> > > be picked up by clock sweep.
> >
> > How's that?
> >
>
> Isn't it its job to find usable buffers from the used buffer list when no
> free ones are available? The next victim buffer can be selected (and
> cleaned if dirty) and then immediately used without touching the free list.
Ah, I see what you mean folks. But I'm talking here only about buffers
which will be allocated after extending shared memory -- they must go
through the freelist first (I don't see why not, any other options?),
and clock sweep will have a chance to pick them up only afterwards. That
makes the freelist sort of an entry point for those buffers.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 12:56 ` Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Andres Freund @ 2025-07-14 12:56 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi,
On 2025-07-14 11:32:25 +0200, Dmitry Dolgov wrote:
> > On Mon, Jul 14, 2025 at 10:24:50AM +0100, Thom Brown wrote:
> > On Mon, 14 Jul 2025, 09:54 Dmitry Dolgov, <9erthalion6@gmail.com> wrote:
> >
> > > > On Mon, Jul 14, 2025 at 01:55:39PM +0530, Ashutosh Bapat wrote:
> > > > > You're right of course, a buffer id could be returned from the
> > > > > ClockSweep and from the custom strategy buffer ring. Buf from what I
> > > see
> > > > > those are picking a buffer from the set of already utilized buffers,
> > > > > meaning that for a buffer to land there it first has to go through
> > > > > StrategyControl->firstFreeBuffer, and hence the idea above will be a
> > > > > requirement for those as well.
> > > >
> > > > That isn't true. A buffer which was never in the free list can still
> > > > be picked up by clock sweep.
> > >
> > > How's that?
> > >
> >
> > Isn't it its job to find usable buffers from the used buffer list when no
> > free ones are available? The next victim buffer can be selected (and
> > cleaned if dirty) and then immediately used without touching the free list.
>
> Ah, I see what you mean folks. But I'm talking here only about buffers
> which will be allocated after extending shared memory -- they must go
> through the freelist first (I don't see why not, any other options?),
> and clock sweep will have a chance to pick them up only afterwards. That
> makes the freelist sort of an entry point for those buffers.
Clock sweep can find any buffer, independent of whether it's on the freelist.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-07-14 13:08 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 13:08 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 08:56:56AM -0400, Andres Freund wrote:
> > Ah, I see what you mean folks. But I'm talking here only about buffers
> > which will be allocated after extending shared memory -- they must go
> > through the freelist first (I don't see why not, any other options?),
> > and clock sweep will have a chance to pick them up only afterwards. That
> > makes the freelist sort of an entry point for those buffers.
>
> Clock sweep can find any buffer, independent of whether it's on the freelist.
It does the search based on nextVictimBuffer, where the actual buffer
will be a modulo of NBuffers, right? If that's correct and I get
everything else right, that would mean as long as NBuffers stays the
same (which is the case for the purposes of the current discussion) new
buffers, allocated on top of NBuffers after shared memory increase, will
not be picked by the clock sweep.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 13:14 ` Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Andres Freund @ 2025-07-14 13:14 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi,
On 2025-07-14 15:08:28 +0200, Dmitry Dolgov wrote:
> > On Mon, Jul 14, 2025 at 08:56:56AM -0400, Andres Freund wrote:
> > > Ah, I see what you mean folks. But I'm talking here only about buffers
> > > which will be allocated after extending shared memory -- they must go
> > > through the freelist first (I don't see why not, any other options?),
> > > and clock sweep will have a chance to pick them up only afterwards. That
> > > makes the freelist sort of an entry point for those buffers.
> >
> > Clock sweep can find any buffer, independent of whether it's on the freelist.
>
> It does the search based on nextVictimBuffer, where the actual buffer
> will be a modulo of NBuffers, right? If that's correct and I get
> everything else right, that would mean as long as NBuffers stays the
> same (which is the case for the purposes of the current discussion) new
> buffers, allocated on top of NBuffers after shared memory increase, will
> not be picked by the clock sweep.
Are you tell me that you'd put "new" buffers onto the freelist, before you
increase NBuffers? That doesn't make sense.
Orthogonaly - there's discussion about simply removing the freelist.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-07-14 13:20 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 13:20 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 09:14:26AM -0400, Andres Freund wrote:
> > > Clock sweep can find any buffer, independent of whether it's on the freelist.
> >
> > It does the search based on nextVictimBuffer, where the actual buffer
> > will be a modulo of NBuffers, right? If that's correct and I get
> > everything else right, that would mean as long as NBuffers stays the
> > same (which is the case for the purposes of the current discussion) new
> > buffers, allocated on top of NBuffers after shared memory increase, will
> > not be picked by the clock sweep.
>
> Are you tell me that you'd put "new" buffers onto the freelist, before you
> increase NBuffers? That doesn't make sense.
Why?
> Orthogonaly - there's discussion about simply removing the freelist.
Good to know, will take a look at that thread, thanks.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 13:42 ` Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Andres Freund @ 2025-07-14 13:42 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi,
On 2025-07-14 15:20:03 +0200, Dmitry Dolgov wrote:
> > On Mon, Jul 14, 2025 at 09:14:26AM -0400, Andres Freund wrote:
> > > > Clock sweep can find any buffer, independent of whether it's on the freelist.
> > >
> > > It does the search based on nextVictimBuffer, where the actual buffer
> > > will be a modulo of NBuffers, right? If that's correct and I get
> > > everything else right, that would mean as long as NBuffers stays the
> > > same (which is the case for the purposes of the current discussion) new
> > > buffers, allocated on top of NBuffers after shared memory increase, will
> > > not be picked by the clock sweep.
> >
> > Are you tell me that you'd put "new" buffers onto the freelist, before you
> > increase NBuffers? That doesn't make sense.
>
> Why?
I think it basically boils down to "That's not how it supposed to work".
If you have buffers that are not in the clock sweep they'll get unfairly high
usage counts, as their usecount won't be decremented by the clock
sweep. Resulting in those buffers potentially being overly sticky after the
s_b resize completed.
It breaks the entirely reasonable check to verify that a buffer returned by
StrategyGetBuffer() is within the buffer pool.
Obviously, if we remove the freelist, not having the clock sweep find the
buffer would mean it's unreachable.
What on earth would be the point of putting a buffer on the freelist but not
make it reachable by the clock sweep? To me that's just nonsensical.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-07-14 14:01 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:22 ` Re: Changing shared_buffers without restart Burd, Greg <greg@burd.me>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
0 siblings, 2 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 14:01 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 09:42:46AM -0400, Andres Freund wrote:
> What on earth would be the point of putting a buffer on the freelist but not
> make it reachable by the clock sweep? To me that's just nonsensical.
To clarify, we're not talking about this scenario as "that's how it
would work after the resize". The point is that to expand shared buffers
they need to be initialized, included into the whole buffer machinery
(freelist, clock sweep, etc.) and NBuffers has to be updated. Those
steps are separated in time, and I'm currently trying to understand what
are the consequences of performing them in different order and whether
there are possible concurrency issues under various scenarios. Does this
make more sense, or still not?
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 14:22 ` Burd, Greg <greg@burd.me>
2025-07-14 14:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Burd, Greg @ 2025-07-14 14:22 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Jul 14, 2025, at 10:01 AM, Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
>> On Mon, Jul 14, 2025 at 09:42:46AM -0400, Andres Freund wrote:
>> What on earth would be the point of putting a buffer on the freelist but not
>> make it reachable by the clock sweep? To me that's just nonsensical.
>
> To clarify, we're not talking about this scenario as "that's how it
> would work after the resize". The point is that to expand shared buffers
> they need to be initialized, included into the whole buffer machinery
> (freelist, clock sweep, etc.) and NBuffers has to be updated. Those
> steps are separated in time, and I'm currently trying to understand what
> are the consequences of performing them in different order and whether
> there are possible concurrency issues under various scenarios. Does this
> make more sense, or still not?
Hello, first off thanks for working on the intricate issues related to resizing
shared_buffers.
Second, I'm new in this code so take that in account but I'm the person trying
to remove the freelist entirely [1] so I have reviewed this code recently.
I'd initialize them, expand BufferDescriptors, and adjust NBuffers. The
clock-sweep algorithm will eventually find them and make use of them. The
buf->freeNext should be FREENEXT_NOT_IN_LIST so that StrategyFreeBuffer() will
do the work required to append it the freelist after use. AFAICT there is no
need to add to the freelist up front.
best.
-greg
[1] https://postgr.es/m/flat/E2D6FCDC-BE98-4F95-B45E-699C3E17BA10%40burd.me
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:22 ` Re: Changing shared_buffers without restart Burd, Greg <greg@burd.me>
@ 2025-07-14 14:43 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 14:43 UTC (permalink / raw)
To: Burd, Greg <greg@burd.me>; +Cc: Andres Freund <andres@anarazel.de>; Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 10:22:17AM -0400, Burd, Greg wrote:
> I'd initialize them, expand BufferDescriptors, and adjust NBuffers. The
> clock-sweep algorithm will eventually find them and make use of them. The
> buf->freeNext should be FREENEXT_NOT_IN_LIST so that StrategyFreeBuffer() will
> do the work required to append it the freelist after use. AFAICT there is no
> need to add to the freelist up front.
Yep, thanks. I think this approach may lead to a problem I'm trying to
address with the buffer lookup table (just have described it in the
message above). But if I'm wrong, that of course would be the way to go.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 14:23 ` Andres Freund <andres@anarazel.de>
2025-07-14 14:39 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:10 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
1 sibling, 2 replies; 167+ messages in thread
From: Andres Freund @ 2025-07-14 14:23 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi,
On 2025-07-14 16:01:50 +0200, Dmitry Dolgov wrote:
> > On Mon, Jul 14, 2025 at 09:42:46AM -0400, Andres Freund wrote:
> > What on earth would be the point of putting a buffer on the freelist but not
> > make it reachable by the clock sweep? To me that's just nonsensical.
>
> To clarify, we're not talking about this scenario as "that's how it
> would work after the resize". The point is that to expand shared buffers
> they need to be initialized, included into the whole buffer machinery
> (freelist, clock sweep, etc.) and NBuffers has to be updated.
It seems pretty obvious to that the order has to be
1) initialize buffer headers
2) update NBuffers
3) put them onto the freelist
(with 3) hopefully becoming obsolete)
> Those steps are separated in time, and I'm currently trying to understand
> what are the consequences of performing them in different order and whether
> there are possible concurrency issues under various scenarios. Does this
> make more sense, or still not?
I still don't understand why it'd ever make sense to put a buffer onto the
freelist before updating NBuffers first.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-07-14 14:39 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:11 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
1 sibling, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 14:39 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 10:23:23AM -0400, Andres Freund wrote:
> > Those steps are separated in time, and I'm currently trying to understand
> > what are the consequences of performing them in different order and whether
> > there are possible concurrency issues under various scenarios. Does this
> > make more sense, or still not?
>
> I still don't understand why it'd ever make sense to put a buffer onto the
> freelist before updating NBuffers first.
Depending on how NBuffers is updated, different backends may have
different value of NBuffers for a short time frame. In that case a
scenario I'm trying to address is when one backend with the new NBuffers
value allocates a new buffer and puts it into the buffer lookup table,
where it could become reachable by another backend, which still has the
old NBuffer value. Correct me if I'm wrong, but initializing buffer
headers + updating NBuffers means clock sweep can now return one of
those new buffers, opening the scenario above, right?
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:39 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 15:11 ` Andres Freund <andres@anarazel.de>
2025-07-14 15:18 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
2025-07-14 15:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Andres Freund @ 2025-07-14 15:11 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi,
On July 14, 2025 10:39:33 AM EDT, Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>> On Mon, Jul 14, 2025 at 10:23:23AM -0400, Andres Freund wrote:
>> > Those steps are separated in time, and I'm currently trying to understand
>> > what are the consequences of performing them in different order and whether
>> > there are possible concurrency issues under various scenarios. Does this
>> > make more sense, or still not?
>>
>> I still don't understand why it'd ever make sense to put a buffer onto the
>> freelist before updating NBuffers first.
>
>Depending on how NBuffers is updated, different backends may have
>different value of NBuffers for a short time frame. In that case a
>scenario I'm trying to address is when one backend with the new NBuffers
>value allocates a new buffer and puts it into the buffer lookup table,
>where it could become reachable by another backend, which still has the
>old NBuffer value. Correct me if I'm wrong, but initializing buffer
>headers + updating NBuffers means clock sweep can now return one of
>those new buffers, opening the scenario above, right?
The same is true if you put buffers into the freelist.
Andres
--
Sent from my Android device with K-9 Mail. Please excuse my brevity.
^ permalink raw reply [nested|flat] 167+ messages in thread
* RE: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:39 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:11 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-07-14 15:18 ` Jack Ng <Jack.Ng@huawei.com>
2025-07-14 16:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Jack Ng @ 2025-07-14 15:18 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Ni Ku <jakkuniku@gmail.com>
Just brain-storming here... would moving NBuffers to shared memory solve this specific issue? Though I'm pretty sure that would open up a new set of synchronization issues elsewhere, so I'm not sure if there's a net gain.
Jack
>-----Original Message-----
>From: Andres Freund <andres@anarazel.de>
>Sent: Monday, July 14, 2025 11:12 AM
>To: Dmitry Dolgov <9erthalion6@gmail.com>
>Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat
><ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>;
>Thomas Munro <thomas.munro@gmail.com>; PostgreSQL-development <pgsql-
>hackers@postgresql.org>; Jack Ng <Jack.Ng@huawei.com>; Ni Ku
><jakkuniku@gmail.com>
>Subject: Re: Changing shared_buffers without restart
>
>Hi,
>
>On July 14, 2025 10:39:33 AM EDT, Dmitry Dolgov <9erthalion6@gmail.com>
>wrote:
>>> On Mon, Jul 14, 2025 at 10:23:23AM -0400, Andres Freund wrote:
>>> > Those steps are separated in time, and I'm currently trying to
>>> > understand what are the consequences of performing them in
>>> > different order and whether there are possible concurrency issues
>>> > under various scenarios. Does this make more sense, or still not?
>>>
>>> I still don't understand why it'd ever make sense to put a buffer
>>> onto the freelist before updating NBuffers first.
>>
>>Depending on how NBuffers is updated, different backends may have
>>different value of NBuffers for a short time frame. In that case a
>>scenario I'm trying to address is when one backend with the new
>>NBuffers value allocates a new buffer and puts it into the buffer
>>lookup table, where it could become reachable by another backend, which
>>still has the old NBuffer value. Correct me if I'm wrong, but
>>initializing buffer headers + updating NBuffers means clock sweep can
>>now return one of those new buffers, opening the scenario above, right?
>
>The same is true if you put buffers into the freelist.
>
>Andres
>--
>Sent from my Android device with K-9 Mail. Please excuse my brevity.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:39 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:11 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 15:18 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
@ 2025-07-14 16:32 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-15 22:52 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 16:32 UTC (permalink / raw)
To: Jack Ng <Jack.Ng@huawei.com>; +Cc: Andres Freund <andres@anarazel.de>; Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 03:18:10PM +0000, Jack Ng wrote:
> Just brain-storming here... would moving NBuffers to shared memory solve this specific issue? Though I'm pretty sure that would open up a new set of synchronization issues elsewhere, so I'm not sure if there's a net gain.
It's in fact already happening, there is a shared structure that
described the resize status. But if I get everything right, it doesn't
solve all the problems.
^ permalink raw reply [nested|flat] 167+ messages in thread
* RE: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:39 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:11 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 15:18 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
2025-07-14 16:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-15 22:52 ` Jack Ng <Jack.Ng@huawei.com>
2025-07-16 14:48 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Jack Ng @ 2025-07-15 22:52 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Ni Ku <jakkuniku@gmail.com>
>> On Mon, Jul 14, 2025 at 03:18:10PM +0000, Jack Ng wrote:
>> Just brain-storming here... would moving NBuffers to shared memory solve
>this specific issue? Though I'm pretty sure that would open up a new set of
>synchronization issues elsewhere, so I'm not sure if there's a net gain.
>
>It's in fact already happening, there is a shared structure that described the
>resize status. But if I get everything right, it doesn't solve all the problems.
Hi Dmitry,
Just to clarify, you're not only referring to the ShmemControl::NSharedBuffers
and related logic in the current patches, but actually getting rid of per-process
NBuffers completely and use ShmemControl::NSharedBuffers everywhere instead (or
something along those lines)? So that when the coordinator updates
ShmemControl::NSharedBuffers, everyone sees the new value right away.
I guess this is part of the "simplified design" you mentioned several posts earlier?
I also thought about that approach more, and there seems to be new synchronization
issues we would need to deal with, like:
1. Mid-execution change of NBuffers in functions like BufferSync and BgBufferSync,
which could cause correctness and performance issues. I suppose most of them
are solvable with atomics and shared r/w locks etc, but at the cost of higher
performance overheads.
2. NBuffers becomes inconsistent with the underlying shared memory mappings for a
period of time for each process. Currently both are updated in AnonymousShmemResize
and AdjustShmemSize "atomically" for a process, so I wonder if letting them get
out-of-sync (even for a brief period) could be problematic.
I agree it doesn't seem to solve all the problems. It can simplify certain aspects
of the design, but may also introduce new issues. Overall not a "silver bullet" :)
Jack
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:39 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:11 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 15:18 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
2025-07-14 16:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-15 22:52 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
@ 2025-07-16 14:48 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-16 14:48 UTC (permalink / raw)
To: Jack Ng <Jack.Ng@huawei.com>; +Cc: Andres Freund <andres@anarazel.de>; Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Ni Ku <jakkuniku@gmail.com>
> On Tue, Jul 15, 2025 at 10:52:01PM +0000, Jack Ng wrote:
> >> On Mon, Jul 14, 2025 at 03:18:10PM +0000, Jack Ng wrote:
> >> Just brain-storming here... would moving NBuffers to shared memory solve
> >this specific issue? Though I'm pretty sure that would open up a new set of
> >synchronization issues elsewhere, so I'm not sure if there's a net gain.
> >
> >It's in fact already happening, there is a shared structure that described the
> >resize status. But if I get everything right, it doesn't solve all the problems.
>
> Just to clarify, you're not only referring to the ShmemControl::NSharedBuffers
> and related logic in the current patches, but actually getting rid of per-process
> NBuffers completely and use ShmemControl::NSharedBuffers everywhere instead (or
> something along those lines)? So that when the coordinator updates
> ShmemControl::NSharedBuffers, everyone sees the new value right away.
> I guess this is part of the "simplified design" you mentioned several posts earlier?
I was thinking more about something like NBuffersAvailable, which would
control how victim buffers are getting picked, but there is a spectrum
of different options to experiment with.
> I also thought about that approach more, and there seems to be new synchronization
> issues we would need to deal with, like:
Potentially tricky change of NBuffers already happens in the current
patch set, e.g. NBuffers is getting updated in ProcessProcSignalBarrier,
which is called at the end of BufferSync loop iteration. By itself I
don't see any obvious problems here except remembering buffer id in
CkptBufferIds (I've mentioned this few messages above).
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:39 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:11 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-07-14 15:35 ` Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-14 15:35 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 11:11:36AM -0400, Andres Freund wrote:
> Hi,
>
> On July 14, 2025 10:39:33 AM EDT, Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> >> On Mon, Jul 14, 2025 at 10:23:23AM -0400, Andres Freund wrote:
> >> > Those steps are separated in time, and I'm currently trying to understand
> >> > what are the consequences of performing them in different order and whether
> >> > there are possible concurrency issues under various scenarios. Does this
> >> > make more sense, or still not?
> >>
> >> I still don't understand why it'd ever make sense to put a buffer onto the
> >> freelist before updating NBuffers first.
> >
> >Depending on how NBuffers is updated, different backends may have
> >different value of NBuffers for a short time frame. In that case a
> >scenario I'm trying to address is when one backend with the new NBuffers
> >value allocates a new buffer and puts it into the buffer lookup table,
> >where it could become reachable by another backend, which still has the
> >old NBuffer value. Correct me if I'm wrong, but initializing buffer
> >headers + updating NBuffers means clock sweep can now return one of
> >those new buffers, opening the scenario above, right?
>
> The same is true if you put buffers into the freelist.
Yep, but the question about clock sweep still stays. Anyway, thanks for
the input, let me digest it and come up with more questions & patch
series.
^ permalink raw reply [nested|flat] 167+ messages in thread
* RE: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Re: Changing shared_buffers without restart Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-07-14 15:10 ` Jack Ng <Jack.Ng@huawei.com>
1 sibling, 0 replies; 167+ messages in thread
From: Jack Ng @ 2025-07-14 15:10 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers; Ni Ku <jakkuniku@gmail.com>
If I understanding correctly, putting a new buffer in the freelist before updating NBuffers could break existing logic that calls BufferIsValid(bufnum) and asserts bufnum <= NBuffers? (since a backend can grab the new buffer and checks its validity before the coordinator can add it to the freelist.)
But it seems updating NBuffers before adding new elements to the freelist could be problematic too? Like if a new buffer is already chosen as a victim and then the coordinator adds it to the freelist, would that lead to "double-use"? (seems possible at least with current logic and serialization in StrategyGetBuffer). If that's a valid concern, would something like this work?
1) initialize buffer headers, with a new state/flag to indicate "add-pending"
2) update NBuffers
-- add a check in clock-sweep logic for "add-pending" and skip them
3) put them onto the freelist
4) when a new element is grabbed from freelist, check for and reset add-pending flag.
This ensure the new element is always obtained from the freelist first I think.
Jack
>-----Original Message-----
>From: Andres Freund <andres@anarazel.de>
>Sent: Monday, July 14, 2025 10:23 AM
>To: Dmitry Dolgov <9erthalion6@gmail.com>
>Cc: Thom Brown <thom@linux.com>; Ashutosh Bapat
><ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>;
>Thomas Munro <thomas.munro@gmail.com>; PostgreSQL-development <pgsql-
>hackers@postgresql.org>; Jack Ng <Jack.Ng@huawei.com>; Ni Ku
><jakkuniku@gmail.com>
>Subject: Re: Changing shared_buffers without restart
>
>Hi,
>
>On 2025-07-14 16:01:50 +0200, Dmitry Dolgov wrote:
>> > On Mon, Jul 14, 2025 at 09:42:46AM -0400, Andres Freund wrote:
>> > What on earth would be the point of putting a buffer on the freelist
>> > but not make it reachable by the clock sweep? To me that's just nonsensical.
>>
>> To clarify, we're not talking about this scenario as "that's how it
>> would work after the resize". The point is that to expand shared
>> buffers they need to be initialized, included into the whole buffer
>> machinery (freelist, clock sweep, etc.) and NBuffers has to be updated.
>
>It seems pretty obvious to that the order has to be
>
>1) initialize buffer headers
>2) update NBuffers
>3) put them onto the freelist
>
>(with 3) hopefully becoming obsolete)
>
>
>> Those steps are separated in time, and I'm currently trying to
>> understand what are the consequences of performing them in different
>> order and whether there are possible concurrency issues under various
>> scenarios. Does this make more sense, or still not?
>
>I still don't understand why it'd ever make sense to put a buffer onto the freelist
>before updating NBuffers first.
>
>Greetings,
>
>Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-07-14 22:55 ` Jim Nasby <jnasby@upgrade.com>
2025-07-16 14:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-16 15:44 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2 siblings, 2 replies; 167+ messages in thread
From: Jim Nasby @ 2025-07-14 22:55 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
On Fri, Jul 4, 2025 at 9:42 AM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > On Fri, Jul 04, 2025 at 02:06:16AM +0200, Tomas Vondra wrote:
>
...
> > 10) what to do about stuck resize?
> >
> > AFAICS the resize can get stuck for various reasons, e.g. because it
> > can't evict pinned buffers, possibly indefinitely. Not great, it's not
> > clear to me if there's a way out (canceling the resize) after a timeout,
> > or something like that? Not great to start an "online resize" only to
> > get stuck with all activity blocked for indefinite amount of time, and
> > get to restart anyway.
> >
> > Seems related to Thomas' message [2], but AFAICS the patch does not do
> > anything about this yet, right? What's the plan here?
>
> It's another open discussion right now, with an idea to eventually allow
> canceling after a timeout. I think canceling when stuck on buffer
> eviction should be pretty straightforward (the evition must take place
> before actual shared memory resize, so we know nothing has changed yet),
> but in some other failure scenarios it would be harder (e.g. if one
> backend is stuck resizing, while other have succeeded -- this would
> require another round of synchronization and some way to figure out what
> is the current status).
From a user standpoint, I would expect any kind of resize like this to be
an online operation that happens in the background. If this is driven by a
GUC I don't see how it could be anything else, but if something else is
decided on I think it'd just be pain to require a session to stay connected
until a resize was complete. (Of course we'd need to provide some means of
monitoring a resize that was in-process, perhaps via a pg_stat_progress
view or a system function.)
Also, while I haven't fully followed discussion about how to synchronize
backends, I will say that I don't think it's at all unreasonable if a
resize doesn't take full effect until every backend has at minimum ended
any running transaction, or potentially even returned back to the
equivalent of `PostgresMain()` for that type of backend. Obviously it'd be
nicer to be more responsive than that, but I don't think the first version
of the feature has to accomplish that.
For that matter, I also feel it'd be fine if the first version didn't even
support shrinking shared buffers.
Finally, while shared buffers is the most visible target here, there are
other shared memory settings that have a *much* smaller surface area, and
in my experience are going to be much more valuable from a tuning
perspective; notably wal_buffers and the MXID SLRUs (and possibly CLOG and
subtrans). I say that because unless you're running a workload that
entirely fits in shared buffers, or a *really* small shared buffers
compared to system memory, increasing shared buffers quickly gets into
diminishing returns. But since the default size for the other fixed sized
areas is so much smaller than normal values for shared_buffers, increasing
those areas can have a much, much larger impact on performance. (Especially
for something like the MXID SLRUs.) I would certainly consider focusing on
one of those areas before trying to tackle shared buffers.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 22:55 ` Re: Changing shared_buffers without restart Jim Nasby <jnasby@upgrade.com>
@ 2025-07-16 14:52 ` Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-07-16 14:52 UTC (permalink / raw)
To: Jim Nasby <jnasby@upgrade.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
> On Mon, Jul 14, 2025 at 05:55:13PM -0500, Jim Nasby wrote:
>
> Finally, while shared buffers is the most visible target here, there are
> other shared memory settings that have a *much* smaller surface area, and
> in my experience are going to be much more valuable from a tuning
> perspective; notably wal_buffers and the MXID SLRUs (and possibly CLOG and
> subtrans). I say that because unless you're running a workload that
> entirely fits in shared buffers, or a *really* small shared buffers
> compared to system memory, increasing shared buffers quickly gets into
> diminishing returns. But since the default size for the other fixed sized
> areas is so much smaller than normal values for shared_buffers, increasing
> those areas can have a much, much larger impact on performance. (Especially
> for something like the MXID SLRUs.) I would certainly consider focusing on
> one of those areas before trying to tackle shared buffers.
That's an interesting idea, thanks for sharing. The reason I'm
concentrating on shared buffers is that it was frequently called out as
a problem when trying to tune PostgreSQL automatically. In this context
shared buffers is usually one of the most impactful knobs, yet one of
the most painful to manage as well. But if the amount of complexity
around resizable shared buffers will be proved unsurmountable, yeah, it
would make sense to consider simpler targets using the same mechanism.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 22:55 ` Re: Changing shared_buffers without restart Jim Nasby <jnasby@upgrade.com>
@ 2025-07-16 15:44 ` Andres Freund <andres@anarazel.de>
1 sibling, 0 replies; 167+ messages in thread
From: Andres Freund @ 2025-07-16 15:44 UTC (permalink / raw)
To: Jim Nasby <jnasby@upgrade.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; Tomas Vondra <tomas@vondra.me>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi,
On 2025-07-14 17:55:13 -0500, Jim Nasby wrote:
> I say that because unless you're running a workload that entirely fits in
> shared buffers, or a *really* small shared buffers compared to system
> memory, increasing shared buffers quickly gets into diminishing returns.
I don't think that's true, at all, today. And it certainly won't be true in a
world where we will be able to use direct_io for real workloads.
Particularly for write heavy workloads, the difference between a small buffer
pool and a large one can be *dramatic*, because the large buffer pool allows
most writes to be done by checkpointer (and thus largely sequentially) or by
backends and bgwriter (and thus largely randomly). Doing more writes
sequentially helps with short-term performance, but *particularly* helps with
sustained performance on SSDs. A larger buffer pool also reduces the *total*
number of writes dramatically, because the same buffer will often be dirtied
repeatedly within one checkpoint window.
r/w/ pgbench is a workload that *undersells* the benefit of a larger
shared_buffers, as each transaction is uncommonly small, making WAL flushes
much more of a bottleneck (the access pattern is too uniform, too). But even
for that the difference can be massive:
A scale 500 pgbench with 48 clients:
s_b= 512MB:
averages 390MB/s of writes in steady state
average TPS: 25072
s_b=8192MB:
averages 48MB/s of writes in steady state
average TPS: 47901
Nearly an order of magnitude difference in writes and nearly a 2x difference
in TPS.
25%, the advice we give for shared_buffers, is literally close to the worst
possible value. The only thing it maximizes is double buffering. While
removing information useful about what to cache for how long from both
postgres and the OS, leading to reduced cache hit rates.
> But since the default size for the other fixed sized areas is so much
> smaller than normal values for shared_buffers, increasing those areas can
> have a much, much larger impact on performance. (Especially for something
> like the MXID SLRUs.) I would certainly consider focusing on one of those
> areas before trying to tackle shared buffers.
I think that'd be a bad idea. There's simply no point in having the complexity
in place to allow for dynamically resizing a few megabytes of buffers. You
just configure them large enough (including probalby increasing some of the
defaults one of these years). Whereas you can't just do that for
shared_buffers, as we're talking really memory. Ahead of time you do not know
how much memory backends themselves need and the amount of memory in the
system may change.
Resizing shared_buffers is particularly important because it's becoming more
important to be able to dynamically increase/decrease the resources of a
running postgres instance to adjust for system load. Memory and CPUs can be
hot added/removed from VMs, but we need to utilize them...
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Re: Changing shared_buffers without restart Tomas Vondra <tomas@vondra.me>
@ 2025-09-18 04:47 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-09-18 04:47 UTC (permalink / raw)
To: Tomas Vondra <tomas@vondra.me>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Jack Ng <Jack.Ng@huawei.com>; Ni Ku <jakkuniku@gmail.com>
Hi Tomas,
Thanks for your detailed feedback. Sorry for replying late.
On Fri, Jul 4, 2025 at 5:36 AM Tomas Vondra <tomas@vondra.me> wrote:
>
> v5-0008-Support-shrinking-shared-buffers.patch
>
> 1) Why is ShmemCtrl->evictor_pid reset in AnonymousShmemResize? Isn't
> there a place starting it and waiting for it to complete? Why
> shouldn't it do EvictExtraBuffers itself?
>
> 3) Seems a bit strange to do it from a random backend. Shouldn't it
> be the responsibility of a process like checkpointer/bgwriter, or
> maybe a dedicated dynamic bgworker? Can we even rely on a backend
> to be available?
I will answer these two together. I don't think we should rely on a
random backend. But that's what the rest of the patches did and
patches to support shrinking followed them. But AFAIK, Dmitry is
working on a set of changes which will make a non-postmaster backend
to be a coordinator for buffer pool resizing process. When that
happens the same backend which initializes the expanded memory when
expanding the buffer pool should also be responsible for evicting the
buffers when shrinking the buffer pool. Will wait for Dmitry's next
set of patches before making this change.
>
> 4) Unsolved issues with buffers pinned for a long time. Could be an
> issue if the buffer is pinned indefinitely (e.g. cursor in idle
> connection), and the resizing blocks some activity (new connections
> or stuff like that).
In such cases we should cancel the operation or kill that backend (per
user preference) after a timeout with (user specified) timeout >= 0.
We haven't yet figured out the details. I think the first version of
the feature would just cancel the operation, if it encounters a pinned
buffer.
> 2) Isn't the change to BufferManagerShmemInit wrong? How do we know the
> last buffer is still at the end of the freelist? Seems unlikely.
> 6) It's not clear to me in what situations this triggers (in the call
> to BufferManagerShmemInit)
>
> if (FirstBufferToInit < NBuffers) ...
>
Will answer these two together. As the comment says FirstBufferToInit
< NBuffers indicates two situations: When FirstBufferToInit = 0, it's
the first time the buffer pool is being initialized. Otherwise it
indicates expanding the buffer pool, in which case the last buffer
will be a newly initialized buffer. All newly initialized buffers are
linked into the freelist one after the other in the increasing order
of their buffer ids by code a few lines above. Now that the free
buffer list has been removed, we don't need to worry about it. In the
next set of patches, I have removed this code.
>
> v5-0009-Reinitialize-StrategyControl-after-resizing-buffe.patch
>
> 1) IMHO this should be included in the earlier resize/shrink patches,
> I don't see a reason to keep it separate (assuming this is the
> correct way, and the "init" is not).
These patches are separate just because me and Dmitry developed them
respectively. Once they are reviewed by Dmitry, we will squash them
into a single patch. I am expecting that Dmitry's next patchset which
will do significant changes to the synchronization will have a single
patch for all code related to and consequential to resizing.
>
> 5) Funny that "AI suggests" something, but doesn't the block fail to
> reset nextVictimBuffer of the clocksweep? It may point to a buffer
> we're removing, and it'll be invalid, no?
>
The TODO no more applies. There's code to reset the clocksweep in a
separate patch. Sorry for not removing it earlier. It will be removed
in the next set of patches.
>
> 2) Doesn't StrategyPurgeFreeList already do some of this for the case
> of shrinking memory?
>
> 3) Not great adding a bunch of static variables to bufmgr.c. Why do we
> need to make "everything" static global? Isn't it enough to make
> only the "valid" flag global? The rest can stay local, no?
>
> If everything needs to be global for some reason, could we at least
> make it a struct, to group the fields, not just separate random
> variables? And maybe at the top, not half-way throught the file?
>
> 4) Isn't the name BgBufferSyncAdjust misleading? It's not adjusting
> anything, it's just invalidating the info about past runs.
I think there's a bit of refactoring possible here. Setting up the
BgBufferSync state, resetting it when bgwriter_lru_maxpages <= 0 and
then re initializing it when bgwriter_lru_maxpages > 0, and actually
performing the buffer sync is all packed into the same function
BgBufferSync() right now. It makes this function harder to read. I
think these functionalities should be separated into their own
functions and use the appropriate one instead of BgBufferSyncAdjust(),
whose name is misleading. The static global variables should all be
packed into a structure which is passed as an argument to these
functions. I need more time to study the code and refactor it that
way. For now I have added a note to the commit message of this patch
so that I will revisit it. I have renamed BgBufferSyncAdjust() to
BgBufferSyncReset().
>
> 5) I don't quite understand why BufferSync needs to do the dance with
> delay_shmem_resize. I mean, we certainly should not run BufferSync
> from the code that resizes buffers, right? Certainly not after the
> eviction, from the part that actually rebuilds shmem structs etc.
Right. But let me answer all three questions together.
> So perhaps something could trigger resize while we're running the
> BufferSync()? Isn't that a bit strange? If this flag is needed, it
> seems more like a band-aid for some issue in the architecture.
>
> 6) Also, why should it be fine to get into situation that some of the
> buffers might not be valid, during shrinking? I mean, why should
> this check (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers).
> It seems better to ensure we never get into "sync" in a way that
> might lead some of the buffers invalid. Seems way too lowlevel to
> care about whether resize is happening.
>
> 7) I don't understand the new condition for "Execute the LRU scan".
> Won't this stop LRU scan even in cases when we want it to happen?
> Don't we want to scan the buffers in the remaining part (after
> shrinking), for example? Also, we already checked this shmem flag at
> the beginning of the function - sure, it could change (if some other
> process modifies it), but does that make sense? Wouldn't it cause
> problems if it can change at an arbitrary point while running the
> BufferSync? IMHO just another sign it may not make sense to allow
> this, i.e. buffer sync should not run during the "actual" resize.
>
ProcessBarrierShmemResize() which does the resizing is part of
ProcessProcSignalBarrier() which in turn gets called from
CHECK_FOR_INTERRUPTS(), which is called from multiple places, even
from elog(). I am not able to find a call stack linking BgBufferSync()
and ProcessProcSignalBarrier(). But I couldn't convince myself that it
is true and will remain true in the future. I mean, the function loops
through a large number of buffers and performs IO, both avenues to
call CHECK_FOR_INTERRUPTS(). Hence that flag. Do you know what (part
of code) guarantees that ProcessProcSignalBarrier() will never be
called from BgBufferSync()?
Note, resizing can not begin till delay_shmem_resize is cleared, so
while BgBufferSync is executing, no buffer can be invalidated or no
new buffers could be added. But at the cost of all other backends to
wait till BgBufferSync finishes. We want to avoid that. The idea here
is to make BgBufferSync stop as soon as it realises that the buffer
resizing is "about to begin". But I think the condition looks wrong. I
think the right condition would be NBufferPending != NBuffers or
NBuffersOld. AFAIK, Dmitry is working on consolidating NBuffers*
variables as you have requested elsewhere. Better even if we could
somehow set a flag in shared memory indicating that the buffer
resizing is "about to begin" and BgBufferSync() checks that flag. So I
will wait for him to make that change and then change this condition.
>
> v5-0010-Additional-validation-for-buffer-in-the-ring.patch
>
> 1) So the problem is we might create a ring before shrinking shared
> buffers, and then GetBufferFromRing will see bogus buffers? OK, but
> we should be more careful with these checks, otherwise we'll miss
> real issues when we incorrectly get an invalid buffer. Can't the
> backends do this only when they for sure know we did shrink the
> shared buffers? Or maybe even handle that during the barrier?
>
> 2) IMHO a sign there's the "transitions" between different NBuffers
> values may not be clear enough, and we're allowing stuff to happen
> in the "blurry" area. I think that's likely to cause bugs (it did
> cause issues for the online checksums patch, I think).
>
I think you are right, that this might hide some bugs. Just like we
remove buffers to be shrunk from freelist only once, I wanted each
backend to remove them buffer rings only once. But I couldn't find a
way to make all the buffer rings for a given backend available to the
barrier handling code. The rings are stored in Scan objects, which
seem local to the executor nodes. Is there a way to make them
available to barrier handling code (even if it has to walk an
execution tree, let's say)?
If there would have been only one scan, we could have set a flag after
shrinking, let GetBufferFromRing() purge all invalid buffers once when
flag is true and reset the flag. But there can be more than one scan
happening and we don't know how many there are and when all of them
have finished calling GetBufferFromRing() after shrinking. Suggestions
to do this only once are welcome.
I will send the next set of patches with my next email.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-09-18 04:55 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-09-28 09:24 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 2 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-09-18 04:55 UTC (permalink / raw)
To: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Mon, Jun 16, 2025 at 6:09 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> >
> > Buffer lookup table resizing
> > ------------------------------------
I looked at the interaction of shared buffer lookup table with buffer
resizing as per the patches in [0]. Here's list of my findings, issues
and fixes.
1. The basic structure of buffer lookup table (directory and control
area etc.) is allocated in a shared memory segment dedicated to the
buffer lookup table. However, the entries are allocated in the shared
memory using ShmemAllocNoError() which allocates the entries in the
main memory segment. In order for ShmemAllocNoError() to allocate
entries in the dedicated shared memory segment, it should know the
shared memory segment. We could do that by setting the segment number
in element_alloc() before calling hashp->alloc(). This is similar to
how ShmemAllocNoError() knows the memory context in which to allocate
the entries on heap. But read on ...
2. When the buffer pool is expanded, an "out of shared memory" error
is thrown when more entries are added to the buffer look up table. We
could temporarily adjust that flag and allocate more entries. But the
directory also needs to be expanded proportionately otherwise it may
lead to more contention. Expanding directory is non-trivial since it's
a contiguous chunk of memory, followed by other data structures.
Further, expanding directory would require rehashing all the existing
entries, which may impact the time taken by the resizing operation and
how long other backends remain blocked.
3. When the buffer pool is shrunk, there is no way to free the extra
entries in such a way that a contiguous chunk of shared memory can be
given back to the OS. In case we implement it, we will need some way
to compact the shrunk entries in contiguous chunk of memory and unmap
remaining chunk. That's some significant code.
Given these things, I think we should set up the buffer lookup table
to hold maximum entries required to expand the buffer pool to its
maximum, right at the beginning. The maximum size to which buffer pool
can grow is given by GUC max_available_memory (which is a misnomer and
should be renamed to max_shared_buffers or something), introduced by
previous set of patches [0]. We don't shrink or expand the buffer
lookup table as we shrink and expand the buffer pool. With that the
buffer lookup table can be located in the main memory segment itself
and we don't have to fix ShmemAllocNoError().
This has two side effects:
1. larger hash table makes hash table operations slower [2]. Its
impact on actual queries needs to be studied.
2. There's increase in the total shared memory allocated upfront.
Currently we allocate 150MB memory with all default GUC values. With
this change we will allocate 250MB memory since max_available_memory
(or rather max_shared_buffers) defaults to allow 524288 shared
buffers. If we make max_shared_buffers to default to shared_buffers,
it won't be a problem. However, when a user sets max_shared_buffers
themselves, they have to be conscious of the fact that it will
allocate more memory than necessary with given shared_buffers value.
This fix is part of patch 0015.
The patchset contains more fixes and improvements as described below.
Per TODO in the prologue of CalculateShmemSize(), more than necessary
shared memory was mapped and allocated in the buffer manager related
memory segments because of an error in that function; the amount of
memory to be allocated in the main shared memory segment was added to
every other shared memory segment. Thus shrinking those memory
segments didn't actually affect the objects allocated in those.
Because of that, we were not seeing SIGBUS even when the objects
supposedly shrunk were accessed, masking bugs in the patches. In this
patchset I have a working fix for CalculateShmemSize(). With that fix
in place we see server crashing with SIGBUS in some resizing
operations. Those cases need to be investigated. The fix changes its
minions to a. return size of shared memory objects to be allocated in
the main memory segment and b. add sizes of the shared memory objects
to be allocated in other memory segments in the respective
AnonymousMapping structures. This assymetry between main segment and
other segment exists so as not to change a lot the minions of
CalculateShmemSize(). But I think we should eliminate the assymetry
and change every minion to add sizes in the respective segment's
AnonymousMapping structure. The patch proposed at [3] would simplify
CalculateShmemSize() which should help eliminating the assymetry.
Along with refactoring CalculateShmemSize() I have added small fixes
to update the total size and end address of shared memory mapping
after resizing them and also to update the new allocated_sizes of
resized structures in ShmemIndex entry. Patch 0009 includes these
changes.
I found that the shared memory resizing synchronization is triggered
even before setting up the shared buffers the first time after
starting the server. That's not required and also can lead to issues
because of trying to resize shared buffers which do not exist. A WIP
fix is included as patch 0012. A TODO in the patch needs to be
addressed. It should be squashed into an earlier patch 0011 when
appropriate.
While debugging the above mentioned issues, I found it useful to have
an insight into the contents of buffer lookup table. Hence I added a
system view exposing the contents of the buffer lookup table. This is
added as patch 0001 in the attached patchset. I think it's useful to
have this independent of this patchset to investigate inconsistencies
between the contents of shared buffer pool and buffer lookup table.
Again for debugging purposes, I have added a new column "segment" in
pg_shmem_allocations reporting the shared memory segment in which the
given allocation has happened. I have also added another view
pg_shmem_segments to provide information about the shared memory
segments. This view definition will change as we design shared memory
mappings and shared memory segments better. So it's WIP and needs doc
changes as well. I have included it in the patchset as patch 0011
since it will be helpful to debug issues found in the patch when
testing. The patch should be merged into patch 0007.
Last but not the least, patch 0016 contains two tests a. stress test
to run buffer resizing while pgbench is running, b. a SQL test to test
the sizes of segments and shared memory allocations after resizing.
The stress test polls "show shared_buffers" output to know when the
resizing is finished. I think we need a better interface to know when
resizing has finished. Thanks a lot my colleague Palak Chaturvedi for
providing initial draft of the test case.
The patches are rebased on top of the latest master, which includes
changes to remove free buffer list. That led to removing all the code
in these patches dealing with free buffer list.
I am intentionally keeping my changes (patches 0001, 0008 to 0012,
0012 to 0016) separate from Dmitry's changes so that Dmitry can review
them easily. The patches are arranged so that my patches are nearer to
Dmitry's patches, into which, they should be squashed.
Dmitry,
I found that max_available_memory is PGC_SIGHUP. Is that intentional?
I thought it's PGC_POSTMASTER since we can not reserve more address
space without restarting postmaster. Left a TODO for this. I think we
also need to change the name and description to better reflect its
actual functionality.
[0] https://www.postgresql.org/message-id/my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj@hmuxsf2ngov...
[1] https://www.postgresql.org/message-id/CAExHW5v0jh3F_wj86yC%3DqBfWk0uiT94qy%3DZ41uzAHLHh0SerRA%40mail...
[2] https://ashutoshpg.blogspot.com/2025/07/efficiency-of-sparse-hash-table.html
[3] https://commitfest.postgresql.org/patch/5997/
--
Best Wishes,
Ashutosh Bapat
Attachments:
[application/x-patch] 0004-Introduce-pss_barrierReceivedGeneration-20250918.patch (7.3K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/2-0004-Introduce-pss_barrierReceivedGeneration-20250918.patch)
download | inline diff:
From 0a55bc15dc3a724f03e674048109dac1f248c406 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 4 Apr 2025 21:46:14 +0200
Subject: [PATCH 04/16] Introduce pss_barrierReceivedGeneration
Currently WaitForProcSignalBarrier allows to make sure the message sent
via EmitProcSignalBarrier was processed by all ProcSignal mechanism
participants.
Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
which will be updated when a process has received the message, but not
processed it yet. This makes it possible to support a new mode of
waiting, when ProcSignal participants want to synchronize message
processing. To do that, a participant can wait via
WaitForProcSignalBarrierReceived when processing a message, effectively
making sure that all processes are going to start processing
ProcSignalBarrier simultaneously.
---
src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------
src/include/storage/procsignal.h | 1 +
2 files changed, 54 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 087821311cc..eb3ceaae809 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -58,7 +58,10 @@
* of it. For such use cases, we set a bit in pss_barrierCheckMask and then
* increment the current "barrier generation"; when the new barrier generation
* (or greater) appears in the pss_barrierGeneration flag of every process,
- * we know that the message has been received everywhere.
+ * we know that the message has been received and processed everywhere. In case
+ * if we only need to know only that the message was received everywhere (e.g.
+ * receiving processes need to handle the message in a coordinated fashion)
+ * use pss_barrierReceivedGeneration in the same way.
*/
typedef struct
{
@@ -70,6 +73,7 @@ typedef struct
/* Barrier-related fields (not protected by pss_mutex) */
pg_atomic_uint64 pss_barrierGeneration;
+ pg_atomic_uint64 pss_barrierReceivedGeneration;
pg_atomic_uint32 pss_barrierCheckMask;
ConditionVariable pss_barrierCV;
} ProcSignalSlot;
@@ -152,6 +156,8 @@ ProcSignalShmemInit(void)
slot->pss_cancel_key_len = 0;
MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration,
+ PG_UINT64_MAX);
pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
ConditionVariableInit(&slot->pss_barrierCV);
}
@@ -199,6 +205,8 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
barrier_generation =
pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration);
pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration,
+ barrier_generation);
if (cancel_key_len > 0)
memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len);
@@ -263,6 +271,7 @@ CleanupProcSignalState(int status, Datum arg)
* no barrier waits block on it.
*/
pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX);
SpinLockRelease(&slot->pss_mutex);
@@ -416,12 +425,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type)
return generation;
}
-/*
- * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
- * requested by a specific call to EmitProcSignalBarrier() have taken effect.
- */
-void
-WaitForProcSignalBarrier(uint64 generation)
+static void
+WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly)
{
Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration));
@@ -436,12 +441,17 @@ WaitForProcSignalBarrier(uint64 generation)
uint64 oldval;
/*
- * It's important that we check only pss_barrierGeneration here and
- * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared
- * before the barrier is actually absorbed, but pss_barrierGeneration
+ * It's important that we check only pss_barrierGeneration &
+ * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in
+ * pss_barrierCheckMask get cleared before the barrier is actually
+ * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration
* is updated only afterward.
*/
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
while (oldval < generation)
{
if (ConditionVariableTimedSleep(&slot->pss_barrierCV,
@@ -450,7 +460,11 @@ WaitForProcSignalBarrier(uint64 generation)
ereport(LOG,
(errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier",
(int) pg_atomic_read_u32(&slot->pss_pid))));
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
}
ConditionVariableCancelSleep();
}
@@ -464,12 +478,33 @@ WaitForProcSignalBarrier(uint64 generation)
* The caller is probably calling this function because it wants to read
* the shared state or perform further writes to shared state once all
* backends are known to have absorbed the barrier. However, the read of
- * pss_barrierGeneration was performed unlocked; insert a memory barrier
- * to separate it from whatever follows.
+ * pss_barrierGeneration & pss_barrierReceivedGeneration was performed
+ * unlocked; insert a memory barrier to separate it from whatever follows.
*/
pg_memory_barrier();
}
+/*
+ * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
+ * requested by a specific call to EmitProcSignalBarrier() have taken effect.
+ */
+void
+WaitForProcSignalBarrier(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, false);
+}
+
+/*
+ * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all
+ * backends have observed the message sent by a specific call to
+ * EmitProcSignalBarrier().
+ */
+void
+WaitForProcSignalBarrierReceived(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, true);
+}
+
/*
* Handle receipt of an interrupt indicating a global barrier event.
*
@@ -523,6 +558,10 @@ ProcessProcSignalBarrier(void)
if (local_gen == shared_gen)
return;
+ /* The message is observed, record that */
+ pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
+ shared_gen);
+
/*
* Get and clear the flags that are set for this backend. Note that
* pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index afeeb1ca019..2733bbb8c5b 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, const uint8 *cancel_key, int cance
extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type);
extern void WaitForProcSignalBarrier(uint64 generation);
+extern void WaitForProcSignalBarrierReceived(uint64 generation);
extern void ProcessProcSignalBarrier(void);
extern void procsignal_sigusr1_handler(SIGNAL_ARGS);
--
2.34.1
[application/x-patch] 0002-Process-config-reload-in-AIO-workers-20250918.patch (1.8K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/3-0002-Process-config-reload-in-AIO-workers-20250918.patch)
download | inline diff:
From d1ed934ccd02fca2c831e582b07a169e17d19f59 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 15:14:33 +0200
Subject: [PATCH 02/16] Process config reload in AIO workers
Currenly AIO workers process interrupts only via CHECK_FOR_INTERRUPTS,
which does not include ConfigReloadPending. Thus we need to check for it
explicitly.
---
src/backend/storage/aio/method_worker.c | 25 +++++++++++++++++++++++++
1 file changed, 25 insertions(+)
diff --git a/src/backend/storage/aio/method_worker.c b/src/backend/storage/aio/method_worker.c
index b5ac073a910..d1c6da89c4b 100644
--- a/src/backend/storage/aio/method_worker.c
+++ b/src/backend/storage/aio/method_worker.c
@@ -80,6 +80,7 @@ static void pgaio_worker_shmem_init(bool first_time);
static bool pgaio_worker_needs_synchronous_execution(PgAioHandle *ioh);
static int pgaio_worker_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+static void pgaio_worker_process_interrupts(void);
const IoMethodOps pgaio_worker_ops = {
.shmem_size = pgaio_worker_shmem_size,
@@ -463,6 +464,8 @@ IoWorkerMain(const void *startup_data, size_t startup_data_len)
int nwakeups = 0;
int worker;
+ pgaio_worker_process_interrupts();
+
/*
* Try to get a job to do.
*
@@ -592,3 +595,25 @@ pgaio_workers_enabled(void)
{
return io_method == IOMETHOD_WORKER;
}
+
+/*
+ * Process any new interrupts.
+ */
+static void
+pgaio_worker_process_interrupts(void)
+{
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
+ if (ConfigReloadPending)
+ {
+ ConfigReloadPending = false;
+ ProcessConfigFile(PGC_SIGHUP);
+ }
+
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+}
--
2.34.1
[application/x-patch] 0003-Introduce-pending-flag-for-GUC-assign-hooks-20250918.patch (12.7K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/4-0003-Introduce-pending-flag-for-GUC-assign-hooks-20250918.patch)
download | inline diff:
From 0a13e56dceea8cc7a2685df7ee8cea434588681b Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:40:32 +0200
Subject: [PATCH 03/16] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 61 insertions(+), 40 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 0baf0ac6160..307ac31b19e 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2197,7 +2197,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 608f10d9412..e40dae2ddf2 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,12 +1161,12 @@ assign_maintenance_io_concurrency(int newval, void *extra)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra)
+assign_io_max_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(newval, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra)
+assign_io_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(io_max_combine_limit, newval);
}
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index 25f739a6a17..1726a7c0993 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1951,7 +1951,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1984,7 +1984,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2007,7 +2007,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2030,7 +2030,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index d356830f756..8d4d6cc3f33 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3596,7 +3596,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 46fdefebe35..0d5e523aaf0 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1680,6 +1680,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1688,9 +1689,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2046,13 +2051,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2429,16 +2439,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3855,18 +3870,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f21ec37da89..c3056cd2da8 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 82ac8646a8d..658c799419e 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,12 +81,12 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
-extern void assign_io_max_combine_limit(int newval, void *extra);
-extern void assign_io_combine_limit(int newval, void *extra);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
+extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -141,13 +141,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -163,7 +163,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.34.1
[application/x-patch] 0005-Allow-to-use-multiple-shared-memory-mapping-20250918.patch (31.3K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/5-0005-Allow-to-use-multiple-shared-memory-mapping-20250918.patch)
download | inline diff:
From 63fe27340656c52b13f4eecebd9e73d24efe5e33 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH 05/16] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 +++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 148 +++++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 15 ++-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 11 +++
12 files changed, 283 insertions(+), 128 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 6ac83ea1a82..7bb363989c4 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -327,7 +327,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -348,7 +348,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..56af0231d24 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 2704e80b3a7..1965b2d3eb4 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2fa045e6b0f..8b38e985327 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -363,7 +368,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index a0770e86796..f185ed28f95 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -76,19 +76,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -101,9 +101,17 @@ Datum pg_numa_available(PG_FUNCTION_ARGS);
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -114,7 +122,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -123,9 +137,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -150,11 +164,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -184,6 +204,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -203,22 +229,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -236,6 +262,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -246,19 +278,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -273,7 +305,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -334,6 +372,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
int64 max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -350,9 +400,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -385,6 +435,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -393,7 +450,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -416,7 +473,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -433,8 +490,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -459,7 +516,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -478,14 +535,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -542,10 +598,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -557,15 +614,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
@@ -630,7 +687,12 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ PGShmemHeader *shmhdr = Segments[segment].ShmemSegHdr;
+ shm_total_page_count += (shmhdr->totalsize / os_page_size) + 1;
+ }
+
page_ptrs = palloc0(sizeof(void *) * shm_total_page_count);
pages_status = palloc(sizeof(int) * shm_total_page_count);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 46c82c63ca5..93792a83af9 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -594,12 +596,15 @@ LWLockNewTrancheId(const char *name)
/*
* We use the ShmemLock spinlock to protect LWLockCounter and
* LWLockTrancheNames.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
*/
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
if (*LWLockCounter - LWTRANCHE_FIRST_USER_DEFINED >= MAX_NAMED_TRANCHES)
{
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
ereport(ERROR,
(errmsg("maximum number of tranches already registered"),
errdetail("No more than %d tranches may be registered.",
@@ -610,7 +615,7 @@ LWLockNewTrancheId(const char *name)
LocalLWLockCounter = *LWLockCounter;
strlcpy(LWLockTrancheNames[result - LWTRANCHE_FIRST_USER_DEFINED], name, NAMEDATALEN);
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
@@ -732,9 +737,9 @@ GetLWTrancheName(uint16 trancheId)
*/
if (trancheId >= LocalLWLockCounter)
{
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
LocalLWLockCounter = *LWLockCounter;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
if (trancheId >= LocalLWLockCounter)
elog(ERROR, "tranche %d is not registered", trancheId);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..6ebda479ced 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 5f7d4b83a60..2348c59b5a0 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -91,4 +106,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index cd683a9d2d9..910c43f54f4 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -30,15 +30,26 @@ extern PGDLLIMPORT slock_t *ShmemLock;
typedef struct PGShmemHeader PGShmemHeader; /* avoid including
* storage/pg_shmem.h here */
extern void InitShmemAccess(PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
--
2.34.1
[application/x-patch] 0001-Add-system-view-for-shared-buffer-lookup-ta-20250918.patch (9.6K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/6-0001-Add-system-view-for-shared-buffer-lookup-ta-20250918.patch)
download | inline diff:
From cc90e3f74fe4a14ba95e11664ac68b8daa3ba056 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 25 Aug 2025 19:23:50 +0530
Subject: [PATCH 01/16] Add system view for shared buffer lookup table
The view exposes the contents of the shared buffer lookup table for
debugging, testing and investigation.
TODO:
It is better to place this view in pg_buffercache. But it's added as a
system view since BufHashTable is not exposed outside buf_table.c. To
move it to pg_buffercache, we should move the function
pg_get_buffer_lookup_table() to pg_buffercache which invokes
BufTableGetContent() by passing it the tuple store and tuple descriptor.
BufTableGetContent fills the tuple store. The partitions are locked by
pg_get_buffer_lookup_table().
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
doc/src/sgml/system-views.sgml | 89 ++++++++++++++++++++++++++
src/backend/catalog/system_views.sql | 7 ++
src/backend/storage/buffer/buf_table.c | 61 ++++++++++++++++++
src/include/catalog/pg_proc.dat | 11 ++++
src/test/regress/expected/rules.out | 7 ++
5 files changed, 175 insertions(+)
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 4187191ea74..89be9bc333f 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -71,6 +71,11 @@
<entry>backend memory contexts</entry>
</row>
+ <row>
+ <entry><link linkend="view-pg-buffer-lookup-table"><structname>pg_buffer_lookup_table</structname></link></entry>
+ <entry>shared buffer lookup table</entry>
+ </row>
+
<row>
<entry><link linkend="view-pg-config"><structname>pg_config</structname></link></entry>
<entry>compile-time configuration parameters</entry>
@@ -896,6 +901,90 @@ AND c1.path[c2.level] = c2.path[c2.level];
</para>
</sect1>
+ <sect1 id="view-pg-buffer-lookup-table">
+ <title><structname>pg_buffer_lookup_table</structname></title>
+ <indexterm>
+ <primary>pg_buffer_lookup_table</primary>
+ </indexterm>
+ <para>
+ The <structname>pg_buffer_lookup_table</structname> view exposes the current
+ contents of the shared buffer lookup table. Each row represents an entry in
+ the lookup table mapping a relation page to the ID of buffer in which it is
+ cached. The shared buffer lookup table is locked for a short duration while
+ reading so as to ensure consistency. This may affect performance if this view
+ is queried very frequently.
+ </para>
+ <table id="pg-buffer-lookup-table-view" xreflabel="pg_buffer_lookup_table">
+ <title><structname>pg_buffer_lookup_table</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>tablespace</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the tablespace containing the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>database</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database containing the relation (zero for shared relations)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relfilenode</structfield> <type>oid</type>
+ </para>
+ <para>
+ relfilenode identifying the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>forknum</structfield> <type>int2</type>
+ </para>
+ <para>
+ Fork number within the relation (see <xref linkend="storage-file-layout"/>)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>blocknum</structfield> <type>int8</type>
+ </para>
+ <para>
+ Block number within the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>bufferid</structfield> <type>int4</type>
+ </para>
+ <para>
+ ID of the buffer caching the page
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ <para>
+ Access to this view is restricted to members of the
+ <literal>pg_read_all_stats</literal> role by default.
+ </para>
+ </sect1>
+
<sect1 id="view-pg-config">
<title><structname>pg_config</structname></title>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index c77fa0234bb..46fc28396de 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1420,3 +1420,10 @@ REVOKE ALL ON pg_aios FROM PUBLIC;
GRANT SELECT ON pg_aios TO pg_read_all_stats;
REVOKE EXECUTE ON FUNCTION pg_get_aios() FROM PUBLIC;
GRANT EXECUTE ON FUNCTION pg_get_aios() TO pg_read_all_stats;
+
+CREATE VIEW pg_buffer_lookup_table AS
+ SELECT * FROM pg_get_buffer_lookup_table();
+REVOKE ALL ON pg_buffer_lookup_table FROM PUBLIC;
+GRANT SELECT ON pg_buffer_lookup_table TO pg_read_all_stats;
+REVOKE EXECUTE ON FUNCTION pg_get_buffer_lookup_table() FROM PUBLIC;
+GRANT EXECUTE ON FUNCTION pg_get_buffer_lookup_table() TO pg_read_all_stats;
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 9d256559bab..1f6e215a2ca 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -21,7 +21,12 @@
*/
#include "postgres.h"
+#include "fmgr.h"
+#include "funcapi.h"
#include "storage/buf_internals.h"
+#include "storage/lwlock.h"
+#include "utils/rel.h"
+#include "utils/builtins.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -159,3 +164,59 @@ BufTableDelete(BufferTag *tagPtr, uint32 hashcode)
if (!result) /* shouldn't happen */
elog(ERROR, "shared buffer hash table corrupted");
}
+
+/*
+ * SQL callable function to report contents of the shared buffer lookup table.
+ */
+Datum
+pg_get_buffer_lookup_table(PG_FUNCTION_ARGS)
+{
+#define PG_GET_BUFFER_LOOKUP_TABLE_COLS 6
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ HASH_SEQ_STATUS hstat;
+ BufferLookupEnt *ent;
+ Datum values[PG_GET_BUFFER_LOOKUP_TABLE_COLS];
+ bool nulls[PG_GET_BUFFER_LOOKUP_TABLE_COLS];
+ int i;
+
+ memset(nulls, 0, sizeof(nulls));
+
+ /*
+ * We put all the tuples into a tuplestore in one scan of the hashtable.
+ * This avoids any issue of the hashtable possibly changing between calls.
+ */
+ InitMaterializedSRF(fcinfo, 0);
+
+ Assert(rsinfo->setDesc->natts == PG_GET_BUFFER_LOOKUP_TABLE_COLS);
+
+ /*
+ * Lock all buffer mapping partitions to ensure a consistent view of the
+ * hash table during the scan. Must grab LWLocks in partition-number order
+ * to avoid LWLock deadlock.
+ */
+ for (i = 0; i < NUM_BUFFER_PARTITIONS; i++)
+ LWLockAcquire(BufMappingPartitionLockByIndex(i), LW_SHARED);
+
+ hash_seq_init(&hstat, SharedBufHash);
+ while ((ent = (BufferLookupEnt *) hash_seq_search(&hstat)) != NULL)
+ {
+ values[0] = ObjectIdGetDatum(ent->key.spcOid);
+ values[1] = ObjectIdGetDatum(ent->key.dbOid);
+ values[2] = ObjectIdGetDatum(ent->key.relNumber);
+ values[3] = ObjectIdGetDatum(ent->key.forkNum);
+ values[4] = UInt32GetDatum(ent->key.blockNum);
+ values[5] = Int32GetDatum(ent->id);
+
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
+
+ /*
+ * Release all buffer mapping partition locks in the reverse order so as
+ * to avoid LWLock deadlock.
+ */
+ for (i = NUM_BUFFER_PARTITIONS - 1; i >= 0; i--)
+ LWLockRelease(BufMappingPartitionLockByIndex(i));
+
+ return (Datum) 0;
+}
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 03e82d28c87..1e53b7a4ae5 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8592,6 +8592,17 @@
proargmodes => '{o,o,o}', proargnames => '{name,type,size}',
prosrc => 'pg_get_dsm_registry_allocations' },
+# buffer lookup table
+{ oid => '5102',
+ descr => 'shared buffer lookup table',
+ proname => 'pg_get_buffer_lookup_table', prorows => '6', proretset => 't',
+ provolatile => 'v', prorettype => 'record',
+ proargtypes => '', proallargtypes => '{oid,oid,oid,int2,int8,int4}',
+ proargmodes => '{o,o,o,o,o,o}',
+ proargnames => '{tablespace,database,relfilenode,forknum,blocknum,bufferid}',
+ prosrc => 'pg_get_buffer_lookup_table'
+},
+
# memory context of local backend
{ oid => '2282',
descr => 'information about all memory contexts of local backend',
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 35e8aad7701..760bb13fe95 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1330,6 +1330,13 @@ pg_backend_memory_contexts| SELECT name,
free_chunks,
used_bytes
FROM pg_get_backend_memory_contexts() pg_get_backend_memory_contexts(name, ident, type, level, path, total_bytes, total_nblocks, free_bytes, free_chunks, used_bytes);
+pg_buffer_lookup_table| SELECT tablespace,
+ database,
+ relfilenode,
+ forknum,
+ blocknum,
+ bufferid
+ FROM pg_get_buffer_lookup_table() pg_get_buffer_lookup_table(tablespace, database, relfilenode, forknum, blocknum, bufferid);
pg_config| SELECT name,
setting
FROM pg_config() pg_config(name, setting);
base-commit: 2e66cae935c2e0f7ce9bab6b65ddeb7806f4de7c
--
2.34.1
[application/x-patch] 0006-Address-space-reservation-for-shared-memory-20250918.patch (24.4K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/7-0006-Address-space-reservation-for-shared-memory-20250918.patch)
download | inline diff:
From e2f48da8a8206711b24e34040d699431910fbf9c Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 11:47:04 +0200
Subject: [PATCH 06/16] Address space reservation for shared memory
Currently the shared memory layout is designed to pack everything tight
together, leaving no space between mappings for resizing. Here is how it
looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
...
Make the layout more dynamic via splitting every shared memory segment
into two parts:
* An anonymous file, which actually contains shared memory content. Such
an anonymous file is created via memfd_create, it lives in memory,
behaves like a regular file and semantically equivalent to an
anonymous memory allocated via mmap with MAP_ANONYMOUS.
* A reservation mapping, which size is much larger than required shared
segment size. This mapping is created with flags PROT_NONE (which
makes sure the reserved space is not used), and MAP_NORESERVE (to not
count the reserved space against memory limits). The anonymous file is
mapped into this reservation mapping.
The resulting layout looks like this:
00400000-00490000 /path/bin/postgres
...
3f526000-3f590000 rw-p [heap]
7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
To resize a shared memory segment in this layout it's possible to use ftruncate
on the anonymous file, adjusting access permissions on the reserved space as
needed.
This approach also do not impact the actual memory usage as reported by
the kernel. Here is the output of /proc/$PID/status for the master
version with shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages
// mapped in mm_struct. It corresponds to the mapped reserved space
// and is the only number that grows with it.
VmPeak: 2043192 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 22908 kB
// Size of resident anonymous memory
RssAnon: 768 kB
// Size of resident file mappings
RssFile: 10364 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 11776 kB
Here is the same for the patch when reserving 20GB of space:
VmPeak: 21255824 kB
VmRSS: 25020 kB
RssAnon: 768 kB
RssFile: 10812 kB
RssShmem: 13440 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
20770816 (~19.8 MB)
To control the amount of space reserved a new GUC max_available_memory
is introduced. Ideally it should be based on the maximum available
memory, hense the name.
There are also few unrelated advantages of using anon files:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps.
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 290 ++++++++++++++++++----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 5 +-
src/backend/storage/ipc/shmem.c | 2 +-
src/backend/utils/init/globals.c | 1 +
src/backend/utils/misc/guc_parameters.dat | 12 +
src/include/miscadmin.h | 1 +
src/include/portability/mem.h | 2 +-
src/include/storage/pg_shmem.h | 5 +-
9 files changed, 260 insertions(+), 60 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 56af0231d24..363ddfd1fca 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -97,10 +97,12 @@ void *UsedShmemSegAddr = NULL;
typedef struct AnonymousMapping
{
int shmem_segment;
- Size shmem_size; /* Size of the mapping */
+ Size shmem_size; /* Size of the actually used memory */
+ Size shmem_reserved; /* Size of the reserved mapping */
Pointer shmem; /* Pointer to the start of the mapped memory */
Pointer seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -108,6 +110,49 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping layout we use looks like this:
+ *
+ * 00400000-00c2a000 r-xp /bin/postgres
+ * ...
+ * 3f526000-3f590000 rw-p [heap]
+ * 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted)
+ * 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted)
+ * 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
+ * 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
+ * ...
+ *
+ * We need to place shared memory mappings in such a way, that there will be
+ * gaps between them in the address space. Those gaps have to be large enough
+ * to resize the mapping up to certain size, without counting towards the total
+ * memory consumption.
+ *
+ * To achieve this, for each shared memory segment we first create an anonymous
+ * file of specified size using memfd_create, which will accomodate actual
+ * shared memory mapping content. It is represented by the first /memfd:main
+ * with rw permissions. Then we create a mapping for this file using mmap, with
+ * size much larger than required and flags PROT_NONE (allows to make sure the
+ * reserved space will not be used) and MAP_NORESERVE (prevents the space from
+ * being counted against memory limits). The mapping serves as an address space
+ * reservation, into which shared memory segment can be extended and is
+ * represented by the second /memfd:main with no permissions.
+ *
+ * The reserved space for each segment is calculated as a fraction of the total
+ * reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
+ * array.
+ */
+static double SHMEM_RESIZE_RATIO[1] = {
+ 1.0, /* MAIN_SHMEM_SLOT */
+};
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -503,19 +548,20 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* hugepage sizes, we might want to think about more invasive strategies,
* such as increasing shared_buffers to absorb the extra space.
*
- * Returns the (real, assumed or config provided) page size into
- * *hugepagesize, and the hugepage-related mmap flags to use into
- * *mmap_flags if requested by the caller. If huge pages are not supported,
- * *hugepagesize and *mmap_flags are set to 0.
+ * Returns the (real, assumed or config provided) page size into *hugepagesize,
+ * the hugepage-related mmap and memfd flags to use into *mmap_flags and
+ * *memfd_flags if requested by the caller. If huge pages are not supported,
+ * *hugepagesize, *mmap_flags and *memfd_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -574,6 +620,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -584,7 +631,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -593,6 +649,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -600,6 +658,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -625,72 +685,90 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
* Creates an anonymous mmap()ed shared memory segment.
*
* This function will modify mapping size to the actual size of the allocation,
- * if it ends up allocating a segment that is larger than requested.
+ * if it ends up allocating a segment that is larger than requested. If needed,
+ * it also rounds up the mapping reserved size to be a multiple of huge page
+ * size.
+ *
+ * Note that we do not fallback from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
*/
static void
CreateAnonymousSegment(AnonymousMapping *mapping)
{
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
- int mmap_errno = 0;
+ int save_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
+
+ elog(DEBUG1, "segment[%s]: size %zu, reserved %zu",
+ MappingName(mapping->shmem_segment), mapping->shmem_size,
+ mapping->shmem_reserved);
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the request size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- {
- DebugMappings();
- elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- MappingName(mapping->shmem_segment), allocsize);
- }
+ /*
+ * The reserved space is multiple of BLCKSZ. We know the huge page
+ * size, round up the reserved space to it.
+ */
+ mapping->shmem_reserved = mapping->shmem_reserved + hugepagesize -
+ (mapping->shmem_reserved % hugepagesize);
+
+ /* Verify that the new size is withing the reserved boundaries */
+ if (mapping->shmem_reserved < mapping->shmem_size)
+ ereport(ERROR,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved"),
+ errhint("You may need to increase \"max_available_memory\".")));
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
}
#endif
/*
- * Report whether huge pages are in use. This needs to be tracked before
- * the second mmap() call if attempting to use huge pages failed
- * previously.
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
*/
- SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
- PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment),
+ memfd_flags);
- if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified value.
+ */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
{
- /*
- * Use the original size, not the rounded-up value, when falling back
- * to non-huge pages.
- */
- allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
- }
+ save_errno = errno;
- if (ptr == MAP_FAILED)
- {
- errno = mmap_errno;
DebugMappings();
+ close(mapping->segment_fd);
+
+ errno = save_errno;
ereport(FATAL,
- (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ (errmsg("segment[%s]: could not truncate anonymous file: %m",
MappingName(mapping->shmem_segment)),
- (mmap_errno == ENOMEM) ?
+ (save_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
"swap space, or huge pages. To reduce the request size "
@@ -700,10 +778,112 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
allocsize) : 0));
}
+ elog(DEBUG1, "segment[%s]: mmap(%zu)",
+ MappingName(mapping->shmem_segment), allocsize);
+
+ /*
+ * Create a reservation mapping.
+ */
+ ptr = mmap(NULL, mapping->shmem_reserved, PROT_NONE,
+ mmap_flags | MAP_NORESERVE, mapping->segment_fd, 0);
+ save_errno = errno;
+
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+
+ errno = save_errno;
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment))));
+ }
+
+ /* Make the memory accessible */
+ if(mprotect(ptr, allocsize, PROT_READ | PROT_WRITE) == -1)
+ {
+ save_errno = errno;
+ DebugMappings();
+
+ errno = save_errno;
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not mprotect anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment))));
+ }
+
mapping->shmem = ptr;
mapping->shmem_size = allocsize;
}
+/*
+ * PrepareHugePages
+ *
+ * Figure out if there are enough huge pages to allocate all shared memory
+ * segments, and report that information via huge_pages_status and
+ * huge_pages_on. It needs to be called before creating shared memory segments.
+ *
+ * It is necessary to maintain the same semantic (simple on/off) for
+ * huge_pages_status, even if there are multiple shared memory segments: all
+ * segments either use huge pages or not, there is no mix of segments with
+ * different page size. The latter might be actually beneficial, in particular
+ * because only some segments may require large amount of memory, but for now
+ * we go with a simple solution.
+ */
+void
+PrepareHugePages()
+{
+ void *ptr = MAP_FAILED;
+
+ /* Reset to handle reinitialization */
+ next_free_segment = 0;
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ */
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ int numSemas;
+ Size segment_size = CalculateShmemSize(&numSemas, segment);
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
+ }
+#endif
+
+ /*
+ * Report whether huge pages are in use. This needs to be tracked before
+ * creating shared memory segments.
+ */
+ SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
+}
+
/*
* AnonymousShmemDetach --- detach from an anonymous mmap'd block
* (called as an on_shmem_exit callback, hence funny argument list)
@@ -746,7 +926,7 @@ PGSharedMemoryCreate(Size size,
void *memAddress;
PGShmemHeader *hdr;
struct stat statbuf;
- Size sysvsize;
+ Size sysvsize, total_reserved;
AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
@@ -760,14 +940,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -776,8 +948,16 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+
+ /* Prepare the mapping information */
mapping->shmem_size = size;
mapping->shmem_segment = next_free_segment;
+ total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+ mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+
+ /* Round up to be a multiple of BLCKSZ */
+ mapping->shmem_reserved = mapping->shmem_reserved + BLCKSZ -
+ (mapping->shmem_reserved % BLCKSZ);
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..732fedee87e 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8b38e985327..b60f7ef9ce2 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -206,6 +206,9 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
+ /* Decide if we use huge pages or regular size pages */
+ PrepareHugePages();
+
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
/* Compute the size of the shared-memory block */
@@ -377,7 +380,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index f185ed28f95..9bb73f31052 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -815,7 +815,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index d31cb45a058..90d3feb547c 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffers = 16384;
+int MaxAvailableMemory = 524288;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index 6bc6be13d2a..c94f3fc3c80 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -1107,6 +1107,18 @@
max => 'INT_MAX / 2',
},
+# TODO: should this be PGC_POSTMASTER?
+{ name => "max_available_memory", type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
+ short_desc => 'Sets the upper limit for the shared_buffers value.',
+ long_desc => 'Shared memory could be resized at runtime, this parameters sets the upper limit for it, beyond which resizing would not be supported. Normally this value would be the same as the total available memory.',
+ flags => 'GUC_UNIT_BLOCKS',
+ variable => 'MaxAvailableMemory',
+ boot_val => '524288',
+ min => '16',
+ max => 'INT_MAX / 2',
+},
+
+
{ name => 'vacuum_buffer_usage_limit', type => 'int', context => 'PGC_USERSET', group => 'RESOURCES_MEM',
short_desc => 'Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum.',
flags => 'GUC_UNIT_KB',
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 1bef98471c3..a0c37a7749e 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int MaxAvailableMemory;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 2348c59b5a0..79b0b1ef9eb 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -61,6 +61,7 @@ extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT int MaxAvailableMemory;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -104,7 +105,9 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+void PrepareHugePages(void);
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
--
2.34.1
[application/x-patch] 0008-Fix-compilation-failures-from-previous-comm-20250918.patch (1.4K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/8-0008-Fix-compilation-failures-from-previous-comm-20250918.patch)
download | inline diff:
From 86ada56e8c48d8111b40d10cae8c96a3286d210a Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 20 Aug 2025 11:35:20 +0530
Subject: [PATCH 08/16] Fix compilation failures from previous commits
shm_total_page_count is used unitialized. If this variable has a random
value to start with, the final sum would be wrong.
Also include pg_shmem.h where shared memory segment macros are used.
Author: Ashutosh Bapat
---
src/backend/storage/buffer/buf_init.c | 1 +
src/backend/storage/ipc/shmem.c | 2 +-
2 files changed, 2 insertions(+), 1 deletion(-)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 5383442e213..6d703e18f8b 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -16,6 +16,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
+#include "storage/pg_shmem.h"
#include "storage/bufmgr.h"
BufferDescPadded *BufferDescriptors;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9bb73f31052..e6cb919f0fc 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -649,7 +649,7 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size os_page_size;
void **page_ptrs;
int *pages_status;
- uint64 shm_total_page_count,
+ uint64 shm_total_page_count = 0,
shm_ent_page_count,
max_nodes;
Size *nodes;
--
2.34.1
[application/x-patch] 0007-Introduce-multiple-shmem-segments-for-share-20250918.patch (11.7K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/9-0007-Introduce-multiple-shmem-segments-for-share-20250918.patch)
download | inline diff:
From 942b69a0876b0e83303e6704da54c4c002a5a2d8 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 11:22:02 +0200
Subject: [PATCH 07/16] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 24 +++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 105 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 363ddfd1fca..dac011b766b 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -139,10 +139,18 @@ static int next_free_segment = 0;
*
* The reserved space for each segment is calculated as a fraction of the total
* reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
- * array.
+ * array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take up to 60% of the whole
+ * space when resizing, based on the fact that it most likely will be the main
+ * consumer of this memory. Those numbers are pulled out of thin air for now,
+ * makes sense to evaluate them more precise.
*/
-static double SHMEM_RESIZE_RATIO[1] = {
- 1.0, /* MAIN_SHMEM_SLOT */
+static double SHMEM_RESIZE_RATIO[6] = {
+ 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.6, /* BUFFERS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -167,6 +175,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6fd3a6bbac5..5383442e213 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +77,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -147,33 +151,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 1f6e215a2ca..18a78967138 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -25,6 +25,7 @@
#include "funcapi.h"
#include "storage/buf_internals.h"
#include "storage/lwlock.h"
+#include "storage/pg_shmem.h"
#include "utils/rel.h"
#include "utils/builtins.h"
@@ -64,10 +65,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 7d59a92bd1a..0bfbbb096d6 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -381,9 +382,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index b60f7ef9ce2..2dbd81afc87 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 47360a3d3d8..f8d34513c7f 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -318,7 +318,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 79b0b1ef9eb..a7b275b4db9 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -109,7 +109,29 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.34.1
[application/x-patch] 0009-Refactor-CalculateShmemSize-20250918.patch (22.4K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/10-0009-Refactor-CalculateShmemSize-20250918.patch)
download | inline diff:
From 2d96fa0ede7381c573ad608d89a90f2a960ceca3 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 21 Aug 2025 11:56:09 +0530
Subject: [PATCH 09/16] Refactor CalculateShmemSize()
This function calls many functions which return the amount of shared
memory required for different shared memory data structures. Up until
now, the returned total of these sizes was used to create a single
shared memory segment. But starting the previous patch, we create
multiple shared memory segments each of which contain one shared memory
structure related to shared buffers and one main memory segment
containing rest of the structures. Since CalculateShmemSize() is called
for every shared memory segment, and its return value is added to the
memory required for all the shared memory segments, we end up allocating
more memory than required.
Instead, CalculateShmemSize() is called only once. Each of its callees
are expected to a. return the size required from the main segment b. add
sizes to the AnonymousMappings corresponding to the other memory
segments.
For individual modules to add memory to their respective
AnonymousMappings, we need to know the different mappings upfront. Hence
ANON_MAPPINGS replaces next_free_segment.
TODOs:
1. This change however requires that the AnonymousMappings array and
macros defining identifiers of each of the segments be
platform-independent. This patch doesn't achieve that goal for all the
platforms for example windows. We need to fix that.
2. If postgres is invoked with -C shared_memory_size, it reports 0.
That's because it report the GUC values before share memory sizes are
set in AnonymousMappings. Fix that too.
3. Eliminate this assymetry in CalculateShmemSize(). See TODO in
prologue of CalculateShmemSize().
4. This is one way to avoid requesting more memory in each segment. But
there may be other ways to design CalculateShmemSize(). Need to think
and implement it better.
Author: Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 48 ++++++--------------
src/backend/port/win32_shmem.c | 7 +--
src/backend/postmaster/postmaster.c | 14 +++---
src/backend/storage/buffer/buf_init.c | 55 ++++++++---------------
src/backend/storage/ipc/ipci.c | 65 ++++++++++++++++++++++-----
src/backend/storage/ipc/shmem.c | 8 ++--
src/backend/tcop/postgres.c | 14 +++---
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_shmem.h | 17 ++++++-
10 files changed, 125 insertions(+), 107 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index dac011b766b..b85911bdfc4 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,21 +94,7 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-typedef struct AnonymousMapping
-{
- int shmem_segment;
- Size shmem_size; /* Size of the actually used memory */
- Size shmem_reserved; /* Size of the reserved mapping */
- Pointer shmem; /* Pointer to the start of the mapped memory */
- Pointer seg_addr; /* SysV shared memory for the header */
- unsigned long seg_id; /* IPC key */
- int segment_fd; /* fd for the backing anon file */
-} AnonymousMapping;
-
-static AnonymousMapping Mappings[ANON_MAPPINGS];
-
-/* Keeps track of used mapping segments */
-static int next_free_segment = 0;
+AnonymousMapping Mappings[ANON_MAPPINGS];
/*
* Anonymous mapping layout we use looks like this:
@@ -168,7 +154,7 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
-static const char*
+const char*
MappingName(int shmem_segment)
{
switch (shmem_segment)
@@ -193,7 +179,7 @@ MappingName(int shmem_segment)
static void
DebugMappings()
{
- for(int i = 0; i < next_free_segment; i++)
+ for(int i = 0; i < ANON_MAPPINGS; i++)
{
AnonymousMapping m = Mappings[i];
elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
@@ -851,9 +837,6 @@ PrepareHugePages()
{
void *ptr = MAP_FAILED;
- /* Reset to handle reinitialization */
- next_free_segment = 0;
-
/* Complain if hugepages demanded but we can't possibly support them */
#if !defined(MAP_HUGETLB)
if (huge_pages == HUGE_PAGES_ON)
@@ -876,8 +859,7 @@ PrepareHugePages()
*/
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
- int numSemas;
- Size segment_size = CalculateShmemSize(&numSemas, segment);
+ Size segment_size = Mappings[segment].shmem_req_size;
if (segment_size % hugepagesize != 0)
segment_size += hugepagesize - (segment_size % hugepagesize);
@@ -909,7 +891,7 @@ PrepareHugePages()
static void
AnonymousShmemDetach(int status, Datum arg)
{
- for(int i = 0; i < next_free_segment; i++)
+ for(int i = 0; i < ANON_MAPPINGS; i++)
{
AnonymousMapping m = Mappings[i];
@@ -927,7 +909,7 @@ AnonymousShmemDetach(int status, Datum arg)
/*
* PGSharedMemoryCreate
*
- * Create a shared memory segment of the given size and initialize its
+ * Create a shared memory segment for the given mapping and initialize its
* standard header. Also, register an on_shmem_exit callback to release
* the storage.
*
@@ -937,7 +919,7 @@ AnonymousShmemDetach(int status, Datum arg)
* postmaster or backend.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(AnonymousMapping *mapping,
PGShmemHeader **shim)
{
IpcMemoryKey NextShmemSegID;
@@ -945,7 +927,6 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize, total_reserved;
- AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -965,13 +946,12 @@ PGSharedMemoryCreate(Size size,
errmsg("huge pages not supported with the current \"shared_memory_type\" setting")));
/* Room for a header? */
- Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ Assert(mapping->shmem_req_size > MAXALIGN(sizeof(PGShmemHeader)));
/* Prepare the mapping information */
- mapping->shmem_size = size;
- mapping->shmem_segment = next_free_segment;
+ mapping->shmem_size = mapping->shmem_req_size;
total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
- mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+ mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[mapping->shmem_segment];
/* Round up to be a multiple of BLCKSZ */
mapping->shmem_reserved = mapping->shmem_reserved + BLCKSZ -
@@ -982,8 +962,6 @@ PGSharedMemoryCreate(Size size,
/* On success, mapping data will be modified. */
CreateAnonymousSegment(mapping);
- next_free_segment++;
-
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -992,7 +970,7 @@ PGSharedMemoryCreate(Size size,
}
else
{
- sysvsize = size;
+ sysvsize = mapping->shmem_req_size;
/* huge pages are only available with mmap */
SetConfigOption("huge_pages_status", "off",
@@ -1005,7 +983,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino + next_free_segment;
+ NextShmemSegID = statbuf.st_ino + mapping->shmem_segment;
for (;;)
{
@@ -1214,7 +1192,7 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- for(int i = 0; i < next_free_segment; i++)
+ for(int i = 0; i < ANON_MAPPINGS; i++)
{
AnonymousMapping m = Mappings[i];
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 732fedee87e..1db07ff65d3 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -204,7 +204,7 @@ EnableLockPagesPrivilege(int elevel)
* standard header.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(AnonymousMapping *mapping,
PGShmemHeader **shim)
{
void *memAddress;
@@ -216,7 +216,7 @@ PGSharedMemoryCreate(Size size,
DWORD size_high;
DWORD size_low;
SIZE_T largePageSize = 0;
- Size orig_size = size;
+ Size size = mapping->shmem_req_size;
DWORD flProtect = PAGE_READWRITE;
DWORD desiredAccess;
@@ -304,7 +304,7 @@ retry:
* Use the original size, not the rounded-up value, when
* falling back to non-huge pages.
*/
- size = orig_size;
+ size = mapping->shmem_req_size;
flProtect = PAGE_READWRITE;
goto retry;
}
@@ -391,6 +391,7 @@ retry:
hdr->totalsize = size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
hdr->dsm_control = 0;
+ mapping->shmem_size = size;
/* Save info for possible future use */
UsedShmemSegAddr = memAddress;
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index e1d643b013d..b59d20b4ac2 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -963,13 +963,6 @@ PostmasterMain(int argc, char *argv[])
*/
process_shmem_requests();
- /*
- * Now that loadable modules have had their chance to request additional
- * shared memory, determine the value of any runtime-computed GUCs that
- * depend on the amount of shared memory required.
- */
- InitializeShmemGUCs();
-
/*
* Now that modules have been loaded, we can process any custom resource
* managers specified in the wal_consistency_checking GUC.
@@ -1005,6 +998,13 @@ PostmasterMain(int argc, char *argv[])
*/
CreateSharedMemoryAndSemaphores();
+ /*
+ * Now that loadable modules have had their chance to request additional
+ * shared memory, determine the value of any runtime-computed GUCs that
+ * depend on the amount of shared memory required.
+ */
+ InitializeShmemGUCs();
+
/*
* Estimate number of openable files. This must happen after setting up
* semaphores, because on some platforms semaphores count as open files.
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6d703e18f8b..6f148d1d80b 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -158,48 +158,31 @@ BufferManagerShmemInit(void)
* data.
*/
Size
-BufferManagerShmemSize(int shmem_segment)
+BufferManagerShmemSize(void)
{
- Size size = 0;
+ size_t size;
- if (shmem_segment == MAIN_SHMEM_SEGMENT)
- return size;
+ /* size of buffer descriptors, plus alignment padding */
+ size = add_size(0, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ Mappings[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_req_size = size;
- if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
- {
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
- }
+ /* size of data pages, plus alignment padding */
+ size = add_size(0, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ Mappings[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
- if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
- {
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
- }
+ /* size of stuff controlled by freelist.c */
+ Mappings[STRATEGY_SHMEM_SEGMENT].shmem_req_size = StrategyShmemSize();
- if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
- {
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
- }
+ /* size of I/O condition variables, plus alignment padding */
+ size = add_size(0, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ Mappings[BUFFER_IOCV_SHMEM_SEGMENT].shmem_req_size = size;
- if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
- {
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
- }
-
- if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
- {
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
- }
+ /* size of checkpoint sort array in bufmgr.c */
+ Mappings[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
return size;
}
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2dbd81afc87..2cd278449f0 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -84,9 +84,23 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * TODO: Right now the minions of this function return the size of shared memory
+ * required in the main shared memory segment but add sizes required from other
+ * segments in the respective mappings. I think we should change this assymetry.
+ * It's only the buffer manager which adds sizes for other segments, but in
+ * future there may be others. Further the buffer manager related other segments
+ * are expected to hold only one resizable structure thus their size should be
+ * set only once when changing shared buffer pool size (i.e. when changin
+ * shared_buffers GUC). We shouldn't allow adding more structures to these
+ * segments, and thus restrict adding sizes to the corresponding mappings after
+ * the initial size is set.
+ *
+ * TODO: Also we should do something about numSemas, which is not required
+ * everywhere CalculateShmemSize is called.
*/
Size
-CalculateShmemSize(int *num_semaphores, int shmem_segment)
+CalculateShmemSize(int *num_semaphores)
{
Size size;
int numSemas;
@@ -113,7 +127,13 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize(shmem_segment));
+
+ /*
+ * Buffer manager adds estimates for memory requirements for every shared
+ * memory segment that it uses in the corresponding AnonymousMappings.
+ * Consider size required from only the main shared memory segment here.
+ */
+ size = add_size(size, BufferManagerShmemSize());
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
@@ -154,8 +174,15 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
+ /*
+ * All the shared memory allocations considered so far happen in the main
+ * shared memory segment.
+ */
+ Mappings[MAIN_SHMEM_SEGMENT].shmem_req_size = size;
+
/* might as well round it off to a multiple of a typical page size */
- size = add_size(size, 8192 - (size % 8192));
+ for (int segment = 0; segment < ANON_MAPPINGS; segment++)
+ Mappings[segment].shmem_req_size = add_size(Mappings[segment].shmem_req_size, 8192 - (Mappings[segment].shmem_req_size % 8192));
return size;
}
@@ -201,26 +228,30 @@ CreateSharedMemoryAndSemaphores(void)
{
PGShmemHeader *shim;
PGShmemHeader *seghdr;
- Size size;
int numSemas;
Assert(!IsUnderPostmaster);
+ CalculateShmemSize(&numSemas);
+
/* Decide if we use huge pages or regular size pages */
PrepareHugePages();
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
+ AnonymousMapping *mapping = &Mappings[segment];
+
+ mapping->shmem_segment = segment;
+
/* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas, segment);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", mapping->shmem_req_size);
/*
* Create the shmem segment.
*
* XXX: Do multiple shims are needed, one per segment?
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(mapping, &shim);
/*
* Make sure that huge pages are never reported as "unknown" while the
@@ -232,9 +263,13 @@ CreateSharedMemoryAndSemaphores(void)
InitShmemAccessInSegment(seghdr, segment);
/*
- * Create semaphores
+ * Shared memory for semaphores is allocated in the main shared memory.
+ * Hence they are allocated after the main segment is created. Patch
+ * proposed at https://commitfest.postgresql.org/patch/5997/ simplifies
+ * this.
*/
- PGReserveSemaphores(numSemas, segment);
+ if (segment == MAIN_SHMEM_SEGMENT)
+ PGReserveSemaphores(numSemas, segment);
/*
* Set up shared memory allocation mechanism
@@ -357,7 +392,9 @@ CreateOrAttachShmemStructs(void)
* InitializeShmemGUCs
*
* This function initializes runtime-computed GUCs related to the amount of
- * shared memory required for the current configuration.
+ * shared memory required for the current configuration. It assumes that the
+ * memory required by the shared memory segments is already calculated and is
+ * available in AnonymousMappings.
*/
void
InitializeShmemGUCs(void)
@@ -366,12 +403,16 @@ InitializeShmemGUCs(void)
Size size_b;
Size size_mb;
Size hp_size;
- int num_semas;
+ int num_semas = ProcGlobalSemas();
+ int i;
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
+ size_b = 0;
+ for (i = 0; i < ANON_MAPPINGS; i++)
+ size_b = add_size(size_b, Mappings[i].shmem_req_size);
+
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index e6cb919f0fc..90c21a97225 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -178,8 +178,8 @@ ShmemAllocInSegment(Size size, int shmem_segment)
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ MappingName(shmem_segment), size)));
return newSpace;
}
@@ -286,8 +286,8 @@ ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ MappingName(shmem_segment), size)));
Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 8d4d6cc3f33..c819608fff6 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4132,13 +4132,6 @@ PostgresSingleUserMain(int argc, char *argv[],
*/
process_shmem_requests();
- /*
- * Now that loadable modules have had their chance to request additional
- * shared memory, determine the value of any runtime-computed GUCs that
- * depend on the amount of shared memory required.
- */
- InitializeShmemGUCs();
-
/*
* Now that modules have been loaded, we can process any custom resource
* managers specified in the wal_consistency_checking GUC.
@@ -4151,6 +4144,13 @@ PostgresSingleUserMain(int argc, char *argv[],
*/
CreateSharedMemoryAndSemaphores();
+ /*
+ * Now that loadable modules have had their chance to request additional
+ * shared memory, determine the value of any runtime-computed GUCs that
+ * depend on the amount of shared memory required.
+ */
+ InitializeShmemGUCs();
+
/*
* Estimate number of openable files. This must happen after setting up
* semaphores, because on some platforms semaphores count as open files.
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index f8d34513c7f..47360a3d3d8 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -318,7 +318,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(int);
+extern Size BufferManagerShmemSize(void);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6ebda479ced..3baf418b3d1 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
+extern Size CalculateShmemSize(int *num_semaphores);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index a7b275b4db9..a1fa6b43fe3 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -27,6 +27,18 @@
#include "storage/dsm_impl.h"
#include "storage/spin.h"
+typedef struct AnonymousMapping
+{
+ int shmem_segment; /* TODO: Do we really need it? */
+ Size shmem_req_size; /* Required size of the segment */
+ Size shmem_size; /* Size of the actually used memory */
+ Size shmem_reserved; /* Size of the reserved mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
+} AnonymousMapping;
+
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
int32 magic; /* magic # to identify Postgres segments */
@@ -55,6 +67,8 @@ typedef struct ShmemSegment
#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
@@ -101,10 +115,11 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(Size size,
+extern PGShmemHeader *PGSharedMemoryCreate(AnonymousMapping *mapping,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
+extern const char *MappingName(int shmem_segment);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
--
2.34.1
[application/x-patch] 0010-WIP-Monitoring-views-20250918.patch (10.5K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/11-0010-WIP-Monitoring-views-20250918.patch)
download | inline diff:
From 86fb36cab8e2079fde361380ae42fe4dbabe7967 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 20 Aug 2025 10:55:27 +0530
Subject: [PATCH 10/16] WIP: Monitoring views
Modifies pg_shmem_allocations to report shared memory segment as well.
Adds pg_shmem_segments to report shared memory segment information.
TODO:
This commit should be merged with the earlier commit introducing
multiple shared memory segments.
Author: Ashutosh Bapat
---
doc/src/sgml/system-views.sgml | 9 +++
src/backend/catalog/system_views.sql | 7 +++
src/backend/storage/ipc/shmem.c | 90 ++++++++++++++++++++++------
src/include/catalog/pg_proc.dat | 12 +++-
src/include/storage/pg_shmem.h | 1 -
src/include/storage/shmem.h | 1 +
src/test/regress/expected/rules.out | 10 +++-
7 files changed, 108 insertions(+), 22 deletions(-)
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 89be9bc333f..7d14a6eca24 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -4167,6 +4167,15 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
</para></entry>
</row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>segment</structfield> <type>text</type>
+ </para>
+ <para>
+ The name of the shared memory segment concerning the allocation.
+ </para></entry>
+ </row>
+
<row>
<entry role="catalog_table_entry"><para role="column_definition">
<structfield>off</structfield> <type>int8</type>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 46fc28396de..f659dbb2f86 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -658,6 +658,13 @@ GRANT SELECT ON pg_shmem_allocations TO pg_read_all_stats;
REVOKE EXECUTE ON FUNCTION pg_get_shmem_allocations() FROM PUBLIC;
GRANT EXECUTE ON FUNCTION pg_get_shmem_allocations() TO pg_read_all_stats;
+CREATE VIEW pg_shmem_segments AS
+ SELECT * FROM pg_get_shmem_segments();
+
+REVOKE ALL ON pg_shmem_segments FROM PUBLIC;
+GRANT SELECT ON pg_shmem_segments TO pg_read_all_stats;
+REVOKE EXECUTE ON FUNCTION pg_get_shmem_segments() FROM PUBLIC;
+GRANT EXECUTE ON FUNCTION pg_get_shmem_segments() TO pg_read_all_stats;
CREATE VIEW pg_shmem_allocations_numa AS
SELECT * FROM pg_get_shmem_allocations_numa();
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 90c21a97225..9499f332e77 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -531,6 +531,7 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
result->size = size;
result->allocated_size = allocated_size;
result->location = structPtr;
+ result->shmem_segment = shmem_segment;
}
LWLockRelease(ShmemIndexLock);
@@ -582,13 +583,14 @@ mul_size(Size s1, Size s2)
Datum
pg_get_shmem_allocations(PG_FUNCTION_ARGS)
{
-#define PG_GET_SHMEM_SIZES_COLS 4
+#define PG_GET_SHMEM_SIZES_COLS 5
ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
HASH_SEQ_STATUS hstat;
ShmemIndexEnt *ent;
- Size named_allocated = 0;
+ Size named_allocated[ANON_MAPPINGS] = {0};
Datum values[PG_GET_SHMEM_SIZES_COLS];
bool nulls[PG_GET_SHMEM_SIZES_COLS];
+ int i;
InitMaterializedSRF(fcinfo, 0);
@@ -598,33 +600,42 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
- /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
- values[2] = Int64GetDatum(ent->size);
- values[3] = Int64GetDatum(ent->allocated_size);
- named_allocated += ent->allocated_size;
+ values[1] = CStringGetTextDatum(MappingName(ent->shmem_segment));
+ values[2] = Int64GetDatum((char *) ent->location - (char *) Segments[ent->shmem_segment].ShmemSegHdr);
+ values[3] = Int64GetDatum(ent->size);
+ values[4] = Int64GetDatum(ent->allocated_size);
+ named_allocated[ent->shmem_segment] += ent->allocated_size;
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
values, nulls);
}
/* output shared memory allocated but not counted via the shmem index */
- values[0] = CStringGetTextDatum("<anonymous>");
- nulls[1] = true;
- values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ for (i = 0; i < ANON_MAPPINGS; i++)
+ {
+ values[0] = CStringGetTextDatum("<anonymous>");
+ values[1] = CStringGetTextDatum(MappingName(i));
+ nulls[2] = true;
+ values[3] = Int64GetDatum(Segments[i].ShmemSegHdr->freeoffset - named_allocated[i]);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
/* output as-of-yet unused shared memory */
- nulls[0] = true;
- values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
- nulls[1] = false;
- values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ memset(nulls, 0, sizeof(nulls));
+
+ for (i = 0; i < ANON_MAPPINGS; i++)
+ {
+ nulls[0] = true;
+ values[1] = CStringGetTextDatum(MappingName(i));
+ values[2] = Int64GetDatum(Segments[i].ShmemSegHdr->freeoffset);
+ values[3] = Int64GetDatum(Segments[i].ShmemSegHdr->totalsize - Segments[i].ShmemSegHdr->freeoffset);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
LWLockRelease(ShmemIndexLock);
@@ -825,3 +836,46 @@ pg_numa_available(PG_FUNCTION_ARGS)
{
PG_RETURN_BOOL(pg_numa_init() != -1);
}
+
+/* SQL SRF showing shared memory segments */
+Datum
+pg_get_shmem_segments(PG_FUNCTION_ARGS)
+{
+#define PG_GET_SHMEM_SEGS_COLS 6
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ Datum values[PG_GET_SHMEM_SEGS_COLS];
+ bool nulls[PG_GET_SHMEM_SEGS_COLS];
+ int i;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* output all allocated entries */
+ for (i = 0; i < ANON_MAPPINGS; i++)
+ {
+ PGShmemHeader *shmhdr = Segments[i].ShmemSegHdr;
+ AnonymousMapping *segmapping = &Mappings[i];
+ int j;
+
+ if (shmhdr == NULL)
+ {
+ for (j = 0; j < PG_GET_SHMEM_SEGS_COLS; j++)
+ nulls[j] = true;
+ }
+ else
+ {
+ memset(nulls, 0, sizeof(nulls));
+ values[0] = Int32GetDatum(i);
+ values[1] = CStringGetTextDatum(MappingName(i));
+ values[2] = Int64GetDatum(shmhdr->totalsize);
+ values[3] = Int64GetDatum(shmhdr->freeoffset);
+ values[4] = Int64GetDatum(segmapping->shmem_size);
+ values[5] = Int64GetDatum(segmapping->shmem_reserved);
+ }
+
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
+
+ return (Datum) 0;
+}
+
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 1e53b7a4ae5..6c37fa90c89 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8568,8 +8568,8 @@
{ oid => '5052', descr => 'allocations from the main shared memory segment',
proname => 'pg_get_shmem_allocations', prorows => '50', proretset => 't',
provolatile => 'v', prorettype => 'record', proargtypes => '',
- proallargtypes => '{text,int8,int8,int8}', proargmodes => '{o,o,o,o}',
- proargnames => '{name,off,size,allocated_size}',
+ proallargtypes => '{text,text,int8,int8,int8}', proargmodes => '{o,o,o,o,o}',
+ proargnames => '{name,segment,off,size,allocated_size}',
prosrc => 'pg_get_shmem_allocations' },
{ oid => '4099', descr => 'Is NUMA support available?',
@@ -8592,6 +8592,14 @@
proargmodes => '{o,o,o}', proargnames => '{name,type,size}',
prosrc => 'pg_get_dsm_registry_allocations' },
+# shared memory segments
+{ oid => '5101', descr => 'shared memory segments',
+ proname => 'pg_get_shmem_segments', prorows => '6', proretset => 't',
+ provolatile => 'v', prorettype => 'record', proargtypes => '',
+ proallargtypes => '{int4,text,int8,int8,int8,int8}', proargmodes => '{o,o,o,o,o,o}',
+ proargnames => '{id,name,size,freeoffset,mapping_size,mapping_reserved_size}',
+ prosrc => 'pg_get_shmem_segments' },
+
# buffer lookup table
{ oid => '5102',
descr => 'shared buffer lookup table',
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index a1fa6b43fe3..715f6acb5dd 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -69,7 +69,6 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
-
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 910c43f54f4..64ff5a286ba 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -71,6 +71,7 @@ typedef struct
void *location; /* location in shared mem */
Size size; /* # bytes requested for the structure */
Size allocated_size; /* # bytes actually allocated */
+ int shmem_segment; /* segment in which the structure is allocated */
} ShmemIndexEnt;
#endif /* SHMEM_H */
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 760bb13fe95..e73314b5ef0 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1764,14 +1764,22 @@ pg_shadow| SELECT pg_authid.rolname AS usename,
LEFT JOIN pg_db_role_setting s ON (((pg_authid.oid = s.setrole) AND (s.setdatabase = (0)::oid))))
WHERE pg_authid.rolcanlogin;
pg_shmem_allocations| SELECT name,
+ segment,
off,
size,
allocated_size
- FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size);
+ FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, segment, off, size, allocated_size);
pg_shmem_allocations_numa| SELECT name,
numa_node,
size
FROM pg_get_shmem_allocations_numa() pg_get_shmem_allocations_numa(name, numa_node, size);
+pg_shmem_segments| SELECT id,
+ name,
+ size,
+ freeoffset,
+ mapping_size,
+ mapping_reserved_size
+ FROM pg_get_shmem_segments() pg_get_shmem_segments(id, name, size, freeoffset, mapping_size, mapping_reserved_size);
pg_stat_activity| SELECT s.datid,
d.datname,
s.pid,
--
2.34.1
[application/x-patch] 0013-Update-sizes-and-addresses-of-shared-memory-20250918.patch (2.5K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/12-0013-Update-sizes-and-addresses-of-shared-memory-20250918.patch)
download | inline diff:
From 7cdcf605c4d67aa35f66e42c98a12c1b97c20b69 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 21 Aug 2025 15:44:24 +0530
Subject: [PATCH 13/16] Update sizes and addresses of shared memory mapping and
shared memory structures
Update totalsize and end address in segment and mapping: Once a shared
memory segment has been resized, the total size and end address of the
same needs to be updated in the corresponding AnonymousMapping and
Segment structure.
Update allocated_size for resized shared memory structure: Reallocating
the shared memory structure after resizing needs a bit more work. But at
least update the allocated_size as well along with the size of shared
memory structure.
Author: Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 4 ++++
src/backend/storage/ipc/shmem.c | 6 +++++-
2 files changed, 9 insertions(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index ba8613678f6..54d335b2e5d 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1021,6 +1021,8 @@ AnonymousShmemResize(void)
for(int i = 0; i < ANON_MAPPINGS; i++)
{
AnonymousMapping *m = &Mappings[i];
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmem_hdr = segment->ShmemSegHdr;
#ifdef MAP_HUGETLB
if (huge_pages_on && (m->shmem_req_size % hugepagesize != 0))
@@ -1067,6 +1069,8 @@ AnonymousShmemResize(void)
reinit = true;
m->shmem_size = m->shmem_req_size;
+ shmem_hdr->totalsize = m->shmem_size;
+ segment->ShmemEnd = m->shmem + m->shmem_size;
}
if (reinit)
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 2a197540300..0f9abf69fd5 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -504,13 +504,17 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
*
* XXX: There is an implicit assumption this can only happen in
* "resizable" segments, where only one shared structure is allowed.
- * This has to be implemented more cleanly.
+ * This has to be implemented more cleanly. Probably we should implement
+ * ShmemReallocRawInSegment functionality just to adjust the size
+ * according to alignment, return the allocated size and update the
+ * mapping offset.
*/
if (result->size != size)
{
Size delta = size - result->size;
result->size = size;
+ result->allocated_size = size;
/* Reflect size change in the shared segment */
SpinLockAcquire(Segments[shmem_segment].ShmemLock);
--
2.34.1
[application/x-patch] 0011-Allow-to-resize-shared-memory-without-resta-20250918.patch (40.7K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/13-0011-Allow-to-resize-shared-memory-without-resta-20250918.patch)
download | inline diff:
From 78bc0a49f8ebe17927abd66164764745ecc6d563 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 14:16:55 +0200
Subject: [PATCH 11/16] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers, adjusts its size using ftruncate and adjust reservation
permissions with mprotect. One elected process signals the postmaster
to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f87909fc000-7f8798248000 rw-s /memfd:strategy (deleted)
7f8798248000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a4e84000 rw-s /memfd:checkpoint (deleted)
7f87a4e84000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1b42000 rw-s /memfd:iocv (deleted)
7f87b1b42000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cb59c000 rw-s /memfd:descriptors (deleted)
7f87cb59c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f87ece38000 rw-s /memfd:buffers (deleted)
^ buffers content, ~247 MB
7f87ece38000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~2210 MB
7f8877066000-7f887e7d0000 rw-s /memfd:main (deleted)
7f887e7d0000-7f8890a00000 ---s /memfd:main (deleted)
-- 512 MB
7f87909fc000-7f879866a000 rw-s /memfd:strategy (deleted)
7f879866a000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a50f4000 rw-s /memfd:checkpoint (deleted)
7f87a50f4000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1d82000 rw-s /memfd:iocv (deleted)
7f87b1d82000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cba1c000 rw-s /memfd:descriptors (deleted)
7f87cba1c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f8804fb8000 rw-s /memfd:buffers (deleted)
^ buffers content, ~632 MB
7f8804fb8000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~1824 MB
7f8877066000-7f887e950000 rw-s /memfd:main (deleted)
7f887e950000-7f8890a00000 ---s /memfd:main (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it receives something.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 443 ++++++++++++++++++
src/backend/postmaster/checkpointer.c | 12 +-
src/backend/postmaster/postmaster.c | 18 +
src/backend/storage/buffer/buf_init.c | 60 ++-
src/backend/storage/ipc/ipci.c | 15 +-
src/backend/storage/ipc/procsignal.c | 46 ++
src/backend/storage/ipc/shmem.c | 23 +-
src/backend/tcop/postgres.c | 10 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_parameters.dat | 3 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 26 +
src/include/storage/pmsignal.h | 1 +
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
17 files changed, 631 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index b85911bdfc4..dc4eeeee56a 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -96,6 +102,13 @@ void *UsedShmemSegAddr = NULL;
AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/*
* Anonymous mapping layout we use looks like this:
*
@@ -147,6 +160,49 @@ static double SHMEM_RESIZE_RATIO[6] = {
*/
static bool huge_pages_on = false;
+/*
+ * Flag telling that we have prepared the memory layout to be resizable. If
+ * false after all shared memory segments creation, it means we failed to setup
+ * needed layout and falled back to the regular non-resizable approach.
+ */
+static bool shmem_resizable = false;
+
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -906,6 +962,346 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * ftruncate, which will fail if is not sufficient space to expand the anon
+ * file. When finished, based on the new and old values initialize new buffer
+ * blocks if any.
+ *
+ * If reinitializing took place, as the last step this function does buffers
+ * reinitialization as well and broadcasts the new value of NSharedBuffers. All
+ * of that needs to be done only by one backend, the first one that managed to
+ * grab the ShmemResizeLock.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ int mmap_flags = PG_MMAP_FLAGS;
+ Size hugepagesize;
+
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+#ifndef MAP_HUGETLB
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the new size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+ }
+#endif
+
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ CalculateShmemSize(&numSemas);
+
+ for(int i = 0; i < ANON_MAPPINGS; i++)
+ {
+ AnonymousMapping *m = &Mappings[i];
+
+#ifdef MAP_HUGETLB
+ if (huge_pages_on && (m->shmem_req_size % hugepagesize != 0))
+ m->shmem_req_size += hugepagesize - (m->shmem_req_size % hugepagesize);
+#endif
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == m->shmem_req_size)
+ continue;
+
+ if (m->shmem_reserved < m->shmem_req_size)
+ ereport(ERROR,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved"),
+ errhint("You may need to increase \"max_available_memory\".")));
+
+ elog(DEBUG1, "segment[%s]: resize from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ m->shmem_req_size, m->shmem);
+
+ /* Resize the backing anon file. */
+ if(ftruncate(m->segment_fd, m->shmem_req_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ /* Adjust memory accessibility */
+ if(mprotect(m->shmem, m->shmem_req_size, PROT_READ | PROT_WRITE) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not mprotect anonymous shared memory for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ /* If shrinking, make reserved space unavailable again */
+ if(m->shmem_req_size < m->shmem_size &&
+ mprotect(m->shmem + m->shmem_req_size, m->shmem_size - m->shmem_req_size, PROT_NONE) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not mprotect reserved shared memory for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ reinit = true;
+ m->shmem_size = m->shmem_req_size;
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * Note that it's enough when only one backend will do that,
+ * even the ShmemInitStruct part. The reason is that resized
+ * shared memory will maintain the same addresses, meaning that
+ * all the pointers are still valid, and we only need to update
+ * structures size in the ShmemIndex once -- any other backend
+ * will pick up this shared structure from the index.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ BufferManagerShmemInit(NBuffersOld);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Wait for all ProcSignal participants
+ * to join the barrier, then do the resize and wait on the barrier until all
+ * participating finish resizing as well -- otherwise we face danger of
+ * inconsistency between backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ Assert(IsUnderPostmaster);
+
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * First thing to do after attaching to the barrier is to wait for others.
+ * We can't simply use BarrierArriveAndWait, because backends might arrive
+ * here in disjoint groups, e.g. first two backends, pause, then second two
+ * backends. If the resize is quick enough that can lead to a situation
+ * when the first group is already finished before the second has appeared,
+ * and the barrier will only synchonize withing those groups.
+ */
+ if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
+ WaitForProcSignalBarrierReceived(
+ pg_atomic_read_u64(&ShmemCtrl->Generation));
+
+ /*
+ * Now start the procedure, and elect one backend to ping postmaster to do
+ * the same.
+ *
+ * XXX: If we need to be able to abort resizing, this has to be done later,
+ * after the SHMEM_RESIZE_DONE.
+ */
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ {
+ Assert(IsUnderPostmaster);
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ }
+
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ BarrierDetach(barrier);
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Start resizing procedure, making sure all existing processes will have
+ * consistent view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * We use dynamic barrier to help dealing with backends that were spawned
+ * during the resize.
+ */
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d",
+ NBuffers, NBuffersOld);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ *
+ * Those phases are ensured by joining the shared barrier associated with
+ * the procedure. Since resizing takes time, we need to take into account
+ * that during that time:
+ *
+ * - New backends can be spawned. They will check status of the barrier
+ * early during the bootstrap, and wait until everything is over to work
+ * with the new NBuffers value.
+ *
+ * - Old backends can exit before attempting to resize. Synchronization
+ * used between backends relies on ProcSignalBarrier and waits for all
+ * participants received the message at the beginning to gather all
+ * existing backends.
+ *
+ * - Some backends might be blocked and not responsing either before or
+ * after receiving the message. In the first case such backend still
+ * have ProcSignalSlot and should be waited for, in the second case
+ * shared barrier will make sure we still waiting for those backends. In
+ * any case there is an unbounded wait.
+ *
+ * - Backends might join barrier in disjoint groups with some time in
+ * between. That means that relying only on the shared dynamic barrier is
+ * not enough -- it will only synchronize resize procedure withing those
+ * groups. That's why we wait first for all participants of ProcSignal
+ * mechanism who received the message.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ pg_atomic_init_u64(&ShmemCtrl->Generation,
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
+
+ /* To order everything after setting Generation value */
+ pg_memory_barrier();
+
+ /*
+ * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
+ * the rest of the pack has started the procedure and it can resize shared
+ * memory as well.
+ *
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier. In fact even if we would like to
+ * wait here, it wouldn't be possible -- we're in the postmaster, without
+ * any waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1217,3 +1613,50 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier()
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ /* Nothing to do if resizing is not started */
+ if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
+ return;
+
+ BarrierAttach(barrier);
+
+ /* Otherwise wait through all available phases */
+ while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
+ {
+ ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
+ BarrierPhase(barrier))));
+
+ BarrierArriveAndWait(barrier, 0);
+ }
+
+ BarrierDetach(barrier);
+}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /*
+ * The barrier is missing here, it will be initialized right before
+ * starting the resizing process as a convenient way to reset it.
+ */
+
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+
+ /* shmem_resizable should be initialized by now */
+ ShmemCtrl->Resizable = shmem_resizable;
+ }
+}
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index e84e8663e96..ef3f84a55f5 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -654,9 +654,12 @@ CheckpointerMain(const void *startup_data, size_t startup_data_len)
static void
ProcessCheckpointerInterrupts(void)
{
- if (ProcSignalBarrierPending)
- ProcessProcSignalBarrier();
-
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
if (ConfigReloadPending)
{
ConfigReloadPending = false;
@@ -676,6 +679,9 @@ ProcessCheckpointerInterrupts(void)
UpdateSharedMemoryConfig();
}
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+
/* Perform logging of memory contexts of this process */
if (LogMemoryContextPending)
ProcessLogMemoryContextInterrupt();
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index b59d20b4ac2..ba9528d5dfa 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -426,6 +426,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1697,6 +1698,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2042,6 +2046,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
@@ -3862,6 +3877,9 @@ process_pm_pmsignal(void)
request_state_update = true;
}
+ if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
+ AnonymousShmemResize();
+
/*
* Try to advance postmaster's state machine, if a child requests it.
*/
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6f148d1d80b..0e72e373193 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -18,6 +18,7 @@
#include "storage/buf_internals.h"
#include "storage/pg_shmem.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -63,18 +64,28 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * postmaster, or in a standalone backend) or during shared-memory resize. Size
+ * of data structures initialized here depends on NBuffers, and to be able to
+ * change NBuffers without a restart we store each structure into a separate
+ * shared memory segment, which could be resized on demand.
+ *
+ * FirstBufferToInit tells where to start initializing buffers. For
+ * initialization it always will be zero, but when resizing shared-memory it
+ * indicates the number of already initialized buffers.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(void)
+BufferManagerShmemInit(int FirstBufferToInit)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
+ elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
+ FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -111,34 +122,35 @@ BufferManagerShmemInit(void)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory, or in all cases when resizing shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
+
+#ifndef EXEC_BACKEND
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = FirstBufferToInit; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
- ClearBufferTag(&buf->tag);
+ ClearBufferTag(&buf->tag);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- buf->buf_id = i;
+ buf->buf_id = i;
- pgaio_wref_clear(&buf->io_wref);
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
+#endif
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2cd278449f0..bd75f06047e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -171,6 +171,14 @@ CalculateShmemSize(int *num_semaphores)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, AioShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -333,7 +341,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit();
+ BufferManagerShmemInit(0);
/*
* Set up lock manager
@@ -345,6 +353,11 @@ CreateOrAttachShmemStructs(void)
*/
PredicateLockShmemInit();
+ /*
+ * Set up shared memory resize manager
+ */
+ ShmemControlInit();
+
/*
* Set up process table
*/
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index eb3ceaae809..2160d258fa7 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -113,6 +114,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -176,6 +181,43 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -615,6 +657,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9499f332e77..2a197540300 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -498,17 +498,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
+ *
+ * XXX: There is an implicit assumption this can only happen in
+ * "resizable" segments, where only one shared structure is allowed.
+ * This has to be implemented more cleanly.
*/
if (result->size != size)
{
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
+ Size delta = size - result->size;
+
+ result->size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
}
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index c819608fff6..15e9dde41d1 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -62,6 +62,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4317,6 +4318,15 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /* Verify the shared barrier, if it's still active: join and wait. */
+ WaitOnShmemBarrier();
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 7553f6eacef..82cee6b8877 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
WAL_RECEIVER_WAIT_START "Waiting for startup process to send initial data for streaming replication."
@@ -355,6 +357,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index c94f3fc3c80..5c534cee2ac 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -1098,13 +1098,14 @@
# We sometimes multiply the number of shared buffers by two without
# checking for overflow, so we mustn't allow more than INT_MAX / 2.
-{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
variable => 'NBuffers',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
+ assign_hook => 'assign_shared_buffers'
},
# TODO: should this be PGC_POSTMASTER?
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 47360a3d3d8..51ce6ebcf6c 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -317,7 +317,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_skipped);
/* in buf_init.c */
-extern void BufferManagerShmemInit(void);
+extern void BufferManagerShmemInit(int);
extern Size BufferManagerShmemSize(void);
/* in localbuf.c */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..847f56a36dc 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index 06a1ffd4b08..cba586027a7 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -85,6 +85,7 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
+PG_LWLOCK(54, ShmemResize)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 715f6acb5dd..eba28ce8a5c 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -69,6 +70,25 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+ pg_atomic_uint64 Generation;
+ bool Resizable;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -123,6 +143,12 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 428aa3fd68a..1a55bf57a70 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,6 +42,7 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 2733bbb8c5b..97033f84dce 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index e90af5b2ad3..ee5c2dd0ad4 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2762,6 +2762,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.34.1
[application/x-patch] 0012-Initial-value-of-shared_buffers-or-NBuffers-20250918.patch (3.7K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/14-0012-Initial-value-of-shared_buffers-or-NBuffers-20250918.patch)
download | inline diff:
From cd2fd975382c629a5fa825a32a1b35e485564001 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 1 Sep 2025 15:40:41 +0530
Subject: [PATCH 12/16] Initial value of shared_buffers (or NBuffers)
The assign_hook for shared_buffers (assign_shared_buffers()) is called twice
during server startup. First time it sets the default value of shared_buffers,
followed by a second time when it sets the value specified in the configuration
file or on the command line. At those times the shared buffer pool is yet to be
initialized. Hence there is no need to keep the GUC change pending or going
through the entire process of resizing memory maps, reinitializing the shared memory
and process synchronization. Instead the given value should be assigned directly to
NBuffers, which will be used when creating the shared
memory and also when initializing the buffer pool the first time. Any changes
to shared_buffer after that will need remapping the shared memory segment and
synchronize buffer pool reinitialization across the backends.
If BufferBlocks is not initilized assign_shared_buffers() sets the given
value to NBuffers directly. Otherwise it marks the change as pending and
sets the flag pending_pm_shmem_resize so that Postmaster can start the
buffer pool reinitialization.
TODO:
1. The change depends upon the C convention that the global pointer variables
being initialized to NULL. May be initialize BufferBlocks to NULL explicitly.
2. We might think of a better way to check whether buffer pool has been
initialized or not. See comment in assign_shared_buffers().
Author: Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 42 ++++++++++++++++++++++++++---------
1 file changed, 32 insertions(+), 10 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index dc4eeeee56a..ba8613678f6 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1168,20 +1168,42 @@ ProcessBarrierShmemResize(Barrier *barrier)
}
/*
- * GUC assign hook for shared_buffers. It's recommended for an assign hook to
- * be as minimal as possible, thus we just request shared memory resize and
- * remember the previous value.
+ * GUC assign hook for shared_buffers.
+ *
+ * When setting the GUC first time after starting the server, the GUC value is
+ * changed immediately since there is not shared memory setup yet.
+ *
+ * After the shared memory is setup, changing the GUC value requires resizing and
+ * reiniatializing (at least parts of) the shared memory structures related to
+ * shared buffers. That's a long and complicated process. It's recommended for
+ * an assign hook to be as minimal as possible, thus we just request shared
+ * memory resize and remember the previous value.
*/
void
assign_shared_buffers(int newval, void *extra, bool *pending)
{
- elog(DEBUG1, "Received SIGHUP for shmem resizing");
-
- pending_pm_shmem_resize = true;
- *pending = true;
- NBuffersPending = newval;
-
- NBuffersOld = NBuffers;
+ /*
+ * TODO: If a backend joins while the buffer resizing is in progress or it
+ * reads a value of shared_buffers from configuration which is different from
+ * the value being used by existing backends, this method may not work. Need
+ * to think of a better solution.
+ */
+ if (BufferBlocks)
+ {
+ elog(DEBUG1, "bufferpool is already initialized with size = %d, reinitializing it with size = %d",
+ NBuffers, newval);
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ NBuffersOld = NBuffers;
+ }
+ else
+ {
+ elog(DEBUG1, "initializing buffer pool with size = %d", newval);
+ NBuffers = newval;
+ *pending = false;
+ pending_pm_shmem_resize = false;
+ }
}
/*
--
2.34.1
[application/x-patch] 0014-Support-shrinking-shared-buffers-20250918.patch (12.5K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/15-0014-Support-shrinking-shared-buffers-20250918.patch)
download | inline diff:
From df1fbd464901d3097d9e74ae06a70bc629948d14 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 19 Jun 2025 17:38:29 +0200
Subject: [PATCH 14/16] Support shrinking shared buffers
Buffer eviction
===============
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
Buffer eviction requires a separate barrier phase for two reasons:
1. No other backend should map a new page to any of buffers being
evicted when eviction is in progress. So they wait while eviction is
in progress.
2. Since a pinned buffer has the pin recorded in the backend local
memory as well as the buffer descriptor (which is in shared memory),
eviction should not coincide with remapping the shared memory of a
backend. Otherwise we might loose consistency of local and shared
pinning records. Hence it needs to be carried out in
ProcessBarrierShmemResize() and not in AnonymousShmemResize() as
indicated by now removed comment.
If a buffer being evicted is pinned, we raise a FATAL error but this should
improve. There are multiple options 1. to wait for the pinned buffer to get
unpinned, 2. the backend is killed or it itself cancels the query or 3.
rollback the operation. Note that option 1 and 2 would require the pinning
related local and shared records to be accessed. But we need infrastructure to
do either of this right now.
Removing the evicted buffers from buffer ring
=============================================
If the buffer pool has been shrunk, the buffers in the buffer ring may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition and masks any bug that
introduces an invalid buffer into the ring. The alternative fix is more
complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure thus accessible when processing buffer
resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work.
Author: Ashutosh Bapat
Reviewed-by: Tomas Vondra
---
src/backend/port/sysv_shmem.c | 42 ++++++---
src/backend/storage/buffer/bufmgr.c | 93 +++++++++++++++++++
src/backend/storage/buffer/freelist.c | 18 +++-
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/bufmgr.h | 1 +
src/include/storage/pg_shmem.h | 1 +
6 files changed, 139 insertions(+), 17 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 54d335b2e5d..9e1b2c3201f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -993,14 +993,6 @@ AnonymousShmemResize(void)
*/
pending_pm_shmem_resize = false;
- /*
- * XXX: Currently only increasing of shared_buffers is supported. For
- * decreasing something similar has to be done, but buffer blocks with
- * data have to be drained first.
- */
- if(NBuffersOld > NBuffers)
- return false;
-
#ifndef MAP_HUGETLB
/* PrepareHugePages should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
@@ -1099,11 +1091,14 @@ AnonymousShmemResize(void)
* all the pointers are still valid, and we only need to update
* structures size in the ShmemIndex once -- any other backend
* will pick up this shared structure from the index.
- *
- * XXX: This is the right place for buffer eviction as well.
*/
BufferManagerShmemInit(NBuffersOld);
+ /*
+ * Wipe out the evictor PID so that it can be used for the next
+ * buffer resizing operation.
+ */
+ ShmemCtrl->evictor_pid = 0;
/* If all fine, broadcast the new value */
pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
}
@@ -1156,11 +1151,31 @@ ProcessBarrierShmemResize(Barrier *barrier)
* XXX: If we need to be able to abort resizing, this has to be done later,
* after the SHMEM_RESIZE_DONE.
*/
- if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+
+ /*
+ * Evict extra buffers when shrinking shared buffers. We need to do this
+ * while the memory for extra buffers is still mapped i.e. before remapping
+ * the shared memory segments to a smaller memory area.
+ */
+ if (NBuffersOld > NBuffersPending)
{
- Assert(IsUnderPostmaster);
- SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
+
+ /*
+ * TODO: If the buffer eviction fails for any reason, we should
+ * gracefully rollback the shared buffer resizing and try again. But the
+ * infrastructure to do so is not available right now. Hence just raise
+ * a FATAL so that the system restarts.
+ */
+ if (!EvictExtraBuffers(NBuffersPending, NBuffersOld))
+ elog(FATAL, "buffer eviction failed");
+
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_EVICT))
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
}
+ else
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
AnonymousShmemResize();
@@ -1684,5 +1699,6 @@ ShmemControlInit(void)
/* shmem_resizable should be initialized by now */
ShmemCtrl->Resizable = shmem_resizable;
+ ShmemCtrl->evictor_pid = 0;
}
}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index fe470de63f2..5424c405b44 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/read_stream.h"
#include "storage/smgr.h"
@@ -7422,3 +7423,95 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int newBufSize, int oldBufSize)
+{
+ bool result = true;
+
+ /*
+ * If the buffer being evicated is locked, this function will need to wait.
+ * This function should not be called from a Postmaster since it can not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * Let only one backend perform eviction. We could split the work across all
+ * the backends but that doesn't seem necessary.
+ *
+ * The first backend to acquire ShmemResizeLock, sets its own PID as the
+ * evictor PID for other backends to know that the eviction is in progress or
+ * has already been performed. The evictor backend releases the lock when it
+ * finishes eviction. While the eviction is in progress, backends other than
+ * evictor backend won't be able to take the lock. They won't perform
+ * eviction. A backend may acquire the lock after eviction has completed, but
+ * it will not perform eviction since the evictor PID is already set. Evictor
+ * PID is reset only when the buffer resizing finishes. Thus only one backend
+ * will perform eviction in a given instance of shared buffers resizing.
+ *
+ * Any backend which acquires this lock will release it before the eviction
+ * phase finishes, hence the same lock can be reused for the next phase of
+ * resizing buffers.
+ */
+ if (LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ if (ShmemCtrl->evictor_pid == 0)
+ {
+ ShmemCtrl->evictor_pid = MyProcPid;
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = newBufSize + 1; buf <= oldBufSize; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u32(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is
+ * going one hence unlocked precheck should be safe and saves
+ * some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ /*
+ * XXX: Looks like CurrentResourceOwner can be NULL here, find
+ * another one in that case?
+ * */
+ if (CurrentResourceOwner)
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+ }
+ LWLockRelease(ShmemResizeLock);
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 0bfbbb096d6..db8aafdaf8c 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -630,12 +630,22 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot has been
+ * invalidated when buffer pool was shrunk, tell the caller to allocate a new
+ * buffer with the normal allocation strategy. He will then fill this slot
+ * by calling AddBufferToRing with the new buffer.
+ *
+ * TODO: Ideally we would want to check for bufnum > NBuffers only once
+ * after every time the buffer pool is shrunk so as to catch any runtime
+ * bugs that introduce invalid buffers in the ring. But that is complicated.
+ * The BufferAccessStrategy objects are not accessible outside the
+ * ScanState. Hence we can not purge the buffers while evicting the buffers.
+ * After the resizing is finished, it's not possible to notice when we touch
+ * the first of those objects and the last of objects. See if this can
+ * fixed.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer || bufnum > NBuffers)
return NULL;
/*
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 82cee6b8877..9a6a6275305 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -156,6 +156,7 @@ REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it c
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_EVICT "Waiting for other backends to finish buffer evication phase."
SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 51ce6ebcf6c..c91a42fc598 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -315,6 +315,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_evicted,
int32 *buffers_flushed,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int fromBuf, int toBuf);
/* in buf_init.c */
extern void BufferManagerShmemInit(int);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index eba28ce8a5c..0a59746b472 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -77,6 +77,7 @@ extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
typedef struct
{
pg_atomic_uint32 NSharedBuffers;
+ pid_t evictor_pid;
Barrier Barrier;
pg_atomic_uint64 Generation;
bool Resizable;
--
2.34.1
[application/x-patch] 0015-Reinitialize-StrategyControl-after-resizing-20250918.patch (19.1K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/16-0015-Reinitialize-StrategyControl-after-resizing-20250918.patch)
download | inline diff:
From d389c5f5948c4c577b480ecf7bf0de551016d447 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 19 Jun 2025 17:38:51 +0200
Subject: [PATCH 15/16] Reinitialize StrategyControl after resizing buffers
... and BgBufferSync and ClockSweepTick adjustments
Reinitializing strategry control area
=====================================
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
2. &StrategyControl->buffer_strategy_lock need not be initialized again.
3. nextVictimBuffer, completePasses and numBufferAllocs are viewed in
the context of NBuffers. Now that NBuffers itself has changed, those
three do not make sense. Reset them as if the server has restarted
again.
Ability to delay resizing operation
===================================
This commit introduces a flag delay_shmem_resize, which postgresql
backends and workers can use to signal the coordinator to delay resizing
operation. Background writer sets this flag when its scanning buffers.
Background writer operation
===========================
Background writer is blocked when the actual resizing is in progress. It
stops a scan in progress when it sees that the resizing has begun or is
about to begin. Once the buffer resizing is finished, before resuming
the regular operation, bgwriter resets the information saved so far.
This information is viewed in the context of NBuffers and hence does not
make sense after resizing which chanegs NBuffers.
Buffer lookup table
===================
Right now there is no way to free shared memory. Even if we shrink the
buffer lookup table when shrinking the buffer pool the unused hash table
entries can not be freed. When we expand the buffer pool, more entries
can be allocated but we can not resize the hash table directory without
rehashing all the entries. Just allocating more entries will lead to
more contention. Hence we setup the buffer lookup table considering the
maximum possible size of the buffer pool which is MaxAvailableMemory
only once at the beginning. Shared buffer lookup table and
StrategyControl are not resized even if the buffer pool is resized hence
they are allocated in the main shared memory segment
TODO:
====
1. The way BgBufferSync is written today, it packs four functionalities:
setting up the buffer sync state, performing the buffer sync,
resetting the buffer sync state when bgwriter_lru_maxpages <= 0 and
setting it up again after bgwriter_lru_maxpages > 0. That makes the
code hard to read. It will be good to divide this function into 3/4
different functions each performing one functionality. Then pack all
the state (the local variables from that function converted to static
global) into a structure, which is passed to these functions. Once
that happens BgBufferSyncReset() will call one of the functions to
reset the state when buffer pool is resized.
2. The condition (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) ==
NBuffers) checked in BgBufferSync() to check whether buffer resizing
is "about to begin" is wrong. NBuffers it not changed, until
AnonymousShmemResize() is called and it wont' be called unless
BgBufferSync() finishes if it has already begun. Need a better
condition to check whether buffer resizing is about to begin.
Author: Ashutosh Bapat
Reviewed-by: Tomas Vondra
---
src/backend/port/sysv_shmem.c | 23 ++++++--
src/backend/storage/buffer/buf_init.c | 19 +++++--
src/backend/storage/buffer/buf_table.c | 9 ++-
src/backend/storage/buffer/bufmgr.c | 72 ++++++++++++++++++------
src/backend/storage/buffer/freelist.c | 77 ++++++++++++++++++++++++--
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 1 +
src/include/storage/ipc.h | 1 +
src/include/storage/pg_shmem.h | 5 +-
9 files changed, 170 insertions(+), 38 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 9e1b2c3201f..3be28e228ae 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -104,6 +104,7 @@ AnonymousMapping Mappings[ANON_MAPPINGS];
/* Flag telling postmaster that resize is needed */
volatile bool pending_pm_shmem_resize = false;
+volatile bool delay_shmem_resize = false;
/* Keeps track of the previous NBuffers value */
static int NBuffersOld = -1;
@@ -144,12 +145,11 @@ static int NBuffersPending = -1;
* makes sense to evaluate them more precise.
*/
static double SHMEM_RESIZE_RATIO[6] = {
- 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.15, /* MAIN_SHMEM_SEGMENT */
0.6, /* BUFFERS_SHMEM_SEGMENT */
0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
- 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -225,8 +225,6 @@ MappingName(int shmem_segment)
return "iocv";
case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
return "checkpoint";
- case STRATEGY_SHMEM_SEGMENT:
- return "strategy";
default:
return "unknown";
}
@@ -1125,13 +1123,17 @@ ProcessBarrierShmemResize(Barrier *barrier)
{
Assert(IsUnderPostmaster);
- elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
- NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize, delay_shmem_resize);
/* Wait until we have seen the new NBuffers value */
if (!pending_pm_shmem_resize)
return false;
+ /* Wait till this process becomes ready to resize buffers. */
+ if (delay_shmem_resize)
+ return false;
+
/*
* First thing to do after attaching to the barrier is to wait for others.
* We can't simply use BarrierArriveAndWait, because backends might arrive
@@ -1182,6 +1184,15 @@ ProcessBarrierShmemResize(Barrier *barrier)
/* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * Before resuming regular background writer activity, adjust the
+ * statistics collected so far.
+ */
+ BgBufferSyncReset(NBuffersOld, NBuffers);
+ }
+
BarrierDetach(barrier);
return true;
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 0e72e373193..be64fa5a136 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -152,8 +152,15 @@ BufferManagerShmemInit(int FirstBufferToInit)
}
#endif
- /* Init other shared buffer-management stuff */
- StrategyInitialize(!foundDescs);
+ /*
+ * Init other shared buffer-management stuff from scratch configuring buffer
+ * pool the first time. If we are just resizing buffer pool adjust only the
+ * required structures.
+ */
+ if (FirstBufferToInit == 0)
+ StrategyInitialize(!foundDescs);
+ else
+ StrategyReInitialize(FirstBufferToInit);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
@@ -184,9 +191,6 @@ BufferManagerShmemSize(void)
size = add_size(size, mul_size(NBuffers, BLCKSZ));
Mappings[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
- /* size of stuff controlled by freelist.c */
- Mappings[STRATEGY_SHMEM_SEGMENT].shmem_req_size = StrategyShmemSize();
-
/* size of I/O condition variables, plus alignment padding */
size = add_size(0, mul_size(NBuffers,
sizeof(ConditionVariableMinimallyPadded)));
@@ -196,5 +200,10 @@ BufferManagerShmemSize(void)
/* size of checkpoint sort array in bufmgr.c */
Mappings[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
+ /* Allocations in the main memory segment, at the end. */
+
+ /* size of stuff controlled by freelist.c */
+ size = add_size(0, StrategyShmemSize());
+
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 18a78967138..e5a97e557d9 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -65,11 +65,18 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
+ /*
+ * The shared buffer look up table is set up only once with maximum possible
+ * entries considering maximum size of the buffer pool. It is not resized
+ * after that even if the buffer pool is resized. Hence it is allocated in
+ * the main shared memory segment and not in a resizeable shared memory
+ * segment.
+ */
SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE,
- STRATEGY_SHMEM_SEGMENT);
+ MAIN_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 5424c405b44..48c46d5b963 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3580,6 +3580,32 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_DONE(NBuffers, num_written, num_to_scan);
}
+/*
+ * Information saved between BgBufferSync() calls so we can determine the
+ * strategy point's advance rate and avoid scanning already-cleaned buffers. The
+ * variables are global instead of static local so that BgBufferSyncReset() can
+ * adjust it when resizing shared buffers.
+ */
+static bool saved_info_valid = false;
+static int prev_strategy_buf_id;
+static uint32 prev_strategy_passes;
+static int next_to_clean;
+static uint32 next_passes;
+
+/* Moving averages of allocation rate and clean-buffer density */
+static float smoothed_alloc = 0;
+static float smoothed_density = 10.0;
+
+void
+BgBufferSyncReset(int NBuffersOld, int NBuffersNew)
+{
+ saved_info_valid = false;
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
+ NBuffersOld, NBuffersNew);
+#endif
+}
+
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
@@ -3599,20 +3625,6 @@ BgBufferSync(WritebackContext *wb_context)
uint32 strategy_passes;
uint32 recent_alloc;
- /*
- * Information saved between calls so we can determine the strategy
- * point's advance rate and avoid scanning already-cleaned buffers.
- */
- static bool saved_info_valid = false;
- static int prev_strategy_buf_id;
- static uint32 prev_strategy_passes;
- static int next_to_clean;
- static uint32 next_passes;
-
- /* Moving averages of allocation rate and clean-buffer density */
- static float smoothed_alloc = 0;
- static float smoothed_density = 10.0;
-
/* Potentially these could be tunables, but for now, not */
float smoothing_samples = 16;
float scan_whole_pool_milliseconds = 120000.0;
@@ -3635,6 +3647,22 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ /*
+ * If buffer pool is being shrunk the buffer being written out may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing out any. Hence wait till
+ * buffer resizing finishes i.e. go into hibernation mode.
+ */
+ if (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ return true;
+
+ /*
+ * Resizing shared buffers while this function is performing an LRU scan on
+ * them may lead to wrong results. Indicate that the resizing should wait for
+ * the LRU scan to complete.
+ */
+ delay_shmem_resize = true;
+
/*
* Find out where the clock-sweep currently is, and how many buffer
* allocations have happened since our last call.
@@ -3811,8 +3839,17 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
- while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
+ /*
+ * Execute the LRU scan.
+ *
+ * If buffer pool is being shrunk, the buffer being written may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing any. Hence stop what we are doing. This
+ * also unblocks other processes that are waiting for buffer resizing to
+ * finish.
+ */
+ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
+ pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) == NBuffers)
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -3871,6 +3908,9 @@ BgBufferSync(WritebackContext *wb_context)
#endif
}
+ /* Let the resizing commence. */
+ delay_shmem_resize = false;
+
/* Return true if OK to hibernate */
return (bufs_to_lap == 0 && recent_alloc == 0);
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index db8aafdaf8c..89269087034 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -371,12 +371,21 @@ StrategyInitialize(bool init)
*
* Since we can't tolerate running out of lookup table entries, we must be
* sure to specify an adequate table size here. The maximum steady-state
- * usage is of course NBuffers entries, but BufferAlloc() tries to insert
- * a new entry before deleting the old. In principle this could be
- * happening in each partition concurrently, so we could need as many as
- * NBuffers + NUM_BUFFER_PARTITIONS entries.
+ * usage is of course is as many number of entries as the number of buffers
+ * in the buffer pool. Right now there is no way to free shared memory. Even
+ * if we shrink the buffer lookup table when shrinking the buffer pool the
+ * unused hash table entries can not be freed. When we expand the buffer
+ * pool, more entries can be allocated but we can not resize the hash table
+ * directory without rehashing all the entries. Just allocating more entries
+ * will lead to more contention. Hence we setup the buffer lookup table
+ * considering the maximum possible size of the buffer pool which is
+ * MaxAvailableMemory.
+ *
+ * Additionally BufferAlloc() tries to insert a new entry before deleting the
+ * old. In principle this could be happening in each partition concurrently,
+ * so we need extra NUM_BUFFER_PARTITIONS entries.
*/
- InitBufTable(NBuffers + NUM_BUFFER_PARTITIONS);
+ InitBufTable(MaxAvailableMemory + NUM_BUFFER_PARTITIONS);
/*
* Get or create the shared strategy control block
@@ -384,7 +393,7 @@ StrategyInitialize(bool init)
StrategyControl = (BufferStrategyControl *)
ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found, STRATEGY_SHMEM_SEGMENT);
+ &found, MAIN_SHMEM_SEGMENT);
if (!found)
{
@@ -409,6 +418,62 @@ StrategyInitialize(bool init)
Assert(!init);
}
+/*
+ * StrategyReInitialize -- re-initialize the buffer cache replacement
+ * strategy.
+ *
+ * To be called when resizing buffer manager and only from the coordinator.
+ * TODO: Assess the differences between this function and StrategyInitialize().
+ */
+void
+StrategyReInitialize(int FirstBufferIdToInit)
+{
+ bool found;
+
+ /*
+ * Resizing memory for buffer pools should not affect the address of
+ * StrategyControl.
+ */
+ if (StrategyControl != (BufferStrategyControl *)
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, MAIN_SHMEM_SEGMENT))
+ elog(FATAL, "something went wrong while re-initializing the buffer strategy");
+
+ Assert(found);
+
+ /* TODO: Buffer lookup table adjustment: There are two options:
+ *
+ * 1. Resize the buffer lookup table to match the new number of buffers. But
+ * this requires rehashing all the entries in the buffer lookup table with
+ * the new table size.
+ *
+ * 2. Allocate maximum size of the buffer lookup table at the beginning and
+ * never resize it. This leaves sparse buffer lookup table which is
+ * inefficient from both memory and time perspective. According to David
+ * Rowley, the sparse entries in the buffer look up table cause frequent
+ * cacheline reload which affect performance. If the impact of that
+ * inefficiency in a benchmark is significant, we will need to consider first
+ * option.
+ */
+ /*
+ * The clock sweep tick pointer might have got invalidated. Reset it as if
+ * starting a fresh server.
+ */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The old statistics is viewed in the context of the number of shared
+ * buffers. It does not make sense now that the number of shared buffers
+ * itself has changed.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_init_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* No pending notification */
+ StrategyControl->bgwprocno = -1;
+}
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index dfd614f7ca4..551479649ca 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -443,6 +443,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyReInitialize(int FirstBufferToInit);
/* buf_table.c */
extern Size BufTableShmemSize(int size);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index c91a42fc598..2fe3202168b 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -299,6 +299,7 @@ extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
extern bool BgBufferSync(WritebackContext *wb_context);
+extern void BgBufferSyncReset(int NBuffersOld, int NBuffersNew);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 847f56a36dc..6e7b0abb625 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -65,6 +65,7 @@ typedef void (*shmem_startup_hook_type) (void);
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
+extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 0a59746b472..704b065f9e9 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -65,7 +65,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 6
+#define ANON_MAPPINGS 5
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -172,7 +172,4 @@ extern void ShmemControlInit(void);
/* Checkpoint BufferIds */
#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
-/* Buffer strategy status */
-#define STRATEGY_SHMEM_SEGMENT 5
-
#endif /* PG_SHMEM_H */
--
2.34.1
[application/x-patch] 0016-Tests-for-dynamic-shared_buffers-resizing-20250918.patch (19.6K, ../../CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k+krxht1yFjA@mail.gmail.com/17-0016-Tests-for-dynamic-shared_buffers-resizing-20250918.patch)
download | inline diff:
From 359800f4e40a4ac97cb5aec1198bd0b8584ce53e Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 3 Sep 2025 10:59:20 +0530
Subject: [PATCH 16/16] Tests for dynamic shared_buffers resizing
The commit adds two tests:
1. TAP test to stress test buffer pool resizing under concurrent load.
2. SQL test to test sanity of shared memory allocations and mappings
after buffer pool resizing operation.
Author: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
src/test/Makefile | 2 +-
src/test/README | 3 +
src/test/buffermgr/Makefile | 27 ++
src/test/buffermgr/README | 26 ++
src/test/buffermgr/expected/buffer_resize.out | 237 ++++++++++++++++++
src/test/buffermgr/meson.build | 17 ++
src/test/buffermgr/sql/buffer_resize.sql | 73 ++++++
src/test/buffermgr/t/001_resize_buffer.pl | 126 ++++++++++
src/test/meson.build | 1 +
9 files changed, 511 insertions(+), 1 deletion(-)
create mode 100644 src/test/buffermgr/Makefile
create mode 100644 src/test/buffermgr/README
create mode 100644 src/test/buffermgr/expected/buffer_resize.out
create mode 100644 src/test/buffermgr/meson.build
create mode 100644 src/test/buffermgr/sql/buffer_resize.sql
create mode 100644 src/test/buffermgr/t/001_resize_buffer.pl
diff --git a/src/test/Makefile b/src/test/Makefile
index 511a72e6238..95f8858a818 100644
--- a/src/test/Makefile
+++ b/src/test/Makefile
@@ -12,7 +12,7 @@ subdir = src/test
top_builddir = ../..
include $(top_builddir)/src/Makefile.global
-SUBDIRS = perl postmaster regress isolation modules authentication recovery subscription
+SUBDIRS = perl postmaster regress isolation modules authentication recovery subscription buffermgr
ifeq ($(with_icu),yes)
SUBDIRS += icu
diff --git a/src/test/README b/src/test/README
index afdc7676519..77f11607ff7 100644
--- a/src/test/README
+++ b/src/test/README
@@ -15,6 +15,9 @@ examples/
Demonstration programs for libpq that double as regression tests via
"make check"
+buffermgr/
+ Tests for resizing buffer pool without restarting the server
+
isolation/
Tests for concurrent behavior at the SQL level
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
new file mode 100644
index 00000000000..97c3da9e20a
--- /dev/null
+++ b/src/test/buffermgr/Makefile
@@ -0,0 +1,27 @@
+#-------------------------------------------------------------------------
+#
+# Makefile for src/test/buffermgr
+#
+# Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+# Portions Copyright (c) 1994, Regents of the University of California
+#
+# src/test/buffermgr/Makefile
+#
+#-------------------------------------------------------------------------
+
+EXTRA_INSTALL = contrib/pg_buffercache
+
+REGRESS = buffer_resize
+
+subdir = src/test/buffermgr
+top_builddir = ../../..
+include $(top_builddir)/src/Makefile.global
+
+check:
+ $(prove_check)
+
+installcheck:
+ $(prove_installcheck)
+
+clean distclean:
+ rm -rf tmp_check
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
new file mode 100644
index 00000000000..c375ad80989
--- /dev/null
+++ b/src/test/buffermgr/README
@@ -0,0 +1,26 @@
+src/test/buffermgr/README
+
+Regression tests for buffer manager
+===================================
+
+This directory contains a test suite for resizing buffer manager without restarting the server.
+
+
+Running the tests
+=================
+
+NOTE: You must have given the --enable-tap-tests argument to configure.
+
+Run
+ make check
+or
+ make installcheck
+You can use "make installcheck" if you previously did "make install".
+In that case, the code in the installation tree is tested. With
+"make check", a temporary installation tree is built from the current
+sources and then tested.
+
+Either way, this test initializes, starts, and stops a test Postgres
+cluster.
+
+See src/test/perl/README for more info about running these tests.
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
new file mode 100644
index 00000000000..a986be9a5da
--- /dev/null
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -0,0 +1,237 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+-- Create a separate schema for this test
+CREATE SCHEMA buffer_resize_test;
+SET search_path TO buffer_resize_test, public;
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, mapping_size, mapping_reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | descriptors | 1048576 | 1048576
+ Buffer IO Condition Variables | iocv | 262144 | 262144
+ Checkpoint BufferIds | checkpoint | 327680 | 327680
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 134225920 | 134225920 | 2576982016
+ checkpoint | 335872 | 335872 | 214753280
+ descriptors | 1056768 | 1056768 | 429498368
+ iocv | 270336 | 270336 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+SELECT pg_sleep(1);
+ pg_sleep
+----------
+
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 64MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+----------+----------------
+ Buffer Blocks | buffers | 67112960 | 67112960
+ Buffer Descriptors | descriptors | 524288 | 524288
+ Buffer IO Condition Variables | iocv | 131072 | 131072
+ Checkpoint BufferIds | checkpoint | 163840 | 163840
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+----------+--------------+-----------------------
+ buffers | 67117056 | 67117056 | 2576982016
+ checkpoint | 172032 | 172032 | 214753280
+ descriptors | 532480 | 532480 | 429498368
+ iocv | 139264 | 139264 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 8192
+(1 row)
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+SELECT pg_sleep(1);
+ pg_sleep
+----------
+
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 268439552 | 268439552
+ Buffer Descriptors | descriptors | 2097152 | 2097152
+ Buffer IO Condition Variables | iocv | 524288 | 524288
+ Checkpoint BufferIds | checkpoint | 655360 | 655360
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 268443648 | 268443648 | 2576982016
+ checkpoint | 663552 | 663552 | 214753280
+ descriptors | 2105344 | 2105344 | 429498368
+ iocv | 532480 | 532480 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 32768
+(1 row)
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+SELECT pg_sleep(1);
+ pg_sleep
+----------
+
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 100MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 104861696 | 104861696
+ Buffer Descriptors | descriptors | 819200 | 819200
+ Buffer IO Condition Variables | iocv | 204800 | 204800
+ Checkpoint BufferIds | checkpoint | 256000 | 256000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 104865792 | 104865792 | 2576982016
+ checkpoint | 262144 | 262144 | 214753280
+ descriptors | 827392 | 827392 | 429498368
+ iocv | 212992 | 212992 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 12800
+(1 row)
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+SELECT pg_sleep(1);
+ pg_sleep
+----------
+
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+--------+----------------
+ Buffer Blocks | buffers | 135168 | 135168
+ Buffer Descriptors | descriptors | 1024 | 1024
+ Buffer IO Condition Variables | iocv | 256 | 256
+ Checkpoint BufferIds | checkpoint | 320 | 320
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+--------+--------------+-----------------------
+ buffers | 139264 | 139264 | 2576982016
+ checkpoint | 8192 | 8192 | 214753280
+ descriptors | 8192 | 8192 | 429498368
+ iocv | 8192 | 8192 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16
+(1 row)
+
+-- Clean up the schema and all its objects
+RESET search_path;
+DROP SCHEMA buffer_resize_test CASCADE;
+NOTICE: drop cascades to 3 other objects
+DETAIL: drop cascades to view buffer_resize_test.buffer_allocations
+drop cascades to view buffer_resize_test.buffer_segments
+drop cascades to extension pg_buffercache
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
new file mode 100644
index 00000000000..e71dcdea685
--- /dev/null
+++ b/src/test/buffermgr/meson.build
@@ -0,0 +1,17 @@
+# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+
+tests += {
+ 'name': 'buffermgr',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'buffer_resize',
+ ],
+ },
+ 'tap': {
+ 'tests': [
+ 't/001_resize_buffer.pl',
+ ],
+ },
+}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
new file mode 100644
index 00000000000..45f5bb6d78b
--- /dev/null
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -0,0 +1,73 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+
+-- Create a separate schema for this test
+CREATE SCHEMA buffer_resize_test;
+SET search_path TO buffer_resize_test, public;
+
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, mapping_size, mapping_reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+SELECT pg_sleep(1);
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+SELECT pg_sleep(1);
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+SELECT pg_sleep(1);
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+SELECT pg_sleep(1);
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Clean up the schema and all its objects
+RESET search_path;
+DROP SCHEMA buffer_resize_test CASCADE;
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
new file mode 100644
index 00000000000..8cf9e4539ab
--- /dev/null
+++ b/src/test/buffermgr/t/001_resize_buffer.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Minimal test testing shared_buffer resizing under load
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Function to resize buffer pool and verify the change.
+sub apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ # Use a single background_psql session for consistency
+ my $psql_session = $node->background_psql('postgres');
+ $psql_session->query_safe("ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $psql_session->query_safe("SELECT pg_reload_conf()");
+
+ # Wait till the resizing finishes using the same session
+ #
+ # TODO: Right now there is no way to know when the resize has finished and
+ # all the backends are using new value of shared_buffers. Hence we poll
+ # manually until we get the expected value in the same session.
+ my $current_size;
+ my $attempts = 0;
+ my $max_attempts = 60; # 60 seconds timeout
+ do {
+ $current_size = $psql_session->query_safe("SHOW shared_buffers");
+ $attempts++;
+
+ # Only sleep if we didn't get the expected result and haven't timed out yet
+ if ($current_size ne $new_size && $attempts < $max_attempts) {
+ sleep(1);
+ }
+ } while ($current_size ne $new_size && $attempts < $max_attempts);
+
+ $psql_session->quit;
+
+ # Check if we succeeded or timed out
+ if ($current_size ne $new_size) {
+ die "Timeout waiting for shared_buffers to change to $new_size (got $current_size after ${attempts}s)";
+ }
+}
+
+# Initialize a cluster and start pgbench in the background for concurrent load.
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->start;
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+my $pgb_scale = 10;
+my $pgb_duration = 120;
+my $pgb_num_clients = 10;
+$node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$pgb_scale --quiet",
+ 0,
+ [qr{^$}],
+ [ # stderr patterns to verify initialization stages
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$pgb_scale)"
+);
+my ($pgbench_stdin, $pgbench_stdout, $pgbench_stderr) = ('', '', '');
+my $pgbench_process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-T', $pgb_duration,
+ '-c', $pgb_num_clients,
+ 'postgres'
+ ],
+ '<' => \$pgbench_stdin,
+ '>' => \$pgbench_stdout,
+ '2>' => \$pgbench_stderr
+);
+
+ok($pgbench_process, "pgbench started successfully");
+
+# Allow pgbench to establish connections and start generating load.
+#
+# TODO: When creating new backends is known to work well with buffer pool
+# resizing, this wait should be removed.
+sleep(1);
+
+# Resize buffer pool to various sizes while pgbench is running in the
+# background.
+#
+# TODO: These are pseudo-randomly picked sizes, but we can do better.
+my $tests_completed = 0;
+my @buffer_sizes = ('900MB', '500MB', '250MB', '400MB', '120MB', '600MB');
+for my $target_size (@buffer_sizes)
+{
+ # Verify workload generator is still running
+ if (!$pgbench_process->pumpable) {
+ ok(0, "pgbench is still running");
+ last;
+ }
+
+ apply_and_verify_buffer_change($node, $target_size);
+ $tests_completed++;
+
+ # Wait for the resized buffer pool to stabilize. If the resized buffer pool
+ # is utilized fully, it might hit any wrongly initialized areas of shared
+ # memory.
+ sleep(2);
+}
+is($tests_completed, scalar(@buffer_sizes), "All buffer sizes were tested");
+
+# Make sure that pgbench can end normally.
+$pgbench_process->signal('TERM');
+IPC::Run::finish $pgbench_process;
+ok(grep { $pgbench_process->result == $_ } (0, 15), "pgbench exited gracefully");
+
+# Log any error output from pgbench for debugging
+diag("pgbench stderr:\n$pgbench_stderr");
+diag("pgbench stdout:\n$pgbench_stdout");
+
+# Ensure database is still functional after all the buffer changes
+$node->connect_ok("dbname=postgres",
+ "Database remains accessible after $tests_completed buffer resize operations");
+
+done_testing();
\ No newline at end of file
diff --git a/src/test/meson.build b/src/test/meson.build
index ccc31d6a86a..2a5ba1dec39 100644
--- a/src/test/meson.build
+++ b/src/test/meson.build
@@ -4,6 +4,7 @@ subdir('regress')
subdir('isolation')
subdir('authentication')
+subdir('buffermgr')
subdir('postmaster')
subdir('recovery')
subdir('subscription')
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-09-18 13:52 ` Andres Freund <andres@anarazel.de>
2025-09-18 14:05 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 2 replies; 167+ messages in thread
From: Andres Freund @ 2025-09-18 13:52 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Hi,
On 2025-09-18 10:25:29 +0530, Ashutosh Bapat wrote:
> From d1ed934ccd02fca2c831e582b07a169e17d19f59 Mon Sep 17 00:00:00 2001
> From: Dmitrii Dolgov <9erthalion6@gmail.com>
> Date: Tue, 17 Jun 2025 15:14:33 +0200
> Subject: [PATCH 02/16] Process config reload in AIO workers
I think this is superfluous due to b8e1f2d96bb9
> Currenly AIO workers process interrupts only via CHECK_FOR_INTERRUPTS,
> which does not include ConfigReloadPending. Thus we need to check for it
> explicitly.
> +/*
> + * Process any new interrupts.
> + */
> +static void
> +pgaio_worker_process_interrupts(void)
> +{
> + /*
> + * Reloading config can trigger further signals, complicating interrupts
> + * processing -- so let it run first.
> + *
> + * XXX: Is there any need in memory barrier after ProcessConfigFile?
> + */
> + if (ConfigReloadPending)
> + {
> + ConfigReloadPending = false;
> + ProcessConfigFile(PGC_SIGHUP);
> + }
> +
> + if (ProcSignalBarrierPending)
> + ProcessProcSignalBarrier();
> +}
Given that even before b8e1f2d96bb9 method_worker.c used
CHECK_FOR_INTERRUPTS(), which contains a ProcessProcSignalBarrier(), I don't
know why that second check was added here?
> From 0a13e56dceea8cc7a2685df7ee8cea434588681b Mon Sep 17 00:00:00 2001
> From: Dmitrii Dolgov <9erthalion6@gmail.com>
> Date: Sun, 6 Apr 2025 16:40:32 +0200
> Subject: [PATCH 03/16] Introduce pending flag for GUC assign hooks
>
> Currently an assing hook can perform some preprocessing of a new value,
> but it cannot change the behavior, which dictates that the new value
> will be applied immediately after the hook. Certain GUC options (like
> shared_buffers, coming in subsequent patches) may need coordinating work
> between backends to change, meaning we cannot apply it right away.
>
> Add a new flag "pending" for an assign hook to allow the hook indicate
> exactly that. If the pending flag is set after the hook, the new value
> will not be applied and it's handling becomes the hook's implementation
> responsibility.
I doubt it makes sense to add this to the GUC system. I think it'd be better
to just use the GUC value as the desired "target" configuration and have a
function or a show-only GUC for reporting the current size.
I don't think you can't just block application of the GUC until the resize is
complete. E.g. what if the value was too big and the new configuration needs
to fixed to be lower?
> From 0a55bc15dc3a724f03e674048109dac1f248c406 Mon Sep 17 00:00:00 2001
> From: Dmitrii Dolgov <9erthalion6@gmail.com>
> Date: Fri, 4 Apr 2025 21:46:14 +0200
> Subject: [PATCH 04/16] Introduce pss_barrierReceivedGeneration
>
> Currently WaitForProcSignalBarrier allows to make sure the message sent
> via EmitProcSignalBarrier was processed by all ProcSignal mechanism
> participants.
>
> Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
> which will be updated when a process has received the message, but not
> processed it yet. This makes it possible to support a new mode of
> waiting, when ProcSignal participants want to synchronize message
> processing. To do that, a participant can wait via
> WaitForProcSignalBarrierReceived when processing a message, effectively
> making sure that all processes are going to start processing
> ProcSignalBarrier simultaneously.
I doubt "online resizing" that requires synchronously processing the same
event, can really be called "online". There can be significant delays in
processing a barrier, stalling the entire server until that is reached seems
like a complete no-go for production systems?
> From 63fe27340656c52b13f4eecebd9e73d24efe5e33 Mon Sep 17 00:00:00 2001
> From: Dmitrii Dolgov <9erthalion6@gmail.com>
> Date: Fri, 28 Feb 2025 19:54:47 +0100
> Subject: [PATCH 05/16] Allow to use multiple shared memory mappings
>
> Currently all the work with shared memory is done via a single anonymous
> memory mapping, which limits ways how the shared memory could be organized.
>
> Introduce possibility to allocate multiple shared memory mappings, where
> a single mapping is associated with a specified shared memory segment.
> There is only fixed amount of available segments, currently only one
> main shared memory segment is allocated. A new shared memory API is
> introduces, extended with a segment as a new parameter. As a path of
> least resistance, the original API is kept in place, utilizing the main
> shared memory segment.
> -#define MAX_ON_EXITS 20
> +#define MAX_ON_EXITS 40
Why does a patch like this contain changes like this mixed in with the rest?
That's clearly not directly related to $subject.
> /* shared memory global variables */
>
> -static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
> +ShmemSegment Segments[ANON_MAPPINGS];
>
> -static void *ShmemBase; /* start address of shared memory */
> -
> -static void *ShmemEnd; /* end+1 address of shared memory */
> -
> -slock_t *ShmemLock; /* spinlock for shared memory and LWLock
> - * allocation */
> -
> -static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
Why do we need a separate ShmemLock for each segment? Besides being
unnecessary, it seems like that prevents locking in a way that provides
consistency across all segments.
> From e2f48da8a8206711b24e34040d699431910fbf9c Mon Sep 17 00:00:00 2001
> From: Dmitrii Dolgov <9erthalion6@gmail.com>
> Date: Tue, 17 Jun 2025 11:47:04 +0200
> Subject: [PATCH 06/16] Address space reservation for shared memory
>
> Currently the shared memory layout is designed to pack everything tight
> together, leaving no space between mappings for resizing. Here is how it
> looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
> anonymous shared memory we talk about:
>
> 00400000-00490000 /path/bin/postgres
> ...
> 012d9000-0133e000 [heap]
> 7f443a800000-7f470a800000 /dev/zero (deleted)
> 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
> 7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
> ...
>
> Make the layout more dynamic via splitting every shared memory segment
> into two parts:
>
> * An anonymous file, which actually contains shared memory content. Such
> an anonymous file is created via memfd_create, it lives in memory,
> behaves like a regular file and semantically equivalent to an
> anonymous memory allocated via mmap with MAP_ANONYMOUS.
>
> * A reservation mapping, which size is much larger than required shared
> segment size. This mapping is created with flags PROT_NONE (which
> makes sure the reserved space is not used), and MAP_NORESERVE (to not
> count the reserved space against memory limits). The anonymous file is
> mapped into this reservation mapping.
The commit message fails to explain why, if we're already relying on
MAP_NORESERVE, we need to anything else? Why can't we just have one maximally
sized allocation that's marked MAP_NORESERVE for all the parts that we don't
yet need?
> There are also few unrelated advantages of using anon files:
>
> * We've got a file descriptor, which could be used for regular file
> operations (modification, truncation, you name it).
What is this an advantage for?
> * The file could be given a name, which improves readability when it
> comes to process maps.
> * By default, Linux will not add file-backed shared mappings into a core dump,
> making it more convenient to work with them in PostgreSQL: no more huge dumps
> to process.
That's just as well a downside, because you now can't investigate some
issues. This was already configurable via coredump_filter.
> From 942b69a0876b0e83303e6704da54c4c002a5a2d8 Mon Sep 17 00:00:00 2001
> From: Dmitrii Dolgov <9erthalion6@gmail.com>
> Date: Tue, 17 Jun 2025 11:22:02 +0200
> Subject: [PATCH 07/16] Introduce multiple shmem segments for shared buffers
>
> Add more shmem segments to split shared buffers into following chunks:
> * BUFFERS_SHMEM_SEGMENT: contains buffer blocks
> * BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
> * BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
> * CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
> * STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Why do all these need to be separate segments? Afaict we'll have to maximally
size everything other than BUFFERS_SHMEM_SEGMENT at start?
> From 78bc0a49f8ebe17927abd66164764745ecc6d563 Mon Sep 17 00:00:00 2001
> From: Dmitrii Dolgov <9erthalion6@gmail.com>
> Date: Tue, 17 Jun 2025 14:16:55 +0200
> Subject: [PATCH 11/16] Allow to resize shared memory without restart
>
> Add assing hook for shared_buffers to resize shared memory using space,
> introduced in the previous commits without requiring PostgreSQL restart.
> Essentially the implementation is based on two mechanisms: a
> ProcSignalBarrier is used to make sure all processes are starting the
> resize procedure simultaneously, and a global Barrier is used to
> coordinate after that and make sure all finished processes are waiting
> for others that are in progress.
>
> The resize process looks like this:
>
> * The GUC assign hook sets a flag to let the Postmaster know that resize
> was requested.
>
> * Postmaster verifies the flag in the event loop, and starts the resize
> by emitting a ProcSignal barrier.
>
> * All processes, that participate in ProcSignal mechanism, begin to
> process ProcSignal barrier. First a process waits until all processes
> have confirmed they received the message and can start simultaneously.
As mentioned above, this basically makes the entire feature not really
online. Besides the latency of some processes not getting to the barrier
immediately, there's also the issue that actually reserving large amounts of
memory can take a long time - during which all processes would be unavailable.
I really don't see that being viable. It'd be one thing if that were a
"temporary" restriction, but the whole design seems to be fairly centered
around that.
> * Every process recalculates shared memory size based on the new
> NBuffers, adjusts its size using ftruncate and adjust reservation
> permissions with mprotect. One elected process signals the postmaster
> to do the same.
If we just used a single memory mapping with all unused parts marked
MAP_NORESERVE, we wouldn't need this (and wouldn't need a fair bit of other
work in this patchset)..
> From experiment it turns out that shared mappings have to be extended
> separately for each process that uses them. Another rough edge is that a
> backend blocked on ReadCommand will not apply shared_buffers change
> until it receives something.
That's not a rough edge, that basically makes the feature unusable, no?
> +-- Test 2: Set to 64MB
> +ALTER SYSTEM SET shared_buffers = '64MB';
> +SELECT pg_reload_conf();
> +SELECT pg_sleep(1);
> +SHOW shared_buffers;
Tests containing sleeps are a significant warning flag imo.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-09-18 14:05 ` Andres Freund <andres@anarazel.de>
2025-09-26 18:04 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Andres Freund @ 2025-09-18 14:05 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Hi,
On 2025-09-18 09:52:03 -0400, Andres Freund wrote:
> On 2025-09-18 10:25:29 +0530, Ashutosh Bapat wrote:
> > From 0a55bc15dc3a724f03e674048109dac1f248c406 Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Fri, 4 Apr 2025 21:46:14 +0200
> > Subject: [PATCH 04/16] Introduce pss_barrierReceivedGeneration
> >
> > Currently WaitForProcSignalBarrier allows to make sure the message sent
> > via EmitProcSignalBarrier was processed by all ProcSignal mechanism
> > participants.
> >
> > Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
> > which will be updated when a process has received the message, but not
> > processed it yet. This makes it possible to support a new mode of
> > waiting, when ProcSignal participants want to synchronize message
> > processing. To do that, a participant can wait via
> > WaitForProcSignalBarrierReceived when processing a message, effectively
> > making sure that all processes are going to start processing
> > ProcSignalBarrier simultaneously.
>
> I doubt "online resizing" that requires synchronously processing the same
> event, can really be called "online". There can be significant delays in
> processing a barrier, stalling the entire server until that is reached seems
> like a complete no-go for production systems?
> [...]
> > From 78bc0a49f8ebe17927abd66164764745ecc6d563 Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Tue, 17 Jun 2025 14:16:55 +0200
> > Subject: [PATCH 11/16] Allow to resize shared memory without restart
> >
> > Add assing hook for shared_buffers to resize shared memory using space,
> > introduced in the previous commits without requiring PostgreSQL restart.
> > Essentially the implementation is based on two mechanisms: a
> > ProcSignalBarrier is used to make sure all processes are starting the
> > resize procedure simultaneously, and a global Barrier is used to
> > coordinate after that and make sure all finished processes are waiting
> > for others that are in progress.
> >
> > The resize process looks like this:
> >
> > * The GUC assign hook sets a flag to let the Postmaster know that resize
> > was requested.
> >
> > * Postmaster verifies the flag in the event loop, and starts the resize
> > by emitting a ProcSignal barrier.
> >
> > * All processes, that participate in ProcSignal mechanism, begin to
> > process ProcSignal barrier. First a process waits until all processes
> > have confirmed they received the message and can start simultaneously.
>
> As mentioned above, this basically makes the entire feature not really
> online. Besides the latency of some processes not getting to the barrier
> immediately, there's also the issue that actually reserving large amounts of
> memory can take a long time - during which all processes would be unavailable.
>
> I really don't see that being viable. It'd be one thing if that were a
> "temporary" restriction, but the whole design seems to be fairly centered
> around that.
Besides not really being online, isn't this a recipe for endless undetected
deadlocks? What if process A waits for a lock held by process B and process B
arrives at the barrier? Process A won't ever get there, because process B
can't make progress, because A is not making progress.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-09-18 14:05 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-09-26 18:04 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-26 18:36 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-09-29 06:57 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-09-26 18:04 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Sorry for late reply folks.
> On Thu, Sep 18, 2025 at 09:52:03AM -0400, Andres Freund wrote:
> > From 0a13e56dceea8cc7a2685df7ee8cea434588681b Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Sun, 6 Apr 2025 16:40:32 +0200
> > Subject: [PATCH 03/16] Introduce pending flag for GUC assign hooks
> >
> > Currently an assing hook can perform some preprocessing of a new value,
> > but it cannot change the behavior, which dictates that the new value
> > will be applied immediately after the hook. Certain GUC options (like
> > shared_buffers, coming in subsequent patches) may need coordinating work
> > between backends to change, meaning we cannot apply it right away.
> >
> > Add a new flag "pending" for an assign hook to allow the hook indicate
> > exactly that. If the pending flag is set after the hook, the new value
> > will not be applied and it's handling becomes the hook's implementation
> > responsibility.
>
> I doubt it makes sense to add this to the GUC system. I think it'd be better
> to just use the GUC value as the desired "target" configuration and have a
> function or a show-only GUC for reporting the current size.
>
> I don't think you can't just block application of the GUC until the resize is
> complete. E.g. what if the value was too big and the new configuration needs
> to fixed to be lower?
I think it was a bit hasty to post another version of the patch without
the design changes we've agreed upon last time. I'm still working on
that (sorry, it takes time, I haven't wrote so much Perl for testing
since forever), the current implementation doesn't include anything with
GUC to simplify the discussion. I'm still convinced that multi-step GUC
changing makes sense, but it has proven to be more complicated than I
anticipated, so I'll spin up another thread to discuss when I come to
it.
> > From 0a55bc15dc3a724f03e674048109dac1f248c406 Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Fri, 4 Apr 2025 21:46:14 +0200
> > Subject: [PATCH 04/16] Introduce pss_barrierReceivedGeneration
> >
> > Currently WaitForProcSignalBarrier allows to make sure the message sent
> > via EmitProcSignalBarrier was processed by all ProcSignal mechanism
> > participants.
> >
> > Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
> > which will be updated when a process has received the message, but not
> > processed it yet. This makes it possible to support a new mode of
> > waiting, when ProcSignal participants want to synchronize message
> > processing. To do that, a participant can wait via
> > WaitForProcSignalBarrierReceived when processing a message, effectively
> > making sure that all processes are going to start processing
> > ProcSignalBarrier simultaneously.
>
> I doubt "online resizing" that requires synchronously processing the same
> event, can really be called "online". There can be significant delays in
> processing a barrier, stalling the entire server until that is reached seems
> like a complete no-go for production systems?
>
> [...]
> As mentioned above, this basically makes the entire feature not really
> online. Besides the latency of some processes not getting to the barrier
> immediately, there's also the issue that actually reserving large amounts of
> memory can take a long time - during which all processes would be unavailable.
>
> I really don't see that being viable. It'd be one thing if that were a
> "temporary" restriction, but the whole design seems to be fairly centered
> around that.
>
> [...]
>
> Besides not really being online, isn't this a recipe for endless undetected
> deadlocks? What if process A waits for a lock held by process B and process B
> arrives at the barrier? Process A won't ever get there, because process B
> can't make progress, because A is not making progress.
Same as above, in the version I'm working right now it's changed in
favor of an approach that looks more like the one from "online checksum
change" patch. I've even stumbled upon a cases when a process was just
killed and never arrive at the barrier, so that was it. The new approach
makes certain parts simpler, but requires managing backends with
different understanding of how large shared memory segments are for some
time interval. Introducing a new parameter "number of available buffers"
seems to be helpful to address all cases I've found so far.
Btw, under "online" resizing I mostly understood "without restart", the
goal was not to make it really "online".
> > -#define MAX_ON_EXITS 20
> > +#define MAX_ON_EXITS 40
>
> Why does a patch like this contain changes like this mixed in with the rest?
> That's clearly not directly related to $subject.
An artifact of rebasing, it belonged to 0007.
> > From e2f48da8a8206711b24e34040d699431910fbf9c Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Tue, 17 Jun 2025 11:47:04 +0200
> > Subject: [PATCH 06/16] Address space reservation for shared memory
> >
> > Currently the shared memory layout is designed to pack everything tight
> > together, leaving no space between mappings for resizing. Here is how it
> > looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
> > anonymous shared memory we talk about:
> >
> > 00400000-00490000 /path/bin/postgres
> > ...
> > 012d9000-0133e000 [heap]
> > 7f443a800000-7f470a800000 /dev/zero (deleted)
> > 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
> > 7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
> > ...
> >
> > Make the layout more dynamic via splitting every shared memory segment
> > into two parts:
> >
> > * An anonymous file, which actually contains shared memory content. Such
> > an anonymous file is created via memfd_create, it lives in memory,
> > behaves like a regular file and semantically equivalent to an
> > anonymous memory allocated via mmap with MAP_ANONYMOUS.
> >
> > * A reservation mapping, which size is much larger than required shared
> > segment size. This mapping is created with flags PROT_NONE (which
> > makes sure the reserved space is not used), and MAP_NORESERVE (to not
> > count the reserved space against memory limits). The anonymous file is
> > mapped into this reservation mapping.
>
> The commit message fails to explain why, if we're already relying on
> MAP_NORESERVE, we need to anything else? Why can't we just have one maximally
> sized allocation that's marked MAP_NORESERVE for all the parts that we don't
> yet need?
How do we return memory to the OS in that case? Currently it's done
explicitly via truncating the anonymous file.
> > * The file could be given a name, which improves readability when it
> > comes to process maps.
>
> > * By default, Linux will not add file-backed shared mappings into a core dump,
> > making it more convenient to work with them in PostgreSQL: no more huge dumps
> > to process.
>
> That's just as well a downside, because you now can't investigate some
> issues. This was already configurable via coredump_filter.
This behaviour is configured via coredump_filter as well, so just the
default value has been changed.
> > From 942b69a0876b0e83303e6704da54c4c002a5a2d8 Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Tue, 17 Jun 2025 11:22:02 +0200
> > Subject: [PATCH 07/16] Introduce multiple shmem segments for shared buffers
> >
> > Add more shmem segments to split shared buffers into following chunks:
> > * BUFFERS_SHMEM_SEGMENT: contains buffer blocks
> > * BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
> > * BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
> > * CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
> > * STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
>
> Why do all these need to be separate segments? Afaict we'll have to maximally
> size everything other than BUFFERS_SHMEM_SEGMENT at start?
Why would they need to me maxed out at the start? So far my rule of
thumb was one segment for one structure which size depends on NBuffers,
so that when changing NBuffers each segment could be adjusted
independently.
> > +-- Test 2: Set to 64MB
> > +ALTER SYSTEM SET shared_buffers = '64MB';
> > +SELECT pg_reload_conf();
> > +SELECT pg_sleep(1);
> > +SHOW shared_buffers;
>
> Tests containing sleeps are a significant warning flag imo.
Tests I'm preparing so far avoiding this by waiting in injection points.
I haven't found anything similar in existing tests, but I assume such
approach is fine.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-09-18 14:05 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-09-26 18:04 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-09-26 18:36 ` Andres Freund <andres@anarazel.de>
1 sibling, 0 replies; 167+ messages in thread
From: Andres Freund @ 2025-09-26 18:36 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Hi,
On 2025-09-26 20:04:21 +0200, Dmitry Dolgov wrote:
> > On Thu, Sep 18, 2025 at 09:52:03AM -0400, Andres Freund wrote:
> > > From 0a13e56dceea8cc7a2685df7ee8cea434588681b Mon Sep 17 00:00:00 2001
> > > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > > Date: Sun, 6 Apr 2025 16:40:32 +0200
> > > Subject: [PATCH 03/16] Introduce pending flag for GUC assign hooks
> > >
> > > Currently an assing hook can perform some preprocessing of a new value,
> > > but it cannot change the behavior, which dictates that the new value
> > > will be applied immediately after the hook. Certain GUC options (like
> > > shared_buffers, coming in subsequent patches) may need coordinating work
> > > between backends to change, meaning we cannot apply it right away.
> > >
> > > Add a new flag "pending" for an assign hook to allow the hook indicate
> > > exactly that. If the pending flag is set after the hook, the new value
> > > will not be applied and it's handling becomes the hook's implementation
> > > responsibility.
> >
> > I doubt it makes sense to add this to the GUC system. I think it'd be better
> > to just use the GUC value as the desired "target" configuration and have a
> > function or a show-only GUC for reporting the current size.
> >
> > I don't think you can't just block application of the GUC until the resize is
> > complete. E.g. what if the value was too big and the new configuration needs
> > to fixed to be lower?
>
> I think it was a bit hasty to post another version of the patch without
> the design changes we've agreed upon last time. I'm still working on
> that (sorry, it takes time, I haven't wrote so much Perl for testing
> since forever), the current implementation doesn't include anything with
> GUC to simplify the discussion. I'm still convinced that multi-step GUC
> changing makes sense, but it has proven to be more complicated than I
> anticipated, so I'll spin up another thread to discuss when I come to
> it.
FWIW, I'm fairly convinced it's a completely dead end.
> > > From e2f48da8a8206711b24e34040d699431910fbf9c Mon Sep 17 00:00:00 2001
> > > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > > Date: Tue, 17 Jun 2025 11:47:04 +0200
> > > Subject: [PATCH 06/16] Address space reservation for shared memory
> > >
> > > Currently the shared memory layout is designed to pack everything tight
> > > together, leaving no space between mappings for resizing. Here is how it
> > > looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
> > > anonymous shared memory we talk about:
> > >
> > > 00400000-00490000 /path/bin/postgres
> > > ...
> > > 012d9000-0133e000 [heap]
> > > 7f443a800000-7f470a800000 /dev/zero (deleted)
> > > 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
> > > 7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
> > > ...
> > >
> > > Make the layout more dynamic via splitting every shared memory segment
> > > into two parts:
> > >
> > > * An anonymous file, which actually contains shared memory content. Such
> > > an anonymous file is created via memfd_create, it lives in memory,
> > > behaves like a regular file and semantically equivalent to an
> > > anonymous memory allocated via mmap with MAP_ANONYMOUS.
> > >
> > > * A reservation mapping, which size is much larger than required shared
> > > segment size. This mapping is created with flags PROT_NONE (which
> > > makes sure the reserved space is not used), and MAP_NORESERVE (to not
> > > count the reserved space against memory limits). The anonymous file is
> > > mapped into this reservation mapping.
> >
> > The commit message fails to explain why, if we're already relying on
> > MAP_NORESERVE, we need to anything else? Why can't we just have one maximally
> > sized allocation that's marked MAP_NORESERVE for all the parts that we don't
> > yet need?
>
> How do we return memory to the OS in that case? Currently it's done
> explicitly via truncating the anonymous file.
madvise with MADV_DONTNEED or MADV_REMOVE.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-09-18 14:05 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-09-26 18:04 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-09-29 06:57 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-09-29 06:57 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Fri, Sep 26, 2025 at 11:34 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> Sorry for late reply folks.
>
> > On Thu, Sep 18, 2025 at 09:52:03AM -0400, Andres Freund wrote:
> > > From 0a13e56dceea8cc7a2685df7ee8cea434588681b Mon Sep 17 00:00:00 2001
> > > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > > Date: Sun, 6 Apr 2025 16:40:32 +0200
> > > Subject: [PATCH 03/16] Introduce pending flag for GUC assign hooks
> > >
> > > Currently an assing hook can perform some preprocessing of a new value,
> > > but it cannot change the behavior, which dictates that the new value
> > > will be applied immediately after the hook. Certain GUC options (like
> > > shared_buffers, coming in subsequent patches) may need coordinating work
> > > between backends to change, meaning we cannot apply it right away.
> > >
> > > Add a new flag "pending" for an assign hook to allow the hook indicate
> > > exactly that. If the pending flag is set after the hook, the new value
> > > will not be applied and it's handling becomes the hook's implementation
> > > responsibility.
> >
> > I doubt it makes sense to add this to the GUC system. I think it'd be better
> > to just use the GUC value as the desired "target" configuration and have a
> > function or a show-only GUC for reporting the current size.
> >
> > I don't think you can't just block application of the GUC until the resize is
> > complete. E.g. what if the value was too big and the new configuration needs
> > to fixed to be lower?
>
> I think it was a bit hasty to post another version of the patch without
> the design changes we've agreed upon last time. I'm still working on
> that (sorry, it takes time, I haven't wrote so much Perl for testing
> since forever), the current implementation doesn't include anything with
> GUC to simplify the discussion. I'm still convinced that multi-step GUC
> changing makes sense, but it has proven to be more complicated than I
> anticipated, so I'll spin up another thread to discuss when I come to
> it.
>
> > > From 0a55bc15dc3a724f03e674048109dac1f248c406 Mon Sep 17 00:00:00 2001
> > > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > > Date: Fri, 4 Apr 2025 21:46:14 +0200
> > > Subject: [PATCH 04/16] Introduce pss_barrierReceivedGeneration
> > >
> > > Currently WaitForProcSignalBarrier allows to make sure the message sent
> > > via EmitProcSignalBarrier was processed by all ProcSignal mechanism
> > > participants.
> > >
> > > Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
> > > which will be updated when a process has received the message, but not
> > > processed it yet. This makes it possible to support a new mode of
> > > waiting, when ProcSignal participants want to synchronize message
> > > processing. To do that, a participant can wait via
> > > WaitForProcSignalBarrierReceived when processing a message, effectively
> > > making sure that all processes are going to start processing
> > > ProcSignalBarrier simultaneously.
> >
> > I doubt "online resizing" that requires synchronously processing the same
> > event, can really be called "online". There can be significant delays in
> > processing a barrier, stalling the entire server until that is reached seems
> > like a complete no-go for production systems?
> >
> > [...]
>
> > As mentioned above, this basically makes the entire feature not really
> > online. Besides the latency of some processes not getting to the barrier
> > immediately, there's also the issue that actually reserving large amounts of
> > memory can take a long time - during which all processes would be unavailable.
> >
> > I really don't see that being viable. It'd be one thing if that were a
> > "temporary" restriction, but the whole design seems to be fairly centered
> > around that.
> >
> > [...]
> >
> > Besides not really being online, isn't this a recipe for endless undetected
> > deadlocks? What if process A waits for a lock held by process B and process B
> > arrives at the barrier? Process A won't ever get there, because process B
> > can't make progress, because A is not making progress.
>
> Same as above, in the version I'm working right now it's changed in
> favor of an approach that looks more like the one from "online checksum
> change" patch. I've even stumbled upon a cases when a process was just
> killed and never arrive at the barrier, so that was it. The new approach
> makes certain parts simpler, but requires managing backends with
> different understanding of how large shared memory segments are for some
> time interval. Introducing a new parameter "number of available buffers"
> seems to be helpful to address all cases I've found so far.
>
> Btw, under "online" resizing I mostly understood "without restart", the
> goal was not to make it really "online".
>
> > > -#define MAX_ON_EXITS 20
> > > +#define MAX_ON_EXITS 40
> >
> > Why does a patch like this contain changes like this mixed in with the rest?
> > That's clearly not directly related to $subject.
>
> An artifact of rebasing, it belonged to 0007.
>
> > > From e2f48da8a8206711b24e34040d699431910fbf9c Mon Sep 17 00:00:00 2001
> > > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > > Date: Tue, 17 Jun 2025 11:47:04 +0200
> > > Subject: [PATCH 06/16] Address space reservation for shared memory
> > >
> > > Currently the shared memory layout is designed to pack everything tight
> > > together, leaving no space between mappings for resizing. Here is how it
> > > looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
> > > anonymous shared memory we talk about:
> > >
> > > 00400000-00490000 /path/bin/postgres
> > > ...
> > > 012d9000-0133e000 [heap]
> > > 7f443a800000-7f470a800000 /dev/zero (deleted)
> > > 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
> > > 7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
> > > ...
> > >
> > > Make the layout more dynamic via splitting every shared memory segment
> > > into two parts:
> > >
> > > * An anonymous file, which actually contains shared memory content. Such
> > > an anonymous file is created via memfd_create, it lives in memory,
> > > behaves like a regular file and semantically equivalent to an
> > > anonymous memory allocated via mmap with MAP_ANONYMOUS.
> > >
> > > * A reservation mapping, which size is much larger than required shared
> > > segment size. This mapping is created with flags PROT_NONE (which
> > > makes sure the reserved space is not used), and MAP_NORESERVE (to not
> > > count the reserved space against memory limits). The anonymous file is
> > > mapped into this reservation mapping.
> >
> > The commit message fails to explain why, if we're already relying on
> > MAP_NORESERVE, we need to anything else? Why can't we just have one maximally
> > sized allocation that's marked MAP_NORESERVE for all the parts that we don't
> > yet need?
>
> How do we return memory to the OS in that case? Currently it's done
> explicitly via truncating the anonymous file.
>
> > > * The file could be given a name, which improves readability when it
> > > comes to process maps.
> >
> > > * By default, Linux will not add file-backed shared mappings into a core dump,
> > > making it more convenient to work with them in PostgreSQL: no more huge dumps
> > > to process.
> >
> > That's just as well a downside, because you now can't investigate some
> > issues. This was already configurable via coredump_filter.
>
> This behaviour is configured via coredump_filter as well, so just the
> default value has been changed.
>
> > > From 942b69a0876b0e83303e6704da54c4c002a5a2d8 Mon Sep 17 00:00:00 2001
> > > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > > Date: Tue, 17 Jun 2025 11:22:02 +0200
> > > Subject: [PATCH 07/16] Introduce multiple shmem segments for shared buffers
> > >
> > > Add more shmem segments to split shared buffers into following chunks:
> > > * BUFFERS_SHMEM_SEGMENT: contains buffer blocks
> > > * BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
> > > * BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
> > > * CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
> > > * STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
> >
> > Why do all these need to be separate segments? Afaict we'll have to maximally
> > size everything other than BUFFERS_SHMEM_SEGMENT at start?
>
> Why would they need to me maxed out at the start? So far my rule of
> thumb was one segment for one structure which size depends on NBuffers,
> so that when changing NBuffers each segment could be adjusted
> independently.
>
Offlist Andres expressed that having multiple shared memory segments
may impact the time it takes to disconnect a backend. If the
application is using all the configured number of backends, a slow
disconnection will lead to a slow connection. If we want to go the
route of multple segments (as many as 5) it would make sense to
measure that impact first.
Maxing out at start avoids using multiple segments. Those segments
have much much lower memory compared to the buffer blocks even when
maxed out with a reasonable max_shared_buffers setting. We avoid
complicating code for a small increase in shared memory.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2025-10-13 15:58 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-14 08:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 2 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2025-10-13 15:58 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
Hi,
I started studying the interaction of the checkpointer process with
buffer pool resizing. Soon I noticed that the checkpointer didn't load
the config as frequently as other backends. When it is executing a
checkpoint, it does not reload the config for the entire duration of
the checkpoint for example. As the synchronization is implemented, in
the set of patches so far, the checkpointer will not see the new value
of shared_buffers and will not acknowledge the proc signal barrier and
thus not enter the synchronized buffer resizing. However, other
backends will notice that the checkpointer has received the proc
signal barrier and will enter the synchronization process. Once the
proc signal barrier is received by all the backends, the backends
which have entered the synchronization process will move forward with
resizing buffer pool leaving behind those who have received but not
acknowledged the proc signal barrier. At the end there will be two
sets of backends, one which have entered synchronization and see the
buffer pool with new size and the other which haven't entered
synchronization and do not see the buffer pool with new size. This
leads to SIGBUS, SIG 11 in the other set of backends. I saw this
mostly with the checkpointer process but we also saw it with other
types of backends.
Every aspect of buffer resizing that I started looking at was blocked
by this behaviour. Since there were already other suggestions and
comments about the current UI as well as synchronization mechanism, I
started implementing a different UI and synchronization as described
below. The WIP implementation is available in the attached set of
patches.
Patches 0001 to 0016 are the same as the previous patchset. I haven't
touched them in case someone would like to see an incremental change.
However, it's getting unwieldy at this point, so I will squash
relevant patches together and provide a patchset with fewer patches
next.
0017 reverts to 0003 and gets rid of the "pending" GUC flag which is
not required by the new UI. They will vanish from the next patchset.
0018 implements the new UI described below.
New UI and synchronization
======================
0018 changes the way "shared_buffers" is handled.
a. A new global variable NBuffersPending is used to hold the value of
this GUC. When the server starts, shared memory required by the buffer
manager is calculated using NBuffersPending instead of NBuffers. Once
the shared memory is allocated, NBuffers is set to NBuffersPending.
NBuffers, thus shows the number of buffers in the buffer pool instead
of the value of the GUC.
b. "shared_buffers" is PGC_SIGHUP now so it can be changed using ALTER
SYSTEM ... SET shared_buffers = ...; followed by SELECT
pg_reload_config(). But this does not resize the buffer pool. It
merely sets NBuffersPending to the new value. A new function
pg_resize_buffer_pool() (described later) can be used to resize the
buffer pool to the pending value.
c. show "shared_buffers" shows the value of NBuffers, and
NBuffersPending if it differs from NBuffers. I think we need some
adjustment here when the resizing is in progress since the value of
NBuffers would be changed to the size of the active buffer pool
(explained later in the email), but I haven't worked out those details
yet.
A new GUC max_shared_buffers sets the upper limit on "shared_buffers".
It is PGC_POSTMASTER; requires a restart to change the value. This GUC
is used a. to reserve the address space for future expansion of the
buffer pool and b. allocate memory for a maximally sized buffer lookup
table at the server start. We may decide to use the GUC to maximally
allocate data structures other than buffer blocks as suggested by
Andres. But these patches don't do that. The default for this GUC is
0, which means it will be the same as shared_buffers. This maintains
backward compatibility and also allows systems, which do not want to
resize shared buffer pool, to allocate minimum memory. When it is set
to a value other than 0, it should be set to a value higher than the
shared_buffers at the start.
We need to support the ALTER SYSTEM ... SET shared_buffers = "" for
backward compatibility. The users will still be able to perform ALTER
SYSTEM and restart the server with a newer size of buffer pool. Also
this allows the new buffer pool size to be written to
postgresql.auto.conf and persist it. With this we can simply use
pg_reload_conf() to load the new value along with other GUC changes.
pg_resize_buffer_pool() merely picks the new value from the backend
where it is executed and resizes the buffer pool. It does not need the
new value to be loaded in all the backends.
We may want to use a new PGC_ for this GUC but PGC_SIGHUP suffices for
the time being and it might be acceptable with clear documentation.
pg_resize_buffer_pool() implements phase wise buffer pool resizing
operation, but it does not block all the backends till the buffer pool
resizing is finished. It works as follows: Pasting from the prologue
in patch 0018.
When resizing the buffer pool is divided into two portions
- active buffer pool, which is the part of the buffer pool which
remains active even during resizing. Its size is given by
activeNBuffers. Newly allocated buffers will have their buffer ids
less than activeNBuffers.
- in-transit buffer pool, which is the part of the buffer pool which
may be accessible to some backends but not others depending upon the
time when a given backend processes a shrink/expand barrier. When
shrinking the buffer pool this is the part of the buffer pool which
will be evicted. When expanding the buffer pool this is the expanded
portion. Its size is given by transitNBuffers. The backends may see
buffer ids upto transitNBuffers till the resizing finishes.
Before starting resizing, activeNBuffers = transitNBuffers = NBuffers
where NBuffers is the size of buffer pool before resizing. NewNBuffers
is the new size of the shared buffer pool. After resizing finishes
activeNBuffers = transitNBuffers = NBuffers = newNBuffers.
In order to synchronize with other running backends, the coordinator
sends following ProcSignalBarriers in the order given below:
1. When shrinking the shared buffer pool the coordinator sends
SHBUF_SHRINK ProcSignalBarrier. Every backend sets activeNBuffers =
NewNBuffers to restrict its buffer pool allocations to the new size of
the buffer pool and acknowledges the ProcSignalBarrrier. Once every
backend has acknowledged, the coordinator evicts the buffers in the
area being shrunk. Note that tansitNBuffers is still NBuffers, so the
backends may see buffer ids upto NBuffers from earlier allocations
till eviction completes.
2. In both cases, when expanding the buffer pool or shrinking the
buffer pool, the coordinator sends SHBUF_RESIZE_MAP_AND_MEM
ProcSignalBarrier after resizing the shared memory segments and
initializing the required data structures if any. Every backend is
expected to adjust their shared memory segment address maps (by
calling AnonymousShmemResize()) and validate that their pointers to
the shared buffers structure are valid and have the right size. When
shrinking shared buffer pool transitNBuffers is set to NewNBuffers and
the backends should no longer see buffer ids beyond NewNBuffers; the
buffer resizing operation is finished at this stage. When expanding
they should set transitNBuffers to NewNBuffers to accommodate for the
backends which may accept the next barrier earlier than the others.
Once every backend acknowledges this barrier, the coordinator sends
the next barrier when expanding the buffer pool.
3. When expanding the buffer pool, the coordinator sends SHBUF_EXPAND
ProcSignalBarrier. The backends are expected to set activeNBuffers =
NewNBuffers and start allocating buffers from the expanded range. The
coordinator uses this barrier to know when all the backends have
settled using the new size of the buffer pool.
For either operation, at most two barriers are sent.
All this together in action looks like (See tests in the patch for
more examples)
SHOW shared_buffers; -- default
shared_buffers
----------------
128MB
(1 row)
ALTER SYSTEM SET shared_buffers = '64MB';
SELECT pg_reload_conf();
pg_reload_conf
----------------
t
(1 row)
SHOW shared_buffers;
shared_buffers
-----------------------
128MB (pending: 64MB)
(1 row)
SELECT pg_resize_shared_buffers();
pg_resize_shared_buffers
--------------------------
t
(1 row)
SHOW shared_buffers;
shared_buffers
----------------
64MB
(1 row)
ALTER SYSTEM SET shared_buffers = '256MB';
SELECT pg_reload_conf();
pg_reload_conf
----------------
t
(1 row)
SHOW shared_buffers;
shared_buffers
-----------------------
64MB (pending: 256MB)
(1 row)
SELECT pg_resize_shared_buffers();
pg_resize_shared_buffers
--------------------------
t
(1 row)
SHOW shared_buffers;
shared_buffers
----------------
256MB
(1 row)
On Thu, Sep 18, 2025 at 7:22 PM Andres Freund <andres@anarazel.de> wrote:
>
> > From 0a13e56dceea8cc7a2685df7ee8cea434588681b Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Sun, 6 Apr 2025 16:40:32 +0200
> > Subject: [PATCH 03/16] Introduce pending flag for GUC assign hooks
> >
> > Currently an assing hook can perform some preprocessing of a new value,
> > but it cannot change the behavior, which dictates that the new value
> > will be applied immediately after the hook. Certain GUC options (like
> > shared_buffers, coming in subsequent patches) may need coordinating work
> > between backends to change, meaning we cannot apply it right away.
> >
> > Add a new flag "pending" for an assign hook to allow the hook indicate
> > exactly that. If the pending flag is set after the hook, the new value
> > will not be applied and it's handling becomes the hook's implementation
> > responsibility.
>
> I doubt it makes sense to add this to the GUC system. I think it'd be better
> to just use the GUC value as the desired "target" configuration and have a
> function or a show-only GUC for reporting the current size.
This has been taken care of in the new implementation with slightly
different approach to show command as described above.
>
> I don't think you can't just block application of the GUC until the resize is
> complete. E.g. what if the value was too big and the new configuration needs
> to fixed to be lower?
>
With the above approach, the application of the GUC won't be blocked
but if the size being applied is taking too long, the operation will
be required to be cancelled before the new resize can happen. That's a
part that needs some work. Chasing a moving target requires a very
complex implementation, which would be good to avoid in the first
version at least. However, we should leave room for that future
enhancement. The current implementation gives that flexibility, I
think.
>
> > From 0a55bc15dc3a724f03e674048109dac1f248c406 Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Fri, 4 Apr 2025 21:46:14 +0200
> > Subject: [PATCH 04/16] Introduce pss_barrierReceivedGeneration
> >
> > Currently WaitForProcSignalBarrier allows to make sure the message sent
> > via EmitProcSignalBarrier was processed by all ProcSignal mechanism
> > participants.
> >
> > Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
> > which will be updated when a process has received the message, but not
> > processed it yet. This makes it possible to support a new mode of
> > waiting, when ProcSignal participants want to synchronize message
> > processing. To do that, a participant can wait via
> > WaitForProcSignalBarrierReceived when processing a message, effectively
> > making sure that all processes are going to start processing
> > ProcSignalBarrier simultaneously.
>
> I doubt "online resizing" that requires synchronously processing the same
> event, can really be called "online". There can be significant delays in
> processing a barrier, stalling the entire server until that is reached seems
> like a complete no-go for production systems?
>
> > From 78bc0a49f8ebe17927abd66164764745ecc6d563 Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Tue, 17 Jun 2025 14:16:55 +0200
> > Subject: [PATCH 11/16] Allow to resize shared memory without restart
> >
> > Add assing hook for shared_buffers to resize shared memory using space,
> > introduced in the previous commits without requiring PostgreSQL restart.
> > Essentially the implementation is based on two mechanisms: a
> > ProcSignalBarrier is used to make sure all processes are starting the
> > resize procedure simultaneously, and a global Barrier is used to
> > coordinate after that and make sure all finished processes are waiting
> > for others that are in progress.
> >
> > The resize process looks like this:
> >
> > * The GUC assign hook sets a flag to let the Postmaster know that resize
> > was requested.
> >
> > * Postmaster verifies the flag in the event loop, and starts the resize
> > by emitting a ProcSignal barrier.
> >
> > * All processes, that participate in ProcSignal mechanism, begin to
> > process ProcSignal barrier. First a process waits until all processes
> > have confirmed they received the message and can start simultaneously.
>
> As mentioned above, this basically makes the entire feature not really
> online. Besides the latency of some processes not getting to the barrier
> immediately, there's also the issue that actually reserving large amounts of
> memory can take a long time - during which all processes would be unavailable.
>
> I really don't see that being viable. It'd be one thing if that were a
> "temporary" restriction, but the whole design seems to be fairly centered
> around that.
In the new implementation regular backends are not stalled when the
resizing is going on. They continue their work with possible temporary
performance degradation (this needs to be measured).
>
> > From experiment it turns out that shared mappings have to be extended
> > separately for each process that uses them. Another rough edge is that a
> > backend blocked on ReadCommand will not apply shared_buffers change
> > until it receives something.
>
> That's not a rough edge, that basically makes the feature unusable, no?
New synchronization doesn't have this problem since it doesn't require
every backend to load the new value. The value being loaded only in
the backend where pg_resize_buffer_pool() is being run is enough.
>
> > From 942b69a0876b0e83303e6704da54c4c002a5a2d8 Mon Sep 17 00:00:00 2001
> > From: Dmitrii Dolgov <9erthalion6@gmail.com>
> > Date: Tue, 17 Jun 2025 11:22:02 +0200
> > Subject: [PATCH 07/16] Introduce multiple shmem segments for shared buffers
> >
> > Add more shmem segments to split shared buffers into following chunks:
> > * BUFFERS_SHMEM_SEGMENT: contains buffer blocks
> > * BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
> > * BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
> > * CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
> > * STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
>
> Why do all these need to be separate segments? Afaict we'll have to maximally
> size everything other than BUFFERS_SHMEM_SEGMENT at start?
>
I am leaning towards that. I will implement that soon.
On Wed, Oct 1, 2025 at 2:40 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
>
> I see you folks are inclined to keep some small segments static and
> allocate maximum allowed memory for it. It's an option, at the end of
> the day we need to experiment and measure both approaches.
I did measure performance with a maximally sized buffer lookup table
(shared_buffers = 128MB, max_shared_buffers = 10GB) on my laptop.
There was no noticeable difference in the performance. I will post
formal numbers with the next patchset.
>
>
> > * Every process recalculates shared memory size based on the new
> > NBuffers, adjusts its size using ftruncate and adjust reservation
> > permissions with mprotect. One elected process signals the postmaster
> > to do the same.
>
> If we just used a single memory mapping with all unused parts marked
> MAP_NORESERVE, we wouldn't need this (and wouldn't need a fair bit of other
> work in this patchset)..
>
On Sat, Sep 27, 2025 at 12:06 AM Andres Freund <andres@anarazel.de> wrote:
>
> > How do we return memory to the OS in that case? Currently it's done
> > explicitly via truncating the anonymous file.
>
> madvise with MADV_DONTNEED or MADV_REMOVE.
The patchset still uses the ftruncate + mprotect. I have questions
apart from portability concerns about your proposal. MADV_DONTNEED
documentation says
After a successful MADV_DONTNEED operation, the
semantics of memory access in the specified region are changed:
subsequent accesses
of pages in the range will succeed, but will result
in either repopulating the memory contents from the up-to-date
contents of the
underlying mapped file (for shared file mappings, shared
anonymous mappings, and shmem-based techniques such as System V shared
mem‐
ory segments) or zero-fill-on-demand pages for anonymous
private mappings.
Note that, when applied to shared mappings,
MADV_DONTNEED might not lead to immediate freeing of the pages in the
range. The kernel
is free to delay freeing the pages until an appropriate
moment. The resident set size (RSS) of the calling process will be
immedi‐
ately reduced however.
MADV_DONTNEED cannot be applied to locked pages, Huge
TLB pages, or VM_PFNMAP pages. (Pages marked with the kernel-internal
VM_PFN‐
MAP flag are special memory areas that are not managed
by the virtual memory subsystem. Such pages are typically created by
device
drivers that map the pages into user space.)
and MADV_REMOVE (since Linux 2.6.16)
Free up a given range of pages and its associated
backing store. This is equivalent to punching a hole in the
corresponding byte
range of the backing store (see fallocate(2)).
Subsequent accesses in the specified address range will see bytes
containing zero.
The specified address range must be mapped shared and
writable. This flag cannot be applied to locked pages, Huge TLB
pages, or
VM_PFNMAP pages.
Combining these two,
1. The access to the freed memory doesn't give any error but returns
0. Won't that lead to silent corruption?
2. Those are not supported with huge tlb pages. So can not be used
when huge pages = on?
With the current approach, we get SIGBUS and SIG 11 when the process
tries to access the freed memory. That protection won't be there with
madvise().
The synchronization mechanism in this patch is inspired from Thomas's
implementation posted in [1].
I still need to go through Tomas's detailed comments and address those
which still apply. And the patches are still WIP, with many TODOs. But
I wanted to get some feedback on the proposed UI and synchronization
as described above.
I will be looking into the cases below one by one
1. New backends join while the synchronization is going on. An
existing backend exiting.
2. Failure or crash in the backend which is executing pg_resize_buffer_pool()
3. Fix crashes in the tests.
[1] postgr.es/m/CA+hUKGL5hW3i_pk5y_gcbF_C5kP-pWFjCuM8bAyCeHo3xUaH8g@mail.gmail.com
--
Best Wishes,
Ashutosh Bapat
Attachments:
[application/x-patch] 0003-Introduce-pending-flag-for-GUC-assign-hooks-20251013.patch (12.7K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/2-0003-Introduce-pending-flag-for-GUC-assign-hooks-20251013.patch)
download | inline diff:
From f616d9a4e88c9edabae143d6f402c3f54730cd0d Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Sun, 6 Apr 2025 16:40:32 +0200
Subject: [PATCH 03/19] Introduce pending flag for GUC assign hooks
Currently an assing hook can perform some preprocessing of a new value,
but it cannot change the behavior, which dictates that the new value
will be applied immediately after the hook. Certain GUC options (like
shared_buffers, coming in subsequent patches) may need coordinating work
between backends to change, meaning we cannot apply it right away.
Add a new flag "pending" for an assign hook to allow the hook indicate
exactly that. If the pending flag is set after the hook, the new value
will not be applied and it's handling becomes the hook's implementation
responsibility.
Note, that this also requires changes in the way how GUCs are getting
reported, but the patch does not cover that yet.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++++++++++++---------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 61 insertions(+), 40 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index eceab341255..cc48b253bc8 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2197,7 +2197,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra)
+assign_max_wal_size(int newval, void *extra, bool *pending)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 608f10d9412..e40dae2ddf2 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra)
+assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,12 +1161,12 @@ assign_maintenance_io_concurrency(int newval, void *extra)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra)
+assign_io_max_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(newval, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra)
+assign_io_combine_limit(int newval, void *extra, bool *pending)
{
io_combine_limit = Min(io_max_combine_limit, newval);
}
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index 25f739a6a17..1726a7c0993 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1951,7 +1951,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra)
+assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1984,7 +1984,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra)
+assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2007,7 +2007,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra)
+assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2030,7 +2030,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra)
+assign_tcp_user_timeout(int newval, void *extra, bool *pending)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 7dd75a490aa..193efeb9022 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3597,7 +3597,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra)
+assign_transaction_timeout(int newval, void *extra, bool *pending)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 8794e26ef1d..c9361a0e423 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1681,6 +1681,7 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
+ bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1689,9 +1690,13 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra);
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
+ conf->assign_hook(newval, extra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
+ }
break;
}
case PGC_REAL:
@@ -2047,13 +2052,18 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
+ bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra);
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
+ conf->reset_extra,
+ &pending);
+ if (!pending)
+ {
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
+ }
break;
}
case PGC_REAL:
@@ -2430,16 +2440,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
+ bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
+ }
}
break;
}
@@ -3856,18 +3871,24 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
+ bool pending = false;
+
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra);
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
+ conf->assign_hook(newval, newextra, &pending);
+
+ if (!pending)
+ {
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
+ }
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index 8f7cf531fbc..ef59ae62008 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra)
+assign_max_stack_depth(int newval, void *extra, bool *pending)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f21ec37da89..c3056cd2da8 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra);
+typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 82ac8646a8d..658c799419e 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,12 +81,12 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra);
-extern void assign_io_max_combine_limit(int newval, void *extra);
-extern void assign_io_combine_limit(int newval, void *extra);
-extern void assign_max_wal_size(int newval, void *extra);
+extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
+extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
+extern void assign_max_wal_size(int newval, void *extra, bool *pending);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra);
+extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -141,13 +141,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra);
+extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra);
+extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra);
+extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra);
+extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -163,7 +163,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra);
+extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.34.1
[application/x-patch] 0002-Process-config-reload-in-AIO-workers-20251013.patch (1.8K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/3-0002-Process-config-reload-in-AIO-workers-20251013.patch)
download | inline diff:
From 8fe9b13edfb2dc84047baa2fed9f48246b42af85 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 15:14:33 +0200
Subject: [PATCH 02/19] Process config reload in AIO workers
Currenly AIO workers process interrupts only via CHECK_FOR_INTERRUPTS,
which does not include ConfigReloadPending. Thus we need to check for it
explicitly.
---
src/backend/storage/aio/method_worker.c | 25 +++++++++++++++++++++++++
1 file changed, 25 insertions(+)
diff --git a/src/backend/storage/aio/method_worker.c b/src/backend/storage/aio/method_worker.c
index b5ac073a910..d1c6da89c4b 100644
--- a/src/backend/storage/aio/method_worker.c
+++ b/src/backend/storage/aio/method_worker.c
@@ -80,6 +80,7 @@ static void pgaio_worker_shmem_init(bool first_time);
static bool pgaio_worker_needs_synchronous_execution(PgAioHandle *ioh);
static int pgaio_worker_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+static void pgaio_worker_process_interrupts(void);
const IoMethodOps pgaio_worker_ops = {
.shmem_size = pgaio_worker_shmem_size,
@@ -463,6 +464,8 @@ IoWorkerMain(const void *startup_data, size_t startup_data_len)
int nwakeups = 0;
int worker;
+ pgaio_worker_process_interrupts();
+
/*
* Try to get a job to do.
*
@@ -592,3 +595,25 @@ pgaio_workers_enabled(void)
{
return io_method == IOMETHOD_WORKER;
}
+
+/*
+ * Process any new interrupts.
+ */
+static void
+pgaio_worker_process_interrupts(void)
+{
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
+ if (ConfigReloadPending)
+ {
+ ConfigReloadPending = false;
+ ProcessConfigFile(PGC_SIGHUP);
+ }
+
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+}
--
2.34.1
[application/x-patch] 0001-Add-system-view-for-shared-buffer-lookup-ta-20251013.patch (9.6K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/4-0001-Add-system-view-for-shared-buffer-lookup-ta-20251013.patch)
download | inline diff:
From 1a13e00fd8b069d653f08132a3d35c7c17fdf5c9 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 25 Aug 2025 19:23:50 +0530
Subject: [PATCH 01/19] Add system view for shared buffer lookup table
The view exposes the contents of the shared buffer lookup table for
debugging, testing and investigation.
TODO:
It is better to place this view in pg_buffercache. But it's added as a
system view since BufHashTable is not exposed outside buf_table.c. To
move it to pg_buffercache, we should move the function
pg_get_buffer_lookup_table() to pg_buffercache which invokes
BufTableGetContent() by passing it the tuple store and tuple descriptor.
BufTableGetContent fills the tuple store. The partitions are locked by
pg_get_buffer_lookup_table().
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
doc/src/sgml/system-views.sgml | 89 ++++++++++++++++++++++++++
src/backend/catalog/system_views.sql | 7 ++
src/backend/storage/buffer/buf_table.c | 61 ++++++++++++++++++
src/include/catalog/pg_proc.dat | 11 ++++
src/test/regress/expected/rules.out | 7 ++
5 files changed, 175 insertions(+)
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 7971498fe75..8f3e2741051 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -71,6 +71,11 @@
<entry>backend memory contexts</entry>
</row>
+ <row>
+ <entry><link linkend="view-pg-buffer-lookup-table"><structname>pg_buffer_lookup_table</structname></link></entry>
+ <entry>shared buffer lookup table</entry>
+ </row>
+
<row>
<entry><link linkend="view-pg-config"><structname>pg_config</structname></link></entry>
<entry>compile-time configuration parameters</entry>
@@ -901,6 +906,90 @@ AND c1.path[c2.level] = c2.path[c2.level];
</para>
</sect1>
+ <sect1 id="view-pg-buffer-lookup-table">
+ <title><structname>pg_buffer_lookup_table</structname></title>
+ <indexterm>
+ <primary>pg_buffer_lookup_table</primary>
+ </indexterm>
+ <para>
+ The <structname>pg_buffer_lookup_table</structname> view exposes the current
+ contents of the shared buffer lookup table. Each row represents an entry in
+ the lookup table mapping a relation page to the ID of buffer in which it is
+ cached. The shared buffer lookup table is locked for a short duration while
+ reading so as to ensure consistency. This may affect performance if this view
+ is queried very frequently.
+ </para>
+ <table id="pg-buffer-lookup-table-view" xreflabel="pg_buffer_lookup_table">
+ <title><structname>pg_buffer_lookup_table</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>tablespace</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the tablespace containing the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>database</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database containing the relation (zero for shared relations)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relfilenode</structfield> <type>oid</type>
+ </para>
+ <para>
+ relfilenode identifying the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>forknum</structfield> <type>int2</type>
+ </para>
+ <para>
+ Fork number within the relation (see <xref linkend="storage-file-layout"/>)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>blocknum</structfield> <type>int8</type>
+ </para>
+ <para>
+ Block number within the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>bufferid</structfield> <type>int4</type>
+ </para>
+ <para>
+ ID of the buffer caching the page
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ <para>
+ Access to this view is restricted to members of the
+ <literal>pg_read_all_stats</literal> role by default.
+ </para>
+ </sect1>
+
<sect1 id="view-pg-config">
<title><structname>pg_config</structname></title>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 823776c1498..c7240250c07 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1436,3 +1436,10 @@ REVOKE ALL ON pg_aios FROM PUBLIC;
GRANT SELECT ON pg_aios TO pg_read_all_stats;
REVOKE EXECUTE ON FUNCTION pg_get_aios() FROM PUBLIC;
GRANT EXECUTE ON FUNCTION pg_get_aios() TO pg_read_all_stats;
+
+CREATE VIEW pg_buffer_lookup_table AS
+ SELECT * FROM pg_get_buffer_lookup_table();
+REVOKE ALL ON pg_buffer_lookup_table FROM PUBLIC;
+GRANT SELECT ON pg_buffer_lookup_table TO pg_read_all_stats;
+REVOKE EXECUTE ON FUNCTION pg_get_buffer_lookup_table() FROM PUBLIC;
+GRANT EXECUTE ON FUNCTION pg_get_buffer_lookup_table() TO pg_read_all_stats;
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 9d256559bab..1f6e215a2ca 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -21,7 +21,12 @@
*/
#include "postgres.h"
+#include "fmgr.h"
+#include "funcapi.h"
#include "storage/buf_internals.h"
+#include "storage/lwlock.h"
+#include "utils/rel.h"
+#include "utils/builtins.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -159,3 +164,59 @@ BufTableDelete(BufferTag *tagPtr, uint32 hashcode)
if (!result) /* shouldn't happen */
elog(ERROR, "shared buffer hash table corrupted");
}
+
+/*
+ * SQL callable function to report contents of the shared buffer lookup table.
+ */
+Datum
+pg_get_buffer_lookup_table(PG_FUNCTION_ARGS)
+{
+#define PG_GET_BUFFER_LOOKUP_TABLE_COLS 6
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ HASH_SEQ_STATUS hstat;
+ BufferLookupEnt *ent;
+ Datum values[PG_GET_BUFFER_LOOKUP_TABLE_COLS];
+ bool nulls[PG_GET_BUFFER_LOOKUP_TABLE_COLS];
+ int i;
+
+ memset(nulls, 0, sizeof(nulls));
+
+ /*
+ * We put all the tuples into a tuplestore in one scan of the hashtable.
+ * This avoids any issue of the hashtable possibly changing between calls.
+ */
+ InitMaterializedSRF(fcinfo, 0);
+
+ Assert(rsinfo->setDesc->natts == PG_GET_BUFFER_LOOKUP_TABLE_COLS);
+
+ /*
+ * Lock all buffer mapping partitions to ensure a consistent view of the
+ * hash table during the scan. Must grab LWLocks in partition-number order
+ * to avoid LWLock deadlock.
+ */
+ for (i = 0; i < NUM_BUFFER_PARTITIONS; i++)
+ LWLockAcquire(BufMappingPartitionLockByIndex(i), LW_SHARED);
+
+ hash_seq_init(&hstat, SharedBufHash);
+ while ((ent = (BufferLookupEnt *) hash_seq_search(&hstat)) != NULL)
+ {
+ values[0] = ObjectIdGetDatum(ent->key.spcOid);
+ values[1] = ObjectIdGetDatum(ent->key.dbOid);
+ values[2] = ObjectIdGetDatum(ent->key.relNumber);
+ values[3] = ObjectIdGetDatum(ent->key.forkNum);
+ values[4] = UInt32GetDatum(ent->key.blockNum);
+ values[5] = Int32GetDatum(ent->id);
+
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
+
+ /*
+ * Release all buffer mapping partition locks in the reverse order so as
+ * to avoid LWLock deadlock.
+ */
+ for (i = NUM_BUFFER_PARTITIONS - 1; i >= 0; i--)
+ LWLockRelease(BufMappingPartitionLockByIndex(i));
+
+ return (Datum) 0;
+}
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index b51d2b17379..e631323a325 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8600,6 +8600,17 @@
proargmodes => '{o,o,o}', proargnames => '{name,type,size}',
prosrc => 'pg_get_dsm_registry_allocations' },
+# buffer lookup table
+{ oid => '5102',
+ descr => 'shared buffer lookup table',
+ proname => 'pg_get_buffer_lookup_table', prorows => '6', proretset => 't',
+ provolatile => 'v', prorettype => 'record',
+ proargtypes => '', proallargtypes => '{oid,oid,oid,int2,int8,int4}',
+ proargmodes => '{o,o,o,o,o,o}',
+ proargnames => '{tablespace,database,relfilenode,forknum,blocknum,bufferid}',
+ prosrc => 'pg_get_buffer_lookup_table'
+},
+
# memory context of local backend
{ oid => '2282',
descr => 'information about all memory contexts of local backend',
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 16753b2e4c0..83f566d3218 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1330,6 +1330,13 @@ pg_backend_memory_contexts| SELECT name,
free_chunks,
used_bytes
FROM pg_get_backend_memory_contexts() pg_get_backend_memory_contexts(name, ident, type, level, path, total_bytes, total_nblocks, free_bytes, free_chunks, used_bytes);
+pg_buffer_lookup_table| SELECT tablespace,
+ database,
+ relfilenode,
+ forknum,
+ blocknum,
+ bufferid
+ FROM pg_get_buffer_lookup_table() pg_get_buffer_lookup_table(tablespace, database, relfilenode, forknum, blocknum, bufferid);
pg_config| SELECT name,
setting
FROM pg_config() pg_config(name, setting);
base-commit: 7a662a46ebf74e9fa15cb62b592b4bf00c96fc94
--
2.34.1
[application/x-patch] 0004-Introduce-pss_barrierReceivedGeneration-20251013.patch (7.3K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/5-0004-Introduce-pss_barrierReceivedGeneration-20251013.patch)
download | inline diff:
From a7cdc1871e0626b0b3f60ea68044ee77eca192c3 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 4 Apr 2025 21:46:14 +0200
Subject: [PATCH 04/19] Introduce pss_barrierReceivedGeneration
Currently WaitForProcSignalBarrier allows to make sure the message sent
via EmitProcSignalBarrier was processed by all ProcSignal mechanism
participants.
Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration,
which will be updated when a process has received the message, but not
processed it yet. This makes it possible to support a new mode of
waiting, when ProcSignal participants want to synchronize message
processing. To do that, a participant can wait via
WaitForProcSignalBarrierReceived when processing a message, effectively
making sure that all processes are going to start processing
ProcSignalBarrier simultaneously.
---
src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------
src/include/storage/procsignal.h | 1 +
2 files changed, 54 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 087821311cc..eb3ceaae809 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -58,7 +58,10 @@
* of it. For such use cases, we set a bit in pss_barrierCheckMask and then
* increment the current "barrier generation"; when the new barrier generation
* (or greater) appears in the pss_barrierGeneration flag of every process,
- * we know that the message has been received everywhere.
+ * we know that the message has been received and processed everywhere. In case
+ * if we only need to know only that the message was received everywhere (e.g.
+ * receiving processes need to handle the message in a coordinated fashion)
+ * use pss_barrierReceivedGeneration in the same way.
*/
typedef struct
{
@@ -70,6 +73,7 @@ typedef struct
/* Barrier-related fields (not protected by pss_mutex) */
pg_atomic_uint64 pss_barrierGeneration;
+ pg_atomic_uint64 pss_barrierReceivedGeneration;
pg_atomic_uint32 pss_barrierCheckMask;
ConditionVariable pss_barrierCV;
} ProcSignalSlot;
@@ -152,6 +156,8 @@ ProcSignalShmemInit(void)
slot->pss_cancel_key_len = 0;
MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration,
+ PG_UINT64_MAX);
pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
ConditionVariableInit(&slot->pss_barrierCV);
}
@@ -199,6 +205,8 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
barrier_generation =
pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration);
pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration,
+ barrier_generation);
if (cancel_key_len > 0)
memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len);
@@ -263,6 +271,7 @@ CleanupProcSignalState(int status, Datum arg)
* no barrier waits block on it.
*/
pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX);
SpinLockRelease(&slot->pss_mutex);
@@ -416,12 +425,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type)
return generation;
}
-/*
- * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
- * requested by a specific call to EmitProcSignalBarrier() have taken effect.
- */
-void
-WaitForProcSignalBarrier(uint64 generation)
+static void
+WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly)
{
Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration));
@@ -436,12 +441,17 @@ WaitForProcSignalBarrier(uint64 generation)
uint64 oldval;
/*
- * It's important that we check only pss_barrierGeneration here and
- * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared
- * before the barrier is actually absorbed, but pss_barrierGeneration
+ * It's important that we check only pss_barrierGeneration &
+ * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in
+ * pss_barrierCheckMask get cleared before the barrier is actually
+ * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration
* is updated only afterward.
*/
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
while (oldval < generation)
{
if (ConditionVariableTimedSleep(&slot->pss_barrierCV,
@@ -450,7 +460,11 @@ WaitForProcSignalBarrier(uint64 generation)
ereport(LOG,
(errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier",
(int) pg_atomic_read_u32(&slot->pss_pid))));
- oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
+
+ if (receivedOnly)
+ oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration);
+ else
+ oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration);
}
ConditionVariableCancelSleep();
}
@@ -464,12 +478,33 @@ WaitForProcSignalBarrier(uint64 generation)
* The caller is probably calling this function because it wants to read
* the shared state or perform further writes to shared state once all
* backends are known to have absorbed the barrier. However, the read of
- * pss_barrierGeneration was performed unlocked; insert a memory barrier
- * to separate it from whatever follows.
+ * pss_barrierGeneration & pss_barrierReceivedGeneration was performed
+ * unlocked; insert a memory barrier to separate it from whatever follows.
*/
pg_memory_barrier();
}
+/*
+ * WaitForProcSignalBarrier - wait until it is guaranteed that all changes
+ * requested by a specific call to EmitProcSignalBarrier() have taken effect.
+ */
+void
+WaitForProcSignalBarrier(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, false);
+}
+
+/*
+ * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all
+ * backends have observed the message sent by a specific call to
+ * EmitProcSignalBarrier().
+ */
+void
+WaitForProcSignalBarrierReceived(uint64 generation)
+{
+ WaitForProcSignalBarrierInternal(generation, true);
+}
+
/*
* Handle receipt of an interrupt indicating a global barrier event.
*
@@ -523,6 +558,10 @@ ProcessProcSignalBarrier(void)
if (local_gen == shared_gen)
return;
+ /* The message is observed, record that */
+ pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration,
+ shared_gen);
+
/*
* Get and clear the flags that are set for this backend. Note that
* pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index afeeb1ca019..2733bbb8c5b 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, const uint8 *cancel_key, int cance
extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type);
extern void WaitForProcSignalBarrier(uint64 generation);
+extern void WaitForProcSignalBarrierReceived(uint64 generation);
extern void ProcessProcSignalBarrier(void);
extern void procsignal_sigusr1_handler(SIGNAL_ARGS);
--
2.34.1
[application/x-patch] 0005-Allow-to-use-multiple-shared-memory-mapping-20251013.patch (31.3K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/6-0005-Allow-to-use-multiple-shared-memory-mapping-20251013.patch)
download | inline diff:
From bcb7a085a92f6b6c3bbd4c75819d5b4c4462ab03 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH 05/19] Allow to use multiple shared memory mappings
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
There is only fixed amount of available segments, currently only one
main shared memory segment is allocated. A new shared memory API is
introduces, extended with a segment as a new parameter. As a path of
least resistance, the original API is kept in place, utilizing the main
shared memory segment.
---
src/backend/port/posix_sema.c | 4 +-
src/backend/port/sysv_sema.c | 4 +-
src/backend/port/sysv_shmem.c | 138 +++++++++++++++++++---------
src/backend/port/win32_sema.c | 2 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 63 +++++++------
src/backend/storage/ipc/shmem.c | 148 +++++++++++++++++++++---------
src/backend/storage/lmgr/lwlock.c | 15 ++-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_sema.h | 2 +-
src/include/storage/pg_shmem.h | 18 ++++
src/include/storage/shmem.h | 11 +++
12 files changed, 283 insertions(+), 128 deletions(-)
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index 269c7460817..401e1113fa1 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas)
* we don't have to expose the counters to other processes.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 6ac83ea1a82..7bb363989c4 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -327,7 +327,7 @@ PGSemaphoreShmemSize(int maxSemas)
* have clobbered.)
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
struct stat statbuf;
@@ -348,7 +348,7 @@ PGReserveSemaphores(int maxSemas)
* ShmemAlloc() won't be ready yet.
*/
sharedSemas = (PGSemaphore)
- ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas));
+ ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment);
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..56af0231d24 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,8 +94,19 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+typedef struct AnonymousMapping
+{
+ int shmem_segment;
+ Size shmem_size; /* Size of the mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+} AnonymousMapping;
+
+static AnonymousMapping Mappings[ANON_MAPPINGS];
+
+/* Keeps track of used mapping segments */
+static int next_free_segment = 0;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+static const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ default:
+ return "unknown";
+ }
+}
+
+static void
+DebugMappings()
+{
+ for(int i = 0; i < next_free_segment; i++)
+ {
+ AnonymousMapping m = Mappings[i];
+ elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
+ MappingName(i), m.shmem, m.shmem_size);
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(AnonymousMapping *mapping)
{
- Size allocsize = *size;
+ Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
@@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size)
PG_MMAP_FLAGS | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ {
+ DebugMappings();
+ elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ MappingName(mapping->shmem_segment), allocsize);
+ }
}
#endif
@@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size)
* Use the original size, not the rounded-up value, when falling back
* to non-huge pages.
*/
- allocsize = *size;
+ allocsize = mapping->shmem_size;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
PG_MMAP_FLAGS, -1, 0);
mmap_errno = errno;
@@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size)
if (ptr == MAP_FAILED)
{
errno = mmap_errno;
+ DebugMappings();
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment)),
(mmap_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
@@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size)
allocsize) : 0));
}
- *size = allocsize;
- return ptr;
+ mapping->shmem = ptr;
+ mapping->shmem_size = allocsize;
}
/*
@@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonymousMapping m = Mappings[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
@@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ mapping->shmem_size = size;
+ mapping->shmem_segment = next_free_segment;
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping);
+
+ next_free_segment++;
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + next_free_segment;
for (;;)
{
@@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ mapping->seg_addr = memAddress;
+ mapping->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (mapping->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) mapping->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < next_free_segment; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ AnonymousMapping m = Mappings[i];
+
+ if (m.seg_addr != NULL)
+ {
+ if ((shmdt(m.seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", m.seg_addr);
+ m.seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (m.shmem != NULL)
+ {
+ if (munmap(m.shmem, m.shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ m.shmem, m.shmem_size);
+ m.shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 2704e80b3a7..1965b2d3eb4 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2fa045e6b0f..8b38e985327 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size)
* required.
*/
Size
-CalculateShmemSize(int *num_semaphores)
+CalculateShmemSize(int *num_semaphores, int shmem_segment)
{
Size size;
int numSemas;
@@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
-
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
-
- /*
- * Create semaphores
- */
- PGReserveSemaphores(numSemas);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ /* Compute the size of the shared-memory block */
+ size = CalculateShmemSize(&numSemas, segment);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(size, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Create semaphores
+ */
+ PGReserveSemaphores(numSemas, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -363,7 +368,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas);
+ size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index a0770e86796..f185ed28f95 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -76,19 +76,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[ANON_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -101,9 +101,17 @@ Datum pg_numa_available(PG_FUNCTION_ARGS);
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -114,7 +122,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -123,9 +137,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -150,11 +164,17 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
@@ -184,6 +204,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -203,22 +229,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -236,6 +262,12 @@ ShmemAllocRaw(Size size, Size *allocated_size)
*/
void *
ShmemAllocUnlocked(Size size)
+{
+ return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -246,19 +278,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
errmsg("out of shared memory (%zu bytes requested)",
size)));
- ShmemSegHdr->freeoffset = newFree;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -273,7 +305,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -334,6 +372,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
int64 max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -350,9 +400,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -385,6 +435,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -393,7 +450,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -416,7 +473,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -433,8 +490,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -459,7 +516,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -478,14 +535,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -542,10 +598,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
+ /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
+ values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
named_allocated += ent->allocated_size;
@@ -557,15 +614,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output shared memory allocated but not counted via the shmem index */
values[0] = CStringGetTextDatum("<anonymous>");
nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
+ values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
+ values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
values[3] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
@@ -630,7 +687,12 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ PGShmemHeader *shmhdr = Segments[segment].ShmemSegHdr;
+ shm_total_page_count += (shmhdr->totalsize / os_page_size) + 1;
+ }
+
page_ptrs = palloc0(sizeof(void *) * shm_total_page_count);
pages_status = palloc(sizeof(int) * shm_total_page_count);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index b017880f5e4..c25dd13b63a 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -612,12 +614,15 @@ LWLockNewTrancheId(const char *name)
/*
* We use the ShmemLock spinlock to protect LWLockCounter and
* LWLockTrancheNames.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
*/
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
if (*LWLockCounter - LWTRANCHE_FIRST_USER_DEFINED >= MAX_NAMED_TRANCHES)
{
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
ereport(ERROR,
(errmsg("maximum number of tranches already registered"),
errdetail("No more than %d tranches may be registered.",
@@ -628,7 +633,7 @@ LWLockNewTrancheId(const char *name)
LocalLWLockCounter = *LWLockCounter;
strlcpy(LWLockTrancheNames[result - LWTRANCHE_FIRST_USER_DEFINED], name, NAMEDATALEN);
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
@@ -750,9 +755,9 @@ GetLWTrancheName(uint16 trancheId)
*/
if (trancheId >= LocalLWLockCounter)
{
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
LocalLWLockCounter = *LWLockCounter;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
if (trancheId >= LocalLWLockCounter)
elog(ERROR, "tranche %d is not registered", trancheId);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..6ebda479ced 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores);
+extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h
index fa6ca35a51f..8ae9637fcd0 100644
--- a/src/include/storage/pg_sema.h
+++ b/src/include/storage/pg_sema.h
@@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore;
extern Size PGSemaphoreShmemSize(int maxSemas);
/* Module initialization (called during postmaster start or shmem reinit) */
-extern void PGReserveSemaphores(int maxSemas);
+extern void PGReserveSemaphores(int maxSemas, int shmem_segment);
/* Allocate a PGSemaphore structure with initial count 1 */
extern PGSemaphore PGSemaphoreCreate(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 5f7d4b83a60..2348c59b5a0 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,7 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define ANON_MAPPINGS 1
+
+extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -91,4 +106,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index cd683a9d2d9..910c43f54f4 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -30,15 +30,26 @@ extern PGDLLIMPORT slock_t *ShmemLock;
typedef struct PGShmemHeader PGShmemHeader; /* avoid including
* storage/pg_shmem.h here */
extern void InitShmemAccess(PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
extern void *ShmemAllocUnlocked(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
--
2.34.1
[application/x-patch] 0008-Fix-compilation-failures-from-previous-comm-20251013.patch (1.4K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/7-0008-Fix-compilation-failures-from-previous-comm-20251013.patch)
download | inline diff:
From 385af90e5bc853f56ed30dd0031c8010f7a45d71 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 20 Aug 2025 11:35:20 +0530
Subject: [PATCH 08/19] Fix compilation failures from previous commits
shm_total_page_count is used unitialized. If this variable has a random
value to start with, the final sum would be wrong.
Also include pg_shmem.h where shared memory segment macros are used.
Author: Ashutosh Bapat
---
src/backend/storage/buffer/buf_init.c | 1 +
src/backend/storage/ipc/shmem.c | 2 +-
2 files changed, 2 insertions(+), 1 deletion(-)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 5383442e213..6d703e18f8b 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -16,6 +16,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
+#include "storage/pg_shmem.h"
#include "storage/bufmgr.h"
BufferDescPadded *BufferDescriptors;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9bb73f31052..e6cb919f0fc 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -649,7 +649,7 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size os_page_size;
void **page_ptrs;
int *pages_status;
- uint64 shm_total_page_count,
+ uint64 shm_total_page_count = 0,
shm_ent_page_count,
max_nodes;
Size *nodes;
--
2.34.1
[application/x-patch] 0006-Address-space-reservation-for-shared-memory-20251013.patch (24.4K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/8-0006-Address-space-reservation-for-shared-memory-20251013.patch)
download | inline diff:
From 62a3e35f4e42bf9c586901e1f9c8f75869b0a13e Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 11:47:04 +0200
Subject: [PATCH 06/19] Address space reservation for shared memory
Currently the shared memory layout is designed to pack everything tight
together, leaving no space between mappings for resizing. Here is how it
looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
...
Make the layout more dynamic via splitting every shared memory segment
into two parts:
* An anonymous file, which actually contains shared memory content. Such
an anonymous file is created via memfd_create, it lives in memory,
behaves like a regular file and semantically equivalent to an
anonymous memory allocated via mmap with MAP_ANONYMOUS.
* A reservation mapping, which size is much larger than required shared
segment size. This mapping is created with flags PROT_NONE (which
makes sure the reserved space is not used), and MAP_NORESERVE (to not
count the reserved space against memory limits). The anonymous file is
mapped into this reservation mapping.
The resulting layout looks like this:
00400000-00490000 /path/bin/postgres
...
3f526000-3f590000 rw-p [heap]
7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
To resize a shared memory segment in this layout it's possible to use ftruncate
on the anonymous file, adjusting access permissions on the reserved space as
needed.
This approach also do not impact the actual memory usage as reported by
the kernel. Here is the output of /proc/$PID/status for the master
version with shared_buffers = 128 MB:
// Peak virtual memory size, which is described as total pages
// mapped in mm_struct. It corresponds to the mapped reserved space
// and is the only number that grows with it.
VmPeak: 2043192 kB
// Size of memory portions. It contains RssAnon + RssFile + RssShmem
VmRSS: 22908 kB
// Size of resident anonymous memory
RssAnon: 768 kB
// Size of resident file mappings
RssFile: 10364 kB
// Size of resident shmem memory (includes SysV shm, mapping of tmpfs and
// shared anonymous mappings)
RssShmem: 11776 kB
Here is the same for the patch when reserving 20GB of space:
VmPeak: 21255824 kB
VmRSS: 25020 kB
RssAnon: 768 kB
RssFile: 10812 kB
RssShmem: 13440 kB
Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched withing
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
20770816 (~19.8 MB)
To control the amount of space reserved a new GUC max_available_memory
is introduced. Ideally it should be based on the maximum available
memory, hense the name.
There are also few unrelated advantages of using anon files:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps.
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process.
The downside is that memfd_create is Linux specific.
---
src/backend/port/sysv_shmem.c | 290 ++++++++++++++++++----
src/backend/port/win32_shmem.c | 2 +-
src/backend/storage/ipc/ipci.c | 5 +-
src/backend/storage/ipc/shmem.c | 2 +-
src/backend/utils/init/globals.c | 1 +
src/backend/utils/misc/guc_parameters.dat | 12 +
src/include/miscadmin.h | 1 +
src/include/portability/mem.h | 2 +-
src/include/storage/pg_shmem.h | 5 +-
9 files changed, 260 insertions(+), 60 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 56af0231d24..363ddfd1fca 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -97,10 +97,12 @@ void *UsedShmemSegAddr = NULL;
typedef struct AnonymousMapping
{
int shmem_segment;
- Size shmem_size; /* Size of the mapping */
+ Size shmem_size; /* Size of the actually used memory */
+ Size shmem_reserved; /* Size of the reserved mapping */
Pointer shmem; /* Pointer to the start of the mapped memory */
Pointer seg_addr; /* SysV shared memory for the header */
unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
} AnonymousMapping;
static AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -108,6 +110,49 @@ static AnonymousMapping Mappings[ANON_MAPPINGS];
/* Keeps track of used mapping segments */
static int next_free_segment = 0;
+/*
+ * Anonymous mapping layout we use looks like this:
+ *
+ * 00400000-00c2a000 r-xp /bin/postgres
+ * ...
+ * 3f526000-3f590000 rw-p [heap]
+ * 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted)
+ * 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted)
+ * 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
+ * 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
+ * ...
+ *
+ * We need to place shared memory mappings in such a way, that there will be
+ * gaps between them in the address space. Those gaps have to be large enough
+ * to resize the mapping up to certain size, without counting towards the total
+ * memory consumption.
+ *
+ * To achieve this, for each shared memory segment we first create an anonymous
+ * file of specified size using memfd_create, which will accomodate actual
+ * shared memory mapping content. It is represented by the first /memfd:main
+ * with rw permissions. Then we create a mapping for this file using mmap, with
+ * size much larger than required and flags PROT_NONE (allows to make sure the
+ * reserved space will not be used) and MAP_NORESERVE (prevents the space from
+ * being counted against memory limits). The mapping serves as an address space
+ * reservation, into which shared memory segment can be extended and is
+ * represented by the second /memfd:main with no permissions.
+ *
+ * The reserved space for each segment is calculated as a fraction of the total
+ * reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
+ * array.
+ */
+static double SHMEM_RESIZE_RATIO[1] = {
+ 1.0, /* MAIN_SHMEM_SLOT */
+};
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -503,19 +548,20 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* hugepage sizes, we might want to think about more invasive strategies,
* such as increasing shared_buffers to absorb the extra space.
*
- * Returns the (real, assumed or config provided) page size into
- * *hugepagesize, and the hugepage-related mmap flags to use into
- * *mmap_flags if requested by the caller. If huge pages are not supported,
- * *hugepagesize and *mmap_flags are set to 0.
+ * Returns the (real, assumed or config provided) page size into *hugepagesize,
+ * the hugepage-related mmap and memfd flags to use into *mmap_flags and
+ * *memfd_flags if requested by the caller. If huge pages are not supported,
+ * *hugepagesize, *mmap_flags and *memfd_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -574,6 +620,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -584,7 +631,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -593,6 +649,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -600,6 +658,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -625,72 +685,90 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
* Creates an anonymous mmap()ed shared memory segment.
*
* This function will modify mapping size to the actual size of the allocation,
- * if it ends up allocating a segment that is larger than requested.
+ * if it ends up allocating a segment that is larger than requested. If needed,
+ * it also rounds up the mapping reserved size to be a multiple of huge page
+ * size.
+ *
+ * Note that we do not fallback from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
*/
static void
CreateAnonymousSegment(AnonymousMapping *mapping)
{
Size allocsize = mapping->shmem_size;
void *ptr = MAP_FAILED;
- int mmap_errno = 0;
+ int save_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
+
+ elog(DEBUG1, "segment[%s]: size %zu, reserved %zu",
+ MappingName(mapping->shmem_segment), mapping->shmem_size,
+ mapping->shmem_reserved);
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the request size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
if (allocsize % hugepagesize != 0)
allocsize += hugepagesize - (allocsize % hugepagesize);
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- {
- DebugMappings();
- elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- MappingName(mapping->shmem_segment), allocsize);
- }
+ /*
+ * The reserved space is multiple of BLCKSZ. We know the huge page
+ * size, round up the reserved space to it.
+ */
+ mapping->shmem_reserved = mapping->shmem_reserved + hugepagesize -
+ (mapping->shmem_reserved % hugepagesize);
+
+ /* Verify that the new size is withing the reserved boundaries */
+ if (mapping->shmem_reserved < mapping->shmem_size)
+ ereport(ERROR,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved"),
+ errhint("You may need to increase \"max_available_memory\".")));
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
}
#endif
/*
- * Report whether huge pages are in use. This needs to be tracked before
- * the second mmap() call if attempting to use huge pages failed
- * previously.
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
*/
- SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
- PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment),
+ memfd_flags);
- if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
+ /*
+ * Specify the segment file size using allocsize, which contains
+ * potentially modified value.
+ */
+ if(ftruncate(mapping->segment_fd, allocsize) == -1)
{
- /*
- * Use the original size, not the rounded-up value, when falling back
- * to non-huge pages.
- */
- allocsize = mapping->shmem_size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
- }
+ save_errno = errno;
- if (ptr == MAP_FAILED)
- {
- errno = mmap_errno;
DebugMappings();
+ close(mapping->segment_fd);
+
+ errno = save_errno;
ereport(FATAL,
- (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ (errmsg("segment[%s]: could not truncate anonymous file: %m",
MappingName(mapping->shmem_segment)),
- (mmap_errno == ENOMEM) ?
+ (save_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
"swap space, or huge pages. To reduce the request size "
@@ -700,10 +778,112 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
allocsize) : 0));
}
+ elog(DEBUG1, "segment[%s]: mmap(%zu)",
+ MappingName(mapping->shmem_segment), allocsize);
+
+ /*
+ * Create a reservation mapping.
+ */
+ ptr = mmap(NULL, mapping->shmem_reserved, PROT_NONE,
+ mmap_flags | MAP_NORESERVE, mapping->segment_fd, 0);
+ save_errno = errno;
+
+ if (ptr == MAP_FAILED)
+ {
+ DebugMappings();
+
+ errno = save_errno;
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment))));
+ }
+
+ /* Make the memory accessible */
+ if(mprotect(ptr, allocsize, PROT_READ | PROT_WRITE) == -1)
+ {
+ save_errno = errno;
+ DebugMappings();
+
+ errno = save_errno;
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not mprotect anonymous shared memory: %m",
+ MappingName(mapping->shmem_segment))));
+ }
+
mapping->shmem = ptr;
mapping->shmem_size = allocsize;
}
+/*
+ * PrepareHugePages
+ *
+ * Figure out if there are enough huge pages to allocate all shared memory
+ * segments, and report that information via huge_pages_status and
+ * huge_pages_on. It needs to be called before creating shared memory segments.
+ *
+ * It is necessary to maintain the same semantic (simple on/off) for
+ * huge_pages_status, even if there are multiple shared memory segments: all
+ * segments either use huge pages or not, there is no mix of segments with
+ * different page size. The latter might be actually beneficial, in particular
+ * because only some segments may require large amount of memory, but for now
+ * we go with a simple solution.
+ */
+void
+PrepareHugePages()
+{
+ void *ptr = MAP_FAILED;
+
+ /* Reset to handle reinitialization */
+ next_free_segment = 0;
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ */
+ for(int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
+ int numSemas;
+ Size segment_size = CalculateShmemSize(&numSemas, segment);
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
+ }
+#endif
+
+ /*
+ * Report whether huge pages are in use. This needs to be tracked before
+ * creating shared memory segments.
+ */
+ SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
+}
+
/*
* AnonymousShmemDetach --- detach from an anonymous mmap'd block
* (called as an on_shmem_exit callback, hence funny argument list)
@@ -746,7 +926,7 @@ PGSharedMemoryCreate(Size size,
void *memAddress;
PGShmemHeader *hdr;
struct stat statbuf;
- Size sysvsize;
+ Size sysvsize, total_reserved;
AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
@@ -760,14 +940,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -776,8 +948,16 @@ PGSharedMemoryCreate(Size size,
/* Room for a header? */
Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+
+ /* Prepare the mapping information */
mapping->shmem_size = size;
mapping->shmem_segment = next_free_segment;
+ total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
+ mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+
+ /* Round up to be a multiple of BLCKSZ */
+ mapping->shmem_reserved = mapping->shmem_reserved + BLCKSZ -
+ (mapping->shmem_reserved % BLCKSZ);
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..732fedee87e 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 8b38e985327..b60f7ef9ce2 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -206,6 +206,9 @@ CreateSharedMemoryAndSemaphores(void)
Assert(!IsUnderPostmaster);
+ /* Decide if we use huge pages or regular size pages */
+ PrepareHugePages();
+
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
/* Compute the size of the shared-memory block */
@@ -377,7 +380,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index f185ed28f95..9bb73f31052 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -815,7 +815,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index d31cb45a058..90d3feb547c 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffers = 16384;
+int MaxAvailableMemory = 524288;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index b176d5130e4..cff8bb815f9 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -1114,6 +1114,18 @@
max => 'INT_MAX / 2',
},
+# TODO: should this be PGC_POSTMASTER?
+{ name => "max_available_memory", type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
+ short_desc => 'Sets the upper limit for the shared_buffers value.',
+ long_desc => 'Shared memory could be resized at runtime, this parameters sets the upper limit for it, beyond which resizing would not be supported. Normally this value would be the same as the total available memory.',
+ flags => 'GUC_UNIT_BLOCKS',
+ variable => 'MaxAvailableMemory',
+ boot_val => '524288',
+ min => '16',
+ max => 'INT_MAX / 2',
+},
+
+
{ name => 'vacuum_buffer_usage_limit', type => 'int', context => 'PGC_USERSET', group => 'RESOURCES_MEM',
short_desc => 'Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum.',
flags => 'GUC_UNIT_KB',
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 1bef98471c3..a0c37a7749e 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int MaxAvailableMemory;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 2348c59b5a0..79b0b1ef9eb 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -61,6 +61,7 @@ extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT int MaxAvailableMemory;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -104,7 +105,9 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+void PrepareHugePages(void);
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
--
2.34.1
[application/x-patch] 0007-Introduce-multiple-shmem-segments-for-share-20251013.patch (11.7K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/9-0007-Introduce-multiple-shmem-segments-for-share-20251013.patch)
download | inline diff:
From 3a419573e0a5fe41e6aa1a2530c0c661d8a7eaa8 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 11:22:02 +0200
Subject: [PATCH 07/19] Introduce multiple shmem segments for shared buffers
Add more shmem segments to split shared buffers into following chunks:
* BUFFERS_SHMEM_SEGMENT: contains buffer blocks
* BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors
* BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers
* CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids
* STRATEGY_SHMEM_SEGMENT: contains buffer strategy status
Size of the corresponding shared data directly depends on NBuffers,
meaning that if we would like to change NBuffers, they have to be
resized correspondingly. Placing each of them in a separate shmem
segment allows to achieve that.
There are some asumptions made about each of shmem segments upper size
limit. The buffer blocks have the largest, while the rest claim less
extra room for resize. Ideally those limits have to be deduced from the
maximum allowed shared memory.
---
src/backend/port/sysv_shmem.c | 24 +++++++-
src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++---------
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipci.c | 2 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/pg_shmem.h | 24 +++++++-
7 files changed, 105 insertions(+), 37 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 363ddfd1fca..dac011b766b 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -139,10 +139,18 @@ static int next_free_segment = 0;
*
* The reserved space for each segment is calculated as a fraction of the total
* reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
- * array.
+ * array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take up to 60% of the whole
+ * space when resizing, based on the fact that it most likely will be the main
+ * consumer of this memory. Those numbers are pulled out of thin air for now,
+ * makes sense to evaluate them more precise.
*/
-static double SHMEM_RESIZE_RATIO[1] = {
- 1.0, /* MAIN_SHMEM_SLOT */
+static double SHMEM_RESIZE_RATIO[6] = {
+ 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.6, /* BUFFERS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
+ 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
+ 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
+ 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -167,6 +175,16 @@ MappingName(int shmem_segment)
{
case MAIN_SHMEM_SEGMENT:
return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
default:
return "unknown";
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6fd3a6bbac5..5383442e213 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +77,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +102,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -147,33 +151,54 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(int shmem_segment)
{
Size size = 0;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == MAIN_SHMEM_SEGMENT)
+ return size;
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
+ {
+ /* size of buffer descriptors */
+ size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ /* to allow aligning buffer descriptors */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of data pages, plus alignment padding */
+ size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ }
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
+ if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
+ {
+ /* size of stuff controlled by freelist.c */
+ size = add_size(size, StrategyShmemSize());
+ }
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
+ {
+ /* size of I/O condition variables */
+ size = add_size(size, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ /* to allow aligning the above */
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ }
+
+ if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
+ {
+ /* size of checkpoint sort array in bufmgr.c */
+ size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ }
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 1f6e215a2ca..18a78967138 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -25,6 +25,7 @@
#include "funcapi.h"
#include "storage/buf_internals.h"
#include "storage/lwlock.h"
+#include "storage/pg_shmem.h"
#include "utils/rel.h"
#include "utils/builtins.h"
@@ -64,10 +65,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 7fe34d3ef4c..299f6aa8e7e 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -418,9 +419,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index b60f7ef9ce2..2dbd81afc87 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(shmem_segment));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 3f37b294af6..a222747b803 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -319,7 +319,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(int);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 79b0b1ef9eb..a7b275b4db9 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -52,7 +52,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 1
+#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
@@ -109,7 +109,29 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
/* The main segment, contains everything except buffer blocks and related data. */
#define MAIN_SHMEM_SEGMENT 0
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
#endif /* PG_SHMEM_H */
--
2.34.1
[application/x-patch] 0010-WIP-Monitoring-views-20251013.patch (10.5K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/10-0010-WIP-Monitoring-views-20251013.patch)
download | inline diff:
From a249feb0e7654865f53f3853c310a5cec58e185e Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 20 Aug 2025 10:55:27 +0530
Subject: [PATCH 10/19] WIP: Monitoring views
Modifies pg_shmem_allocations to report shared memory segment as well.
Adds pg_shmem_segments to report shared memory segment information.
TODO:
This commit should be merged with the earlier commit introducing
multiple shared memory segments.
Author: Ashutosh Bapat
---
doc/src/sgml/system-views.sgml | 9 +++
src/backend/catalog/system_views.sql | 7 +++
src/backend/storage/ipc/shmem.c | 90 ++++++++++++++++++++++------
src/include/catalog/pg_proc.dat | 12 +++-
src/include/storage/pg_shmem.h | 1 -
src/include/storage/shmem.h | 1 +
src/test/regress/expected/rules.out | 10 +++-
7 files changed, 108 insertions(+), 22 deletions(-)
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 8f3e2741051..bc70a3ee6c9 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -4233,6 +4233,15 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
</para></entry>
</row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>segment</structfield> <type>text</type>
+ </para>
+ <para>
+ The name of the shared memory segment concerning the allocation.
+ </para></entry>
+ </row>
+
<row>
<entry role="catalog_table_entry"><para role="column_definition">
<structfield>off</structfield> <type>int8</type>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index c7240250c07..94a2b5a9a67 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -668,6 +668,13 @@ GRANT SELECT ON pg_shmem_allocations TO pg_read_all_stats;
REVOKE EXECUTE ON FUNCTION pg_get_shmem_allocations() FROM PUBLIC;
GRANT EXECUTE ON FUNCTION pg_get_shmem_allocations() TO pg_read_all_stats;
+CREATE VIEW pg_shmem_segments AS
+ SELECT * FROM pg_get_shmem_segments();
+
+REVOKE ALL ON pg_shmem_segments FROM PUBLIC;
+GRANT SELECT ON pg_shmem_segments TO pg_read_all_stats;
+REVOKE EXECUTE ON FUNCTION pg_get_shmem_segments() FROM PUBLIC;
+GRANT EXECUTE ON FUNCTION pg_get_shmem_segments() TO pg_read_all_stats;
CREATE VIEW pg_shmem_allocations_numa AS
SELECT * FROM pg_get_shmem_allocations_numa();
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 90c21a97225..9499f332e77 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -531,6 +531,7 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
result->size = size;
result->allocated_size = allocated_size;
result->location = structPtr;
+ result->shmem_segment = shmem_segment;
}
LWLockRelease(ShmemIndexLock);
@@ -582,13 +583,14 @@ mul_size(Size s1, Size s2)
Datum
pg_get_shmem_allocations(PG_FUNCTION_ARGS)
{
-#define PG_GET_SHMEM_SIZES_COLS 4
+#define PG_GET_SHMEM_SIZES_COLS 5
ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
HASH_SEQ_STATUS hstat;
ShmemIndexEnt *ent;
- Size named_allocated = 0;
+ Size named_allocated[ANON_MAPPINGS] = {0};
Datum values[PG_GET_SHMEM_SIZES_COLS];
bool nulls[PG_GET_SHMEM_SIZES_COLS];
+ int i;
InitMaterializedSRF(fcinfo, 0);
@@ -598,33 +600,42 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
/* output all allocated entries */
memset(nulls, 0, sizeof(nulls));
- /* XXX: take all shared memory segments into account. */
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr);
- values[2] = Int64GetDatum(ent->size);
- values[3] = Int64GetDatum(ent->allocated_size);
- named_allocated += ent->allocated_size;
+ values[1] = CStringGetTextDatum(MappingName(ent->shmem_segment));
+ values[2] = Int64GetDatum((char *) ent->location - (char *) Segments[ent->shmem_segment].ShmemSegHdr);
+ values[3] = Int64GetDatum(ent->size);
+ values[4] = Int64GetDatum(ent->allocated_size);
+ named_allocated[ent->shmem_segment] += ent->allocated_size;
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
values, nulls);
}
/* output shared memory allocated but not counted via the shmem index */
- values[0] = CStringGetTextDatum("<anonymous>");
- nulls[1] = true;
- values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ for (i = 0; i < ANON_MAPPINGS; i++)
+ {
+ values[0] = CStringGetTextDatum("<anonymous>");
+ values[1] = CStringGetTextDatum(MappingName(i));
+ nulls[2] = true;
+ values[3] = Int64GetDatum(Segments[i].ShmemSegHdr->freeoffset - named_allocated[i]);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
/* output as-of-yet unused shared memory */
- nulls[0] = true;
- values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
- nulls[1] = false;
- values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ memset(nulls, 0, sizeof(nulls));
+
+ for (i = 0; i < ANON_MAPPINGS; i++)
+ {
+ nulls[0] = true;
+ values[1] = CStringGetTextDatum(MappingName(i));
+ values[2] = Int64GetDatum(Segments[i].ShmemSegHdr->freeoffset);
+ values[3] = Int64GetDatum(Segments[i].ShmemSegHdr->totalsize - Segments[i].ShmemSegHdr->freeoffset);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
LWLockRelease(ShmemIndexLock);
@@ -825,3 +836,46 @@ pg_numa_available(PG_FUNCTION_ARGS)
{
PG_RETURN_BOOL(pg_numa_init() != -1);
}
+
+/* SQL SRF showing shared memory segments */
+Datum
+pg_get_shmem_segments(PG_FUNCTION_ARGS)
+{
+#define PG_GET_SHMEM_SEGS_COLS 6
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ Datum values[PG_GET_SHMEM_SEGS_COLS];
+ bool nulls[PG_GET_SHMEM_SEGS_COLS];
+ int i;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* output all allocated entries */
+ for (i = 0; i < ANON_MAPPINGS; i++)
+ {
+ PGShmemHeader *shmhdr = Segments[i].ShmemSegHdr;
+ AnonymousMapping *segmapping = &Mappings[i];
+ int j;
+
+ if (shmhdr == NULL)
+ {
+ for (j = 0; j < PG_GET_SHMEM_SEGS_COLS; j++)
+ nulls[j] = true;
+ }
+ else
+ {
+ memset(nulls, 0, sizeof(nulls));
+ values[0] = Int32GetDatum(i);
+ values[1] = CStringGetTextDatum(MappingName(i));
+ values[2] = Int64GetDatum(shmhdr->totalsize);
+ values[3] = Int64GetDatum(shmhdr->freeoffset);
+ values[4] = Int64GetDatum(segmapping->shmem_size);
+ values[5] = Int64GetDatum(segmapping->shmem_reserved);
+ }
+
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
+
+ return (Datum) 0;
+}
+
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index e631323a325..8f1d0b7c031 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8576,8 +8576,8 @@
{ oid => '5052', descr => 'allocations from the main shared memory segment',
proname => 'pg_get_shmem_allocations', prorows => '50', proretset => 't',
provolatile => 'v', prorettype => 'record', proargtypes => '',
- proallargtypes => '{text,int8,int8,int8}', proargmodes => '{o,o,o,o}',
- proargnames => '{name,off,size,allocated_size}',
+ proallargtypes => '{text,text,int8,int8,int8}', proargmodes => '{o,o,o,o,o}',
+ proargnames => '{name,segment,off,size,allocated_size}',
prosrc => 'pg_get_shmem_allocations' },
{ oid => '4099', descr => 'Is NUMA support available?',
@@ -8600,6 +8600,14 @@
proargmodes => '{o,o,o}', proargnames => '{name,type,size}',
prosrc => 'pg_get_dsm_registry_allocations' },
+# shared memory segments
+{ oid => '5101', descr => 'shared memory segments',
+ proname => 'pg_get_shmem_segments', prorows => '6', proretset => 't',
+ provolatile => 'v', prorettype => 'record', proargtypes => '',
+ proallargtypes => '{int4,text,int8,int8,int8,int8}', proargmodes => '{o,o,o,o,o,o}',
+ proargnames => '{id,name,size,freeoffset,mapping_size,mapping_reserved_size}',
+ prosrc => 'pg_get_shmem_segments' },
+
# buffer lookup table
{ oid => '5102',
descr => 'shared buffer lookup table',
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index a1fa6b43fe3..715f6acb5dd 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -69,7 +69,6 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
-
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 910c43f54f4..64ff5a286ba 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -71,6 +71,7 @@ typedef struct
void *location; /* location in shared mem */
Size size; /* # bytes requested for the structure */
Size allocated_size; /* # bytes actually allocated */
+ int shmem_segment; /* segment in which the structure is allocated */
} ShmemIndexEnt;
#endif /* SHMEM_H */
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 83f566d3218..60c08081b69 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1772,14 +1772,22 @@ pg_shadow| SELECT pg_authid.rolname AS usename,
LEFT JOIN pg_db_role_setting s ON (((pg_authid.oid = s.setrole) AND (s.setdatabase = (0)::oid))))
WHERE pg_authid.rolcanlogin;
pg_shmem_allocations| SELECT name,
+ segment,
off,
size,
allocated_size
- FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size);
+ FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, segment, off, size, allocated_size);
pg_shmem_allocations_numa| SELECT name,
numa_node,
size
FROM pg_get_shmem_allocations_numa() pg_get_shmem_allocations_numa(name, numa_node, size);
+pg_shmem_segments| SELECT id,
+ name,
+ size,
+ freeoffset,
+ mapping_size,
+ mapping_reserved_size
+ FROM pg_get_shmem_segments() pg_get_shmem_segments(id, name, size, freeoffset, mapping_size, mapping_reserved_size);
pg_stat_activity| SELECT s.datid,
d.datname,
s.pid,
--
2.34.1
[application/x-patch] 0009-Refactor-CalculateShmemSize-20251013.patch (22.4K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/11-0009-Refactor-CalculateShmemSize-20251013.patch)
download | inline diff:
From 07377b5f6722dcfd60b91458ca03cef8a5230e4c Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 21 Aug 2025 11:56:09 +0530
Subject: [PATCH 09/19] Refactor CalculateShmemSize()
This function calls many functions which return the amount of shared
memory required for different shared memory data structures. Up until
now, the returned total of these sizes was used to create a single
shared memory segment. But starting the previous patch, we create
multiple shared memory segments each of which contain one shared memory
structure related to shared buffers and one main memory segment
containing rest of the structures. Since CalculateShmemSize() is called
for every shared memory segment, and its return value is added to the
memory required for all the shared memory segments, we end up allocating
more memory than required.
Instead, CalculateShmemSize() is called only once. Each of its callees
are expected to a. return the size required from the main segment b. add
sizes to the AnonymousMappings corresponding to the other memory
segments.
For individual modules to add memory to their respective
AnonymousMappings, we need to know the different mappings upfront. Hence
ANON_MAPPINGS replaces next_free_segment.
TODOs:
1. This change however requires that the AnonymousMappings array and
macros defining identifiers of each of the segments be
platform-independent. This patch doesn't achieve that goal for all the
platforms for example windows. We need to fix that.
2. If postgres is invoked with -C shared_memory_size, it reports 0.
That's because it report the GUC values before share memory sizes are
set in AnonymousMappings. Fix that too.
3. Eliminate this assymetry in CalculateShmemSize(). See TODO in
prologue of CalculateShmemSize().
4. This is one way to avoid requesting more memory in each segment. But
there may be other ways to design CalculateShmemSize(). Need to think
and implement it better.
Author: Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 48 ++++++--------------
src/backend/port/win32_shmem.c | 7 +--
src/backend/postmaster/postmaster.c | 14 +++---
src/backend/storage/buffer/buf_init.c | 55 ++++++++---------------
src/backend/storage/ipc/ipci.c | 65 ++++++++++++++++++++++-----
src/backend/storage/ipc/shmem.c | 8 ++--
src/backend/tcop/postgres.c | 14 +++---
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_shmem.h | 17 ++++++-
10 files changed, 125 insertions(+), 107 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index dac011b766b..b85911bdfc4 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -94,21 +94,7 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-typedef struct AnonymousMapping
-{
- int shmem_segment;
- Size shmem_size; /* Size of the actually used memory */
- Size shmem_reserved; /* Size of the reserved mapping */
- Pointer shmem; /* Pointer to the start of the mapped memory */
- Pointer seg_addr; /* SysV shared memory for the header */
- unsigned long seg_id; /* IPC key */
- int segment_fd; /* fd for the backing anon file */
-} AnonymousMapping;
-
-static AnonymousMapping Mappings[ANON_MAPPINGS];
-
-/* Keeps track of used mapping segments */
-static int next_free_segment = 0;
+AnonymousMapping Mappings[ANON_MAPPINGS];
/*
* Anonymous mapping layout we use looks like this:
@@ -168,7 +154,7 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
-static const char*
+const char*
MappingName(int shmem_segment)
{
switch (shmem_segment)
@@ -193,7 +179,7 @@ MappingName(int shmem_segment)
static void
DebugMappings()
{
- for(int i = 0; i < next_free_segment; i++)
+ for(int i = 0; i < ANON_MAPPINGS; i++)
{
AnonymousMapping m = Mappings[i];
elog(DEBUG1, "Mapping[%s]: addr %p, size %zu",
@@ -851,9 +837,6 @@ PrepareHugePages()
{
void *ptr = MAP_FAILED;
- /* Reset to handle reinitialization */
- next_free_segment = 0;
-
/* Complain if hugepages demanded but we can't possibly support them */
#if !defined(MAP_HUGETLB)
if (huge_pages == HUGE_PAGES_ON)
@@ -876,8 +859,7 @@ PrepareHugePages()
*/
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
- int numSemas;
- Size segment_size = CalculateShmemSize(&numSemas, segment);
+ Size segment_size = Mappings[segment].shmem_req_size;
if (segment_size % hugepagesize != 0)
segment_size += hugepagesize - (segment_size % hugepagesize);
@@ -909,7 +891,7 @@ PrepareHugePages()
static void
AnonymousShmemDetach(int status, Datum arg)
{
- for(int i = 0; i < next_free_segment; i++)
+ for(int i = 0; i < ANON_MAPPINGS; i++)
{
AnonymousMapping m = Mappings[i];
@@ -927,7 +909,7 @@ AnonymousShmemDetach(int status, Datum arg)
/*
* PGSharedMemoryCreate
*
- * Create a shared memory segment of the given size and initialize its
+ * Create a shared memory segment for the given mapping and initialize its
* standard header. Also, register an on_shmem_exit callback to release
* the storage.
*
@@ -937,7 +919,7 @@ AnonymousShmemDetach(int status, Datum arg)
* postmaster or backend.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(AnonymousMapping *mapping,
PGShmemHeader **shim)
{
IpcMemoryKey NextShmemSegID;
@@ -945,7 +927,6 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize, total_reserved;
- AnonymousMapping *mapping = &Mappings[next_free_segment];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -965,13 +946,12 @@ PGSharedMemoryCreate(Size size,
errmsg("huge pages not supported with the current \"shared_memory_type\" setting")));
/* Room for a header? */
- Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ Assert(mapping->shmem_req_size > MAXALIGN(sizeof(PGShmemHeader)));
/* Prepare the mapping information */
- mapping->shmem_size = size;
- mapping->shmem_segment = next_free_segment;
+ mapping->shmem_size = mapping->shmem_req_size;
total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
- mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[next_free_segment];
+ mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[mapping->shmem_segment];
/* Round up to be a multiple of BLCKSZ */
mapping->shmem_reserved = mapping->shmem_reserved + BLCKSZ -
@@ -982,8 +962,6 @@ PGSharedMemoryCreate(Size size,
/* On success, mapping data will be modified. */
CreateAnonymousSegment(mapping);
- next_free_segment++;
-
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -992,7 +970,7 @@ PGSharedMemoryCreate(Size size,
}
else
{
- sysvsize = size;
+ sysvsize = mapping->shmem_req_size;
/* huge pages are only available with mmap */
SetConfigOption("huge_pages_status", "off",
@@ -1005,7 +983,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino + next_free_segment;
+ NextShmemSegID = statbuf.st_ino + mapping->shmem_segment;
for (;;)
{
@@ -1214,7 +1192,7 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- for(int i = 0; i < next_free_segment; i++)
+ for(int i = 0; i < ANON_MAPPINGS; i++)
{
AnonymousMapping m = Mappings[i];
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 732fedee87e..1db07ff65d3 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -204,7 +204,7 @@ EnableLockPagesPrivilege(int elevel)
* standard header.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(AnonymousMapping *mapping,
PGShmemHeader **shim)
{
void *memAddress;
@@ -216,7 +216,7 @@ PGSharedMemoryCreate(Size size,
DWORD size_high;
DWORD size_low;
SIZE_T largePageSize = 0;
- Size orig_size = size;
+ Size size = mapping->shmem_req_size;
DWORD flProtect = PAGE_READWRITE;
DWORD desiredAccess;
@@ -304,7 +304,7 @@ retry:
* Use the original size, not the rounded-up value, when
* falling back to non-huge pages.
*/
- size = orig_size;
+ size = mapping->shmem_req_size;
flProtect = PAGE_READWRITE;
goto retry;
}
@@ -391,6 +391,7 @@ retry:
hdr->totalsize = size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
hdr->dsm_control = 0;
+ mapping->shmem_size = size;
/* Save info for possible future use */
UsedShmemSegAddr = memAddress;
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index e1d643b013d..b59d20b4ac2 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -963,13 +963,6 @@ PostmasterMain(int argc, char *argv[])
*/
process_shmem_requests();
- /*
- * Now that loadable modules have had their chance to request additional
- * shared memory, determine the value of any runtime-computed GUCs that
- * depend on the amount of shared memory required.
- */
- InitializeShmemGUCs();
-
/*
* Now that modules have been loaded, we can process any custom resource
* managers specified in the wal_consistency_checking GUC.
@@ -1005,6 +998,13 @@ PostmasterMain(int argc, char *argv[])
*/
CreateSharedMemoryAndSemaphores();
+ /*
+ * Now that loadable modules have had their chance to request additional
+ * shared memory, determine the value of any runtime-computed GUCs that
+ * depend on the amount of shared memory required.
+ */
+ InitializeShmemGUCs();
+
/*
* Estimate number of openable files. This must happen after setting up
* semaphores, because on some platforms semaphores count as open files.
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6d703e18f8b..6f148d1d80b 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -158,48 +158,31 @@ BufferManagerShmemInit(void)
* data.
*/
Size
-BufferManagerShmemSize(int shmem_segment)
+BufferManagerShmemSize(void)
{
- Size size = 0;
+ size_t size;
- if (shmem_segment == MAIN_SHMEM_SEGMENT)
- return size;
+ /* size of buffer descriptors, plus alignment padding */
+ size = add_size(0, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ Mappings[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_req_size = size;
- if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT)
- {
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
- size = add_size(size, PG_CACHE_LINE_SIZE);
- }
+ /* size of data pages, plus alignment padding */
+ size = add_size(0, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ Mappings[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
- if (shmem_segment == BUFFERS_SHMEM_SEGMENT)
- {
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
- }
+ /* size of stuff controlled by freelist.c */
+ Mappings[STRATEGY_SHMEM_SEGMENT].shmem_req_size = StrategyShmemSize();
- if (shmem_segment == STRATEGY_SHMEM_SEGMENT)
- {
- /* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
- }
+ /* size of I/O condition variables, plus alignment padding */
+ size = add_size(0, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ Mappings[BUFFER_IOCV_SHMEM_SEGMENT].shmem_req_size = size;
- if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT)
- {
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
- size = add_size(size, PG_CACHE_LINE_SIZE);
- }
-
- if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT)
- {
- /* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
- }
+ /* size of checkpoint sort array in bufmgr.c */
+ Mappings[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
return size;
}
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2dbd81afc87..2cd278449f0 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -84,9 +84,23 @@ RequestAddinShmemSpace(Size size)
*
* If num_semaphores is not NULL, it will be set to the number of semaphores
* required.
+ *
+ * TODO: Right now the minions of this function return the size of shared memory
+ * required in the main shared memory segment but add sizes required from other
+ * segments in the respective mappings. I think we should change this assymetry.
+ * It's only the buffer manager which adds sizes for other segments, but in
+ * future there may be others. Further the buffer manager related other segments
+ * are expected to hold only one resizable structure thus their size should be
+ * set only once when changing shared buffer pool size (i.e. when changin
+ * shared_buffers GUC). We shouldn't allow adding more structures to these
+ * segments, and thus restrict adding sizes to the corresponding mappings after
+ * the initial size is set.
+ *
+ * TODO: Also we should do something about numSemas, which is not required
+ * everywhere CalculateShmemSize is called.
*/
Size
-CalculateShmemSize(int *num_semaphores, int shmem_segment)
+CalculateShmemSize(int *num_semaphores)
{
Size size;
int numSemas;
@@ -113,7 +127,13 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize(shmem_segment));
+
+ /*
+ * Buffer manager adds estimates for memory requirements for every shared
+ * memory segment that it uses in the corresponding AnonymousMappings.
+ * Consider size required from only the main shared memory segment here.
+ */
+ size = add_size(size, BufferManagerShmemSize());
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
@@ -154,8 +174,15 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment)
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
+ /*
+ * All the shared memory allocations considered so far happen in the main
+ * shared memory segment.
+ */
+ Mappings[MAIN_SHMEM_SEGMENT].shmem_req_size = size;
+
/* might as well round it off to a multiple of a typical page size */
- size = add_size(size, 8192 - (size % 8192));
+ for (int segment = 0; segment < ANON_MAPPINGS; segment++)
+ Mappings[segment].shmem_req_size = add_size(Mappings[segment].shmem_req_size, 8192 - (Mappings[segment].shmem_req_size % 8192));
return size;
}
@@ -201,26 +228,30 @@ CreateSharedMemoryAndSemaphores(void)
{
PGShmemHeader *shim;
PGShmemHeader *seghdr;
- Size size;
int numSemas;
Assert(!IsUnderPostmaster);
+ CalculateShmemSize(&numSemas);
+
/* Decide if we use huge pages or regular size pages */
PrepareHugePages();
for(int segment = 0; segment < ANON_MAPPINGS; segment++)
{
+ AnonymousMapping *mapping = &Mappings[segment];
+
+ mapping->shmem_segment = segment;
+
/* Compute the size of the shared-memory block */
- size = CalculateShmemSize(&numSemas, segment);
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+ elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", mapping->shmem_req_size);
/*
* Create the shmem segment.
*
* XXX: Do multiple shims are needed, one per segment?
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(mapping, &shim);
/*
* Make sure that huge pages are never reported as "unknown" while the
@@ -232,9 +263,13 @@ CreateSharedMemoryAndSemaphores(void)
InitShmemAccessInSegment(seghdr, segment);
/*
- * Create semaphores
+ * Shared memory for semaphores is allocated in the main shared memory.
+ * Hence they are allocated after the main segment is created. Patch
+ * proposed at https://commitfest.postgresql.org/patch/5997/ simplifies
+ * this.
*/
- PGReserveSemaphores(numSemas, segment);
+ if (segment == MAIN_SHMEM_SEGMENT)
+ PGReserveSemaphores(numSemas, segment);
/*
* Set up shared memory allocation mechanism
@@ -357,7 +392,9 @@ CreateOrAttachShmemStructs(void)
* InitializeShmemGUCs
*
* This function initializes runtime-computed GUCs related to the amount of
- * shared memory required for the current configuration.
+ * shared memory required for the current configuration. It assumes that the
+ * memory required by the shared memory segments is already calculated and is
+ * available in AnonymousMappings.
*/
void
InitializeShmemGUCs(void)
@@ -366,12 +403,16 @@ InitializeShmemGUCs(void)
Size size_b;
Size size_mb;
Size hp_size;
- int num_semas;
+ int num_semas = ProcGlobalSemas();
+ int i;
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT);
+ size_b = 0;
+ for (i = 0; i < ANON_MAPPINGS; i++)
+ size_b = add_size(size_b, Mappings[i].shmem_req_size);
+
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index e6cb919f0fc..90c21a97225 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -178,8 +178,8 @@ ShmemAllocInSegment(Size size, int shmem_segment)
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ MappingName(shmem_segment), size)));
return newSpace;
}
@@ -286,8 +286,8 @@ ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ MappingName(shmem_segment), size)));
Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 193efeb9022..86ffe020c01 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4133,13 +4133,6 @@ PostgresSingleUserMain(int argc, char *argv[],
*/
process_shmem_requests();
- /*
- * Now that loadable modules have had their chance to request additional
- * shared memory, determine the value of any runtime-computed GUCs that
- * depend on the amount of shared memory required.
- */
- InitializeShmemGUCs();
-
/*
* Now that modules have been loaded, we can process any custom resource
* managers specified in the wal_consistency_checking GUC.
@@ -4152,6 +4145,13 @@ PostgresSingleUserMain(int argc, char *argv[],
*/
CreateSharedMemoryAndSemaphores();
+ /*
+ * Now that loadable modules have had their chance to request additional
+ * shared memory, determine the value of any runtime-computed GUCs that
+ * depend on the amount of shared memory required.
+ */
+ InitializeShmemGUCs();
+
/*
* Estimate number of openable files. This must happen after setting up
* semaphores, because on some platforms semaphores count as open files.
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index a222747b803..3f37b294af6 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -319,7 +319,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(int);
+extern Size BufferManagerShmemSize(void);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6ebda479ced..3baf418b3d1 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment);
+extern Size CalculateShmemSize(int *num_semaphores);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index a7b275b4db9..a1fa6b43fe3 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -27,6 +27,18 @@
#include "storage/dsm_impl.h"
#include "storage/spin.h"
+typedef struct AnonymousMapping
+{
+ int shmem_segment; /* TODO: Do we really need it? */
+ Size shmem_req_size; /* Required size of the segment */
+ Size shmem_size; /* Size of the actually used memory */
+ Size shmem_reserved; /* Size of the reserved mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+ unsigned long seg_id; /* IPC key */
+ int segment_fd; /* fd for the backing anon file */
+} AnonymousMapping;
+
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
int32 magic; /* magic # to identify Postgres segments */
@@ -55,6 +67,8 @@ typedef struct ShmemSegment
#define ANON_MAPPINGS 6
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
+extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
@@ -101,10 +115,11 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(Size size,
+extern PGShmemHeader *PGSharedMemoryCreate(AnonymousMapping *mapping,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
+extern const char *MappingName(int shmem_segment);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
--
2.34.1
[application/x-patch] 0012-Initial-value-of-shared_buffers-or-NBuffers-20251013.patch (3.7K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/12-0012-Initial-value-of-shared_buffers-or-NBuffers-20251013.patch)
download | inline diff:
From 132e1155b6b2ad1522087a11deb1029fa0dcdb4b Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 1 Sep 2025 15:40:41 +0530
Subject: [PATCH 12/19] Initial value of shared_buffers (or NBuffers)
The assign_hook for shared_buffers (assign_shared_buffers()) is called twice
during server startup. First time it sets the default value of shared_buffers,
followed by a second time when it sets the value specified in the configuration
file or on the command line. At those times the shared buffer pool is yet to be
initialized. Hence there is no need to keep the GUC change pending or going
through the entire process of resizing memory maps, reinitializing the shared memory
and process synchronization. Instead the given value should be assigned directly to
NBuffers, which will be used when creating the shared
memory and also when initializing the buffer pool the first time. Any changes
to shared_buffer after that will need remapping the shared memory segment and
synchronize buffer pool reinitialization across the backends.
If BufferBlocks is not initilized assign_shared_buffers() sets the given
value to NBuffers directly. Otherwise it marks the change as pending and
sets the flag pending_pm_shmem_resize so that Postmaster can start the
buffer pool reinitialization.
TODO:
1. The change depends upon the C convention that the global pointer variables
being initialized to NULL. May be initialize BufferBlocks to NULL explicitly.
2. We might think of a better way to check whether buffer pool has been
initialized or not. See comment in assign_shared_buffers().
Author: Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 42 ++++++++++++++++++++++++++---------
1 file changed, 32 insertions(+), 10 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index dc4eeeee56a..ba8613678f6 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1168,20 +1168,42 @@ ProcessBarrierShmemResize(Barrier *barrier)
}
/*
- * GUC assign hook for shared_buffers. It's recommended for an assign hook to
- * be as minimal as possible, thus we just request shared memory resize and
- * remember the previous value.
+ * GUC assign hook for shared_buffers.
+ *
+ * When setting the GUC first time after starting the server, the GUC value is
+ * changed immediately since there is not shared memory setup yet.
+ *
+ * After the shared memory is setup, changing the GUC value requires resizing and
+ * reiniatializing (at least parts of) the shared memory structures related to
+ * shared buffers. That's a long and complicated process. It's recommended for
+ * an assign hook to be as minimal as possible, thus we just request shared
+ * memory resize and remember the previous value.
*/
void
assign_shared_buffers(int newval, void *extra, bool *pending)
{
- elog(DEBUG1, "Received SIGHUP for shmem resizing");
-
- pending_pm_shmem_resize = true;
- *pending = true;
- NBuffersPending = newval;
-
- NBuffersOld = NBuffers;
+ /*
+ * TODO: If a backend joins while the buffer resizing is in progress or it
+ * reads a value of shared_buffers from configuration which is different from
+ * the value being used by existing backends, this method may not work. Need
+ * to think of a better solution.
+ */
+ if (BufferBlocks)
+ {
+ elog(DEBUG1, "bufferpool is already initialized with size = %d, reinitializing it with size = %d",
+ NBuffers, newval);
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+ NBuffersOld = NBuffers;
+ }
+ else
+ {
+ elog(DEBUG1, "initializing buffer pool with size = %d", newval);
+ NBuffers = newval;
+ *pending = false;
+ pending_pm_shmem_resize = false;
+ }
}
/*
--
2.34.1
[application/x-patch] 0013-Update-sizes-and-addresses-of-shared-memory-20251013.patch (2.5K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/13-0013-Update-sizes-and-addresses-of-shared-memory-20251013.patch)
download | inline diff:
From 9548894435757b4a542ba172490b757f46eae8fa Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 21 Aug 2025 15:44:24 +0530
Subject: [PATCH 13/19] Update sizes and addresses of shared memory mapping and
shared memory structures
Update totalsize and end address in segment and mapping: Once a shared
memory segment has been resized, the total size and end address of the
same needs to be updated in the corresponding AnonymousMapping and
Segment structure.
Update allocated_size for resized shared memory structure: Reallocating
the shared memory structure after resizing needs a bit more work. But at
least update the allocated_size as well along with the size of shared
memory structure.
Author: Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 4 ++++
src/backend/storage/ipc/shmem.c | 6 +++++-
2 files changed, 9 insertions(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index ba8613678f6..54d335b2e5d 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1021,6 +1021,8 @@ AnonymousShmemResize(void)
for(int i = 0; i < ANON_MAPPINGS; i++)
{
AnonymousMapping *m = &Mappings[i];
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmem_hdr = segment->ShmemSegHdr;
#ifdef MAP_HUGETLB
if (huge_pages_on && (m->shmem_req_size % hugepagesize != 0))
@@ -1067,6 +1069,8 @@ AnonymousShmemResize(void)
reinit = true;
m->shmem_size = m->shmem_req_size;
+ shmem_hdr->totalsize = m->shmem_size;
+ segment->ShmemEnd = m->shmem + m->shmem_size;
}
if (reinit)
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 2a197540300..0f9abf69fd5 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -504,13 +504,17 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
*
* XXX: There is an implicit assumption this can only happen in
* "resizable" segments, where only one shared structure is allowed.
- * This has to be implemented more cleanly.
+ * This has to be implemented more cleanly. Probably we should implement
+ * ShmemReallocRawInSegment functionality just to adjust the size
+ * according to alignment, return the allocated size and update the
+ * mapping offset.
*/
if (result->size != size)
{
Size delta = size - result->size;
result->size = size;
+ result->allocated_size = size;
/* Reflect size change in the shared segment */
SpinLockAcquire(Segments[shmem_segment].ShmemLock);
--
2.34.1
[application/x-patch] 0011-Allow-to-resize-shared-memory-without-resta-20251013.patch (40.8K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/14-0011-Allow-to-resize-shared-memory-without-resta-20251013.patch)
download | inline diff:
From 46a7cabd8c0e8b92c5f7515856ad1d8f11a410a7 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 14:16:55 +0200
Subject: [PATCH 11/19] Allow to resize shared memory without restart
Add assing hook for shared_buffers to resize shared memory using space,
introduced in the previous commits without requiring PostgreSQL restart.
Essentially the implementation is based on two mechanisms: a
ProcSignalBarrier is used to make sure all processes are starting the
resize procedure simultaneously, and a global Barrier is used to
coordinate after that and make sure all finished processes are waiting
for others that are in progress.
The resize process looks like this:
* The GUC assign hook sets a flag to let the Postmaster know that resize
was requested.
* Postmaster verifies the flag in the event loop, and starts the resize
by emitting a ProcSignal barrier.
* All processes, that participate in ProcSignal mechanism, begin to
process ProcSignal barrier. First a process waits until all processes
have confirmed they received the message and can start simultaneously.
* Every process recalculates shared memory size based on the new
NBuffers, adjusts its size using ftruncate and adjust reservation
permissions with mprotect. One elected process signals the postmaster
to do the same.
* When finished, every process waits on a global ShmemControl barrier,
untill all others are finished as well. This way we ensure three
stages with clear boundaries: before the resize, when all processes
use old NBuffers; during the resize, when processes have mix of old
and new NBuffers, and wait until it's done; after the resize, when all
processes use new NBuffers.
* After all processes are using new value, one of them will initialize
new shared structures (buffer blocks, descriptors, etc) as needed and
broadcast new value of NBuffers via ShmemControl in shared memory.
Other backends are waiting for this operation to finish as well. Then
the barrier is lifted and everything goes as usual.
Since resizing takes time, we need to take into account that during that time:
- New backends can be spawned. They will check status of the barrier
early during the bootstrap, and wait until everything is over to work
with the new NBuffers value.
- Old backends can exit before attempting to resize. Synchronization
used between backends relies on ProcSignalBarrier and waits for all
participants received the message at the beginning to gather all
existing backends.
- Some backends might be blocked and not responsing either before or
after receiving the message. In the first case such backend still
have ProcSignalSlot and should be waited for, in the second case
shared barrier will make sure we still waiting for those backends. In
any case there is an unbounded wait.
- Backends might join barrier in disjoint groups with some time in
between. That means that relying only on the shared dynamic barrier is
not enough -- it will only synchronize resize procedure withing those
groups. That's why we wait first for all participants of ProcSignal
mechanism who received the message.
Here is how it looks like after raising shared_buffers from 128 MB to
512 MB and calling pg_reload_conf():
-- 128 MB
7f87909fc000-7f8798248000 rw-s /memfd:strategy (deleted)
7f8798248000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a4e84000 rw-s /memfd:checkpoint (deleted)
7f87a4e84000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1b42000 rw-s /memfd:iocv (deleted)
7f87b1b42000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cb59c000 rw-s /memfd:descriptors (deleted)
7f87cb59c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f87ece38000 rw-s /memfd:buffers (deleted)
^ buffers content, ~247 MB
7f87ece38000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~2210 MB
7f8877066000-7f887e7d0000 rw-s /memfd:main (deleted)
7f887e7d0000-7f8890a00000 ---s /memfd:main (deleted)
-- 512 MB
7f87909fc000-7f879866a000 rw-s /memfd:strategy (deleted)
7f879866a000-7f879d6ca000 ---s /memfd:strategy (deleted)
7f879d6ca000-7f87a50f4000 rw-s /memfd:checkpoint (deleted)
7f87a50f4000-7f87aa398000 ---s /memfd:checkpoint (deleted)
7f87aa398000-7f87b1d82000 rw-s /memfd:iocv (deleted)
7f87b1d82000-7f87c3d32000 ---s /memfd:iocv (deleted)
7f87c3d32000-7f87cba1c000 rw-s /memfd:descriptors (deleted)
7f87cba1c000-7f87dd6cc000 ---s /memfd:descriptors (deleted)
7f87dd6cc000-7f8804fb8000 rw-s /memfd:buffers (deleted)
^ buffers content, ~632 MB
7f8804fb8000-7f8877066000 ---s /memfd:buffers (deleted)
^ reserved space, ~1824 MB
7f8877066000-7f887e950000 rw-s /memfd:main (deleted)
7f887e950000-7f8890a00000 ---s /memfd:main (deleted)
The implementation supports only increasing of shared_buffers. For
decreasing the value a similar procedure is needed. But the buffer
blocks with data have to be drained first, so that the actual data set
fits into the new smaller space.
From experiment it turns out that shared mappings have to be extended
separately for each process that uses them. Another rough edge is that a
backend blocked on ReadCommand will not apply shared_buffers change
until it receives something.
Authors: Dmitrii Dolgov, Ashutosh Bapat
---
src/backend/port/sysv_shmem.c | 443 ++++++++++++++++++
src/backend/postmaster/checkpointer.c | 12 +-
src/backend/postmaster/postmaster.c | 18 +
src/backend/storage/buffer/buf_init.c | 60 ++-
src/backend/storage/ipc/ipci.c | 15 +-
src/backend/storage/ipc/procsignal.c | 46 ++
src/backend/storage/ipc/shmem.c | 23 +-
src/backend/tcop/postgres.c | 10 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/misc/guc_parameters.dat | 3 +-
src/include/storage/bufmgr.h | 2 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 26 +
src/include/storage/pmsignal.h | 3 +-
src/include/storage/procsignal.h | 1 +
src/tools/pgindent/typedefs.list | 1 +
17 files changed, 632 insertions(+), 38 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index b85911bdfc4..dc4eeeee56a 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -96,6 +102,13 @@ void *UsedShmemSegAddr = NULL;
AnonymousMapping Mappings[ANON_MAPPINGS];
+/* Flag telling postmaster that resize is needed */
+volatile bool pending_pm_shmem_resize = false;
+
+/* Keeps track of the previous NBuffers value */
+static int NBuffersOld = -1;
+static int NBuffersPending = -1;
+
/*
* Anonymous mapping layout we use looks like this:
*
@@ -147,6 +160,49 @@ static double SHMEM_RESIZE_RATIO[6] = {
*/
static bool huge_pages_on = false;
+/*
+ * Flag telling that we have prepared the memory layout to be resizable. If
+ * false after all shared memory segments creation, it means we failed to setup
+ * needed layout and falled back to the regular non-resizable approach.
+ */
+static bool shmem_resizable = false;
+
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -906,6 +962,346 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the current NBuffers value, which
+ * is is applied from NBuffersPending. The actual segment resizing is done via
+ * ftruncate, which will fail if is not sufficient space to expand the anon
+ * file. When finished, based on the new and old values initialize new buffer
+ * blocks if any.
+ *
+ * If reinitializing took place, as the last step this function does buffers
+ * reinitialization as well and broadcasts the new value of NSharedBuffers. All
+ * of that needs to be done only by one backend, the first one that managed to
+ * grab the ShmemResizeLock.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int numSemas;
+ bool reinit = false;
+ int mmap_flags = PG_MMAP_FLAGS;
+ Size hugepagesize;
+
+ NBuffers = NBuffersPending;
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
+
+ /*
+ * XXX: Where to reset the flag is still an open question. E.g. do we
+ * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
+ * and reset the flag?
+ */
+ pending_pm_shmem_resize = false;
+
+ /*
+ * XXX: Currently only increasing of shared_buffers is supported. For
+ * decreasing something similar has to be done, but buffer blocks with
+ * data have to be drained first.
+ */
+ if(NBuffersOld > NBuffers)
+ return false;
+
+#ifndef MAP_HUGETLB
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+
+ /* Round up the new size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+ }
+#endif
+
+ /* Note that CalculateShmemSize indirectly depends on NBuffers */
+ CalculateShmemSize(&numSemas);
+
+ for(int i = 0; i < ANON_MAPPINGS; i++)
+ {
+ AnonymousMapping *m = &Mappings[i];
+
+#ifdef MAP_HUGETLB
+ if (huge_pages_on && (m->shmem_req_size % hugepagesize != 0))
+ m->shmem_req_size += hugepagesize - (m->shmem_req_size % hugepagesize);
+#endif
+
+ if (m->shmem == NULL)
+ continue;
+
+ if (m->shmem_size == m->shmem_req_size)
+ continue;
+
+ if (m->shmem_reserved < m->shmem_req_size)
+ ereport(ERROR,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved"),
+ errhint("You may need to increase \"max_available_memory\".")));
+
+ elog(DEBUG1, "segment[%s]: resize from %zu to %zu at address %p",
+ MappingName(m->shmem_segment), m->shmem_size,
+ m->shmem_req_size, m->shmem);
+
+ /* Resize the backing anon file. */
+ if(ftruncate(m->segment_fd, m->shmem_req_size) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncase anonymous file for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ /* Adjust memory accessibility */
+ if(mprotect(m->shmem, m->shmem_req_size, PROT_READ | PROT_WRITE) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not mprotect anonymous shared memory for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ /* If shrinking, make reserved space unavailable again */
+ if(m->shmem_req_size < m->shmem_size &&
+ mprotect(m->shmem + m->shmem_req_size, m->shmem_size - m->shmem_req_size, PROT_NONE) == -1)
+ ereport(FATAL,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not mprotect reserved shared memory for \"%s\": %m",
+ MappingName(m->shmem_segment))));
+
+ reinit = true;
+ m->shmem_size = m->shmem_req_size;
+ }
+
+ if (reinit)
+ {
+ if(IsUnderPostmaster &&
+ LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ /*
+ * If the new NBuffers was already broadcasted, the buffer pool was
+ * already initialized before.
+ *
+ * Since we're not on a hot path, we use lwlocks and do not need to
+ * involve memory barrier.
+ */
+ if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ {
+ /*
+ * Allow the first backend that managed to get the lock to
+ * reinitialize the new portion of buffer pool. Every other
+ * process will wait on the shared barrier for that to finish,
+ * since it's a part of the SHMEM_RESIZE_DONE phase.
+ *
+ * Note that it's enough when only one backend will do that,
+ * even the ShmemInitStruct part. The reason is that resized
+ * shared memory will maintain the same addresses, meaning that
+ * all the pointers are still valid, and we only need to update
+ * structures size in the ShmemIndex once -- any other backend
+ * will pick up this shared structure from the index.
+ *
+ * XXX: This is the right place for buffer eviction as well.
+ */
+ BufferManagerShmemInit(NBuffersOld);
+
+ /* If all fine, broadcast the new value */
+ pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+ }
+
+ LWLockRelease(ShmemResizeLock);
+ }
+ }
+
+ return true;
+}
+
+/*
+ * We are asked to resize shared memory. Wait for all ProcSignal participants
+ * to join the barrier, then do the resize and wait on the barrier until all
+ * participating finish resizing as well -- otherwise we face danger of
+ * inconsistency between backends.
+ *
+ * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
+ * proceed with AnonymousShmemResize after receiving SIGHUP, until something
+ * will be sent.
+ */
+bool
+ProcessBarrierShmemResize(Barrier *barrier)
+{
+ Assert(IsUnderPostmaster);
+
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+
+ /* Wait until we have seen the new NBuffers value */
+ if (!pending_pm_shmem_resize)
+ return false;
+
+ /*
+ * First thing to do after attaching to the barrier is to wait for others.
+ * We can't simply use BarrierArriveAndWait, because backends might arrive
+ * here in disjoint groups, e.g. first two backends, pause, then second two
+ * backends. If the resize is quick enough that can lead to a situation
+ * when the first group is already finished before the second has appeared,
+ * and the barrier will only synchonize withing those groups.
+ */
+ if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
+ WaitForProcSignalBarrierReceived(
+ pg_atomic_read_u64(&ShmemCtrl->Generation));
+
+ /*
+ * Now start the procedure, and elect one backend to ping postmaster to do
+ * the same.
+ *
+ * XXX: If we need to be able to abort resizing, this has to be done later,
+ * after the SHMEM_RESIZE_DONE.
+ */
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ {
+ Assert(IsUnderPostmaster);
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ }
+
+ AnonymousShmemResize();
+
+ /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+
+ BarrierDetach(barrier);
+ return true;
+}
+
+/*
+ * GUC assign hook for shared_buffers. It's recommended for an assign hook to
+ * be as minimal as possible, thus we just request shared memory resize and
+ * remember the previous value.
+ */
+void
+assign_shared_buffers(int newval, void *extra, bool *pending)
+{
+ elog(DEBUG1, "Received SIGHUP for shmem resizing");
+
+ pending_pm_shmem_resize = true;
+ *pending = true;
+ NBuffersPending = newval;
+
+ NBuffersOld = NBuffers;
+}
+
+/*
+ * Test if we have somehow missed a shmem resize signal and NBuffers value
+ * differs from NSharedBuffers. If yes, catchup and do resize.
+ */
+void
+AdjustShmemSize(void)
+{
+ uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
+
+ if (NSharedBuffers != NBuffers)
+ {
+ /*
+ * If the broadcasted shared_buffers is different from the one we see,
+ * it could be that the backend has missed a resize signal. To avoid
+ * any inconsistency, adjust the shared mappings, before having a
+ * chance to access the buffer pool.
+ */
+ ereport(LOG,
+ (errmsg("shared_buffers has been changed from %d to %d, "
+ "resize shared memory",
+ NBuffers, NSharedBuffers)));
+ NBuffers = NSharedBuffers;
+ AnonymousShmemResize();
+ }
+}
+
+/*
+ * Start resizing procedure, making sure all existing processes will have
+ * consistent view of shared memory size. Must be called only in postmaster.
+ */
+void
+CoordinateShmemResize(void)
+{
+ elog(DEBUG1, "Coordinating shmem resize from %d to %d",
+ NBuffersOld, NBuffers);
+ Assert(!IsUnderPostmaster);
+
+ /*
+ * We use dynamic barrier to help dealing with backends that were spawned
+ * during the resize.
+ */
+ BarrierInit(&ShmemCtrl->Barrier, 0);
+
+ /*
+ * If the value did not change, or shared memory segments are not
+ * initialized yet, skip the resize.
+ */
+ if (NBuffersPending == NBuffersOld)
+ {
+ elog(DEBUG1, "Skip resizing, new %d, old %d",
+ NBuffers, NBuffersOld);
+ return;
+ }
+
+ /*
+ * Shared memory resize requires some coordination done by postmaster,
+ * and consists of three phases:
+ *
+ * - Before the resize all existing backends have the same old NBuffers.
+ * - When resize is in progress, backends are expected to have a
+ * mixture of old a new values. They're not allowed to touch buffer
+ * pool during this time frame.
+ * - After resize has been finished, all existing backends, that can access
+ * the buffer pool, are expected to have the same new value of NBuffers.
+ *
+ * Those phases are ensured by joining the shared barrier associated with
+ * the procedure. Since resizing takes time, we need to take into account
+ * that during that time:
+ *
+ * - New backends can be spawned. They will check status of the barrier
+ * early during the bootstrap, and wait until everything is over to work
+ * with the new NBuffers value.
+ *
+ * - Old backends can exit before attempting to resize. Synchronization
+ * used between backends relies on ProcSignalBarrier and waits for all
+ * participants received the message at the beginning to gather all
+ * existing backends.
+ *
+ * - Some backends might be blocked and not responsing either before or
+ * after receiving the message. In the first case such backend still
+ * have ProcSignalSlot and should be waited for, in the second case
+ * shared barrier will make sure we still waiting for those backends. In
+ * any case there is an unbounded wait.
+ *
+ * - Backends might join barrier in disjoint groups with some time in
+ * between. That means that relying only on the shared dynamic barrier is
+ * not enough -- it will only synchronize resize procedure withing those
+ * groups. That's why we wait first for all participants of ProcSignal
+ * mechanism who received the message.
+ */
+ elog(DEBUG1, "Emit a barrier for shmem resizing");
+ pg_atomic_init_u64(&ShmemCtrl->Generation,
+ EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
+
+ /* To order everything after setting Generation value */
+ pg_memory_barrier();
+
+ /*
+ * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
+ * the rest of the pack has started the procedure and it can resize shared
+ * memory as well.
+ *
+ * Normally we would call WaitForProcSignalBarrier here to wait until every
+ * backend has reported on the ProcSignalBarrier. But for shared memory
+ * resize we don't need this, as every participating backend will
+ * synchronize on the ProcSignal barrier. In fact even if we would like to
+ * wait here, it wouldn't be possible -- we're in the postmaster, without
+ * any waiting infrastructure available.
+ *
+ * If at some point it will turn out that waiting is essential, we would
+ * need to consider some alternatives. E.g. it could be a designated
+ * coordination process, which is not a postmaster. Another option would be
+ * to introduce a CoordinateShmemResize lock and allow only one process to
+ * take it (this probably would have to be something different than
+ * LWLocks, since they block interrupts, and coordination relies on them).
+ */
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1217,3 +1613,50 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+WaitOnShmemBarrier()
+{
+ Barrier *barrier = &ShmemCtrl->Barrier;
+
+ /* Nothing to do if resizing is not started */
+ if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
+ return;
+
+ BarrierAttach(barrier);
+
+ /* Otherwise wait through all available phases */
+ while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
+ {
+ ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
+ BarrierPhase(barrier))));
+
+ BarrierArriveAndWait(barrier, 0);
+ }
+
+ BarrierDetach(barrier);
+}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ /*
+ * The barrier is missing here, it will be initialized right before
+ * starting the resizing process as a convenient way to reset it.
+ */
+
+ /* Initialize with the currently known value */
+ pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
+
+ /* shmem_resizable should be initialized by now */
+ ShmemCtrl->Resizable = shmem_resizable;
+ }
+}
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index e84e8663e96..ef3f84a55f5 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -654,9 +654,12 @@ CheckpointerMain(const void *startup_data, size_t startup_data_len)
static void
ProcessCheckpointerInterrupts(void)
{
- if (ProcSignalBarrierPending)
- ProcessProcSignalBarrier();
-
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
if (ConfigReloadPending)
{
ConfigReloadPending = false;
@@ -676,6 +679,9 @@ ProcessCheckpointerInterrupts(void)
UpdateSharedMemoryConfig();
}
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+
/* Perform logging of memory contexts of this process */
if (LogMemoryContextPending)
ProcessLogMemoryContextInterrupt();
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index b59d20b4ac2..ba9528d5dfa 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -426,6 +426,7 @@ static void process_pm_pmsignal(void);
static void process_pm_child_exit(void);
static void process_pm_reload_request(void);
static void process_pm_shutdown_request(void);
+static void process_pm_shmem_resize(void);
static void dummy_handler(SIGNAL_ARGS);
static void CleanupBackend(PMChild *bp, int exitstatus);
static void HandleChildCrash(int pid, int exitstatus, const char *procname);
@@ -1697,6 +1698,9 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
+ if (pending_pm_shmem_resize)
+ process_pm_shmem_resize();
+
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2042,6 +2046,17 @@ process_pm_reload_request(void)
}
}
+static void
+process_pm_shmem_resize(void)
+{
+ /*
+ * Failure to resize is considered to be fatal and will not be
+ * retried, which means we can disable pending flag right here.
+ */
+ pending_pm_shmem_resize = false;
+ CoordinateShmemResize();
+}
+
/*
* pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of
* shutdown.
@@ -3862,6 +3877,9 @@ process_pm_pmsignal(void)
request_state_update = true;
}
+ if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
+ AnonymousShmemResize();
+
/*
* Try to advance postmaster's state machine, if a child requests it.
*/
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6f148d1d80b..0e72e373193 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -18,6 +18,7 @@
#include "storage/buf_internals.h"
#include "storage/pg_shmem.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -63,18 +64,28 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * postmaster, or in a standalone backend) or during shared-memory resize. Size
+ * of data structures initialized here depends on NBuffers, and to be able to
+ * change NBuffers without a restart we store each structure into a separate
+ * shared memory segment, which could be resized on demand.
+ *
+ * FirstBufferToInit tells where to start initializing buffers. For
+ * initialization it always will be zero, but when resizing shared-memory it
+ * indicates the number of already initialized buffers.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(void)
+BufferManagerShmemInit(int FirstBufferToInit)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
+ elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
+ FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
@@ -111,34 +122,35 @@ BufferManagerShmemInit(void)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory, or in all cases when resizing shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
+
+#ifndef EXEC_BACKEND
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = FirstBufferToInit; i < NBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
- ClearBufferTag(&buf->tag);
+ ClearBufferTag(&buf->tag);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- buf->buf_id = i;
+ buf->buf_id = i;
- pgaio_wref_clear(&buf->io_wref);
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
+#endif
/* Init other shared buffer-management stuff */
StrategyInitialize(!foundDescs);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2cd278449f0..bd75f06047e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -171,6 +171,14 @@ CalculateShmemSize(int *num_semaphores)
size = add_size(size, SlotSyncShmemSize());
size = add_size(size, AioShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -333,7 +341,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit();
+ BufferManagerShmemInit(0);
/*
* Set up lock manager
@@ -345,6 +353,11 @@ CreateOrAttachShmemStructs(void)
*/
PredicateLockShmemInit();
+ /*
+ * Set up shared memory resize manager
+ */
+ ShmemControlInit();
+
/*
* Set up process table
*/
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index eb3ceaae809..2160d258fa7 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -27,6 +27,7 @@
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -113,6 +114,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -176,6 +181,43 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -615,6 +657,10 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
+ processed = ProcessBarrierShmemResize(
+ &ShmemCtrl->Barrier);
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9499f332e77..2a197540300 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -498,17 +498,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. Verify the structure's size:
+ * - If it's the same, we've found the expected structure.
+ * - If it's different, we're resizing the expected structure.
+ *
+ * XXX: There is an implicit assumption this can only happen in
+ * "resizable" segments, where only one shared structure is allowed.
+ * This has to be implemented more cleanly.
*/
if (result->size != size)
{
- LWLockRelease(ShmemIndexLock);
- ereport(ERROR,
- (errmsg("ShmemIndex entry size is wrong for data structure"
- " \"%s\": expected %zu, actual %zu",
- name, size, result->size)));
+ Size delta = size - result->size;
+
+ result->size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
}
+
structPtr = result->location;
}
else
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 86ffe020c01..81881ef56c1 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -63,6 +63,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4318,6 +4319,15 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /* Verify the shared barrier, if it's still active: join and wait. */
+ WaitOnShmemBarrier();
+
+ /*
+ * After waiting on the barrier above we guaranteed to have NSharedBuffers
+ * broadcasted, so we can use it in the function below.
+ */
+ AdjustShmemSize();
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 7553f6eacef..82cee6b8877 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
+SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
WAL_RECEIVER_WAIT_START "Waiting for startup process to send initial data for streaming replication."
@@ -355,6 +357,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index cff8bb815f9..7b3ac5f3716 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -1105,13 +1105,14 @@
# We sometimes multiply the number of shared buffers by two without
# checking for overflow, so we mustn't allow more than INT_MAX / 2.
-{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
variable => 'NBuffers',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
+ assign_hook => 'assign_shared_buffers'
},
# TODO: should this be PGC_POSTMASTER?
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 3f37b294af6..e2e97866b40 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -318,7 +318,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_skipped);
/* in buf_init.c */
-extern void BufferManagerShmemInit(void);
+extern void BufferManagerShmemInit(int);
extern Size BufferManagerShmemSize(void);
/* in localbuf.c */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 3baf418b3d1..847f56a36dc 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index 06a1ffd4b08..cba586027a7 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -85,6 +85,7 @@ PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
+PG_LWLOCK(54, ShmemResize)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 715f6acb5dd..eba28ce8a5c 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,6 +24,7 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
@@ -69,6 +70,25 @@ typedef struct ShmemSegment
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ */
+typedef struct
+{
+ pg_atomic_uint32 NSharedBuffers;
+ Barrier Barrier;
+ pg_atomic_uint64 Generation;
+ bool Resizable;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -123,6 +143,12 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+void assign_shared_buffers(int newval, void *extra, bool *pending);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 428aa3fd68a..5ced2a83537 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,9 +42,10 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
-#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
+#define NUM_PMSIGNALS (PMSIGNAL_SHMEM_RESIZE+1)
/*
* Reasons why the postmaster would send SIGQUIT to its children.
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 2733bbb8c5b..97033f84dce 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,7 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
} ProcSignalBarrierType;
/*
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 5290b91e83e..691c39c8ad2 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2770,6 +2770,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.34.1
[application/x-patch] 0014-Support-shrinking-shared-buffers-20251013.patch (12.5K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/15-0014-Support-shrinking-shared-buffers-20251013.patch)
download | inline diff:
From 63dd4ccbb0ae0de2eefb72ae5a4a7bf7b5a6455b Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 19 Jun 2025 17:38:29 +0200
Subject: [PATCH 14/19] Support shrinking shared buffers
Buffer eviction
===============
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
Buffer eviction requires a separate barrier phase for two reasons:
1. No other backend should map a new page to any of buffers being
evicted when eviction is in progress. So they wait while eviction is
in progress.
2. Since a pinned buffer has the pin recorded in the backend local
memory as well as the buffer descriptor (which is in shared memory),
eviction should not coincide with remapping the shared memory of a
backend. Otherwise we might loose consistency of local and shared
pinning records. Hence it needs to be carried out in
ProcessBarrierShmemResize() and not in AnonymousShmemResize() as
indicated by now removed comment.
If a buffer being evicted is pinned, we raise a FATAL error but this should
improve. There are multiple options 1. to wait for the pinned buffer to get
unpinned, 2. the backend is killed or it itself cancels the query or 3.
rollback the operation. Note that option 1 and 2 would require the pinning
related local and shared records to be accessed. But we need infrastructure to
do either of this right now.
Removing the evicted buffers from buffer ring
=============================================
If the buffer pool has been shrunk, the buffers in the buffer ring may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition and masks any bug that
introduces an invalid buffer into the ring. The alternative fix is more
complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure thus accessible when processing buffer
resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work.
Author: Ashutosh Bapat
Reviewed-by: Tomas Vondra
---
src/backend/port/sysv_shmem.c | 42 ++++++---
src/backend/storage/buffer/bufmgr.c | 93 +++++++++++++++++++
src/backend/storage/buffer/freelist.c | 18 +++-
.../utils/activity/wait_event_names.txt | 1 +
src/include/storage/bufmgr.h | 1 +
src/include/storage/pg_shmem.h | 1 +
6 files changed, 139 insertions(+), 17 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 54d335b2e5d..9e1b2c3201f 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -993,14 +993,6 @@ AnonymousShmemResize(void)
*/
pending_pm_shmem_resize = false;
- /*
- * XXX: Currently only increasing of shared_buffers is supported. For
- * decreasing something similar has to be done, but buffer blocks with
- * data have to be drained first.
- */
- if(NBuffersOld > NBuffers)
- return false;
-
#ifndef MAP_HUGETLB
/* PrepareHugePages should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
@@ -1099,11 +1091,14 @@ AnonymousShmemResize(void)
* all the pointers are still valid, and we only need to update
* structures size in the ShmemIndex once -- any other backend
* will pick up this shared structure from the index.
- *
- * XXX: This is the right place for buffer eviction as well.
*/
BufferManagerShmemInit(NBuffersOld);
+ /*
+ * Wipe out the evictor PID so that it can be used for the next
+ * buffer resizing operation.
+ */
+ ShmemCtrl->evictor_pid = 0;
/* If all fine, broadcast the new value */
pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
}
@@ -1156,11 +1151,31 @@ ProcessBarrierShmemResize(Barrier *barrier)
* XXX: If we need to be able to abort resizing, this has to be done later,
* after the SHMEM_RESIZE_DONE.
*/
- if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+
+ /*
+ * Evict extra buffers when shrinking shared buffers. We need to do this
+ * while the memory for extra buffers is still mapped i.e. before remapping
+ * the shared memory segments to a smaller memory area.
+ */
+ if (NBuffersOld > NBuffersPending)
{
- Assert(IsUnderPostmaster);
- SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
+
+ /*
+ * TODO: If the buffer eviction fails for any reason, we should
+ * gracefully rollback the shared buffer resizing and try again. But the
+ * infrastructure to do so is not available right now. Hence just raise
+ * a FATAL so that the system restarts.
+ */
+ if (!EvictExtraBuffers(NBuffersPending, NBuffersOld))
+ elog(FATAL, "buffer eviction failed");
+
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_EVICT))
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
}
+ else
+ if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
AnonymousShmemResize();
@@ -1684,5 +1699,6 @@ ShmemControlInit(void)
/* shmem_resizable should be initialized by now */
ShmemCtrl->Resizable = shmem_resizable;
+ ShmemCtrl->evictor_pid = 0;
}
}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index edf17ce3ea1..467d9880f7b 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/read_stream.h"
#include "storage/smgr.h"
@@ -7457,3 +7458,95 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int newBufSize, int oldBufSize)
+{
+ bool result = true;
+
+ /*
+ * If the buffer being evicated is locked, this function will need to wait.
+ * This function should not be called from a Postmaster since it can not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * Let only one backend perform eviction. We could split the work across all
+ * the backends but that doesn't seem necessary.
+ *
+ * The first backend to acquire ShmemResizeLock, sets its own PID as the
+ * evictor PID for other backends to know that the eviction is in progress or
+ * has already been performed. The evictor backend releases the lock when it
+ * finishes eviction. While the eviction is in progress, backends other than
+ * evictor backend won't be able to take the lock. They won't perform
+ * eviction. A backend may acquire the lock after eviction has completed, but
+ * it will not perform eviction since the evictor PID is already set. Evictor
+ * PID is reset only when the buffer resizing finishes. Thus only one backend
+ * will perform eviction in a given instance of shared buffers resizing.
+ *
+ * Any backend which acquires this lock will release it before the eviction
+ * phase finishes, hence the same lock can be reused for the next phase of
+ * resizing buffers.
+ */
+ if (LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ {
+ if (ShmemCtrl->evictor_pid == 0)
+ {
+ ShmemCtrl->evictor_pid = MyProcPid;
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = newBufSize + 1; buf <= oldBufSize; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u32(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is
+ * going one hence unlocked precheck should be safe and saves
+ * some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ /*
+ * XXX: Looks like CurrentResourceOwner can be NULL here, find
+ * another one in that case?
+ * */
+ if (CurrentResourceOwner)
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+ }
+ LWLockRelease(ShmemResizeLock);
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 299f6aa8e7e..0da8fbb580e 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -669,12 +669,22 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot has been
+ * invalidated when buffer pool was shrunk, tell the caller to allocate a new
+ * buffer with the normal allocation strategy. He will then fill this slot
+ * by calling AddBufferToRing with the new buffer.
+ *
+ * TODO: Ideally we would want to check for bufnum > NBuffers only once
+ * after every time the buffer pool is shrunk so as to catch any runtime
+ * bugs that introduce invalid buffers in the ring. But that is complicated.
+ * The BufferAccessStrategy objects are not accessible outside the
+ * ScanState. Hence we can not purge the buffers while evicting the buffers.
+ * After the resizing is finished, it's not possible to notice when we touch
+ * the first of those objects and the last of objects. See if this can
+ * fixed.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer || bufnum > NBuffers)
return NULL;
buf = GetBufferDescriptor(bufnum - 1);
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 82cee6b8877..9a6a6275305 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -156,6 +156,7 @@ REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it c
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
+SHMEM_RESIZE_EVICT "Waiting for other backends to finish buffer evication phase."
SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index e2e97866b40..e7c973adca8 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -316,6 +316,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_evicted,
int32 *buffers_flushed,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int fromBuf, int toBuf);
/* in buf_init.c */
extern void BufferManagerShmemInit(int);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index eba28ce8a5c..0a59746b472 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -77,6 +77,7 @@ extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
typedef struct
{
pg_atomic_uint32 NSharedBuffers;
+ pid_t evictor_pid;
Barrier Barrier;
pg_atomic_uint64 Generation;
bool Resizable;
--
2.34.1
[application/x-patch] 0015-Reinitialize-StrategyControl-after-resizing-20251013.patch (19.1K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/16-0015-Reinitialize-StrategyControl-after-resizing-20251013.patch)
download | inline diff:
From 34c5477568b4a31863aeba57c4c0cb700e274f4e Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 19 Jun 2025 17:38:51 +0200
Subject: [PATCH 15/19] Reinitialize StrategyControl after resizing buffers
... and BgBufferSync and ClockSweepTick adjustments
Reinitializing strategry control area
=====================================
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
2. &StrategyControl->buffer_strategy_lock need not be initialized again.
3. nextVictimBuffer, completePasses and numBufferAllocs are viewed in
the context of NBuffers. Now that NBuffers itself has changed, those
three do not make sense. Reset them as if the server has restarted
again.
Ability to delay resizing operation
===================================
This commit introduces a flag delay_shmem_resize, which postgresql
backends and workers can use to signal the coordinator to delay resizing
operation. Background writer sets this flag when its scanning buffers.
Background writer operation
===========================
Background writer is blocked when the actual resizing is in progress. It
stops a scan in progress when it sees that the resizing has begun or is
about to begin. Once the buffer resizing is finished, before resuming
the regular operation, bgwriter resets the information saved so far.
This information is viewed in the context of NBuffers and hence does not
make sense after resizing which chanegs NBuffers.
Buffer lookup table
===================
Right now there is no way to free shared memory. Even if we shrink the
buffer lookup table when shrinking the buffer pool the unused hash table
entries can not be freed. When we expand the buffer pool, more entries
can be allocated but we can not resize the hash table directory without
rehashing all the entries. Just allocating more entries will lead to
more contention. Hence we setup the buffer lookup table considering the
maximum possible size of the buffer pool which is MaxAvailableMemory
only once at the beginning. Shared buffer lookup table and
StrategyControl are not resized even if the buffer pool is resized hence
they are allocated in the main shared memory segment
TODO:
====
1. The way BgBufferSync is written today, it packs four functionalities:
setting up the buffer sync state, performing the buffer sync,
resetting the buffer sync state when bgwriter_lru_maxpages <= 0 and
setting it up again after bgwriter_lru_maxpages > 0. That makes the
code hard to read. It will be good to divide this function into 3/4
different functions each performing one functionality. Then pack all
the state (the local variables from that function converted to static
global) into a structure, which is passed to these functions. Once
that happens BgBufferSyncReset() will call one of the functions to
reset the state when buffer pool is resized.
2. The condition (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) ==
NBuffers) checked in BgBufferSync() to check whether buffer resizing
is "about to begin" is wrong. NBuffers it not changed, until
AnonymousShmemResize() is called and it wont' be called unless
BgBufferSync() finishes if it has already begun. Need a better
condition to check whether buffer resizing is about to begin.
Author: Ashutosh Bapat
Reviewed-by: Tomas Vondra
---
src/backend/port/sysv_shmem.c | 23 ++++++--
src/backend/storage/buffer/buf_init.c | 19 +++++--
src/backend/storage/buffer/buf_table.c | 9 ++-
src/backend/storage/buffer/bufmgr.c | 72 ++++++++++++++++++------
src/backend/storage/buffer/freelist.c | 77 ++++++++++++++++++++++++--
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 1 +
src/include/storage/ipc.h | 1 +
src/include/storage/pg_shmem.h | 5 +-
9 files changed, 170 insertions(+), 38 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 9e1b2c3201f..3be28e228ae 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -104,6 +104,7 @@ AnonymousMapping Mappings[ANON_MAPPINGS];
/* Flag telling postmaster that resize is needed */
volatile bool pending_pm_shmem_resize = false;
+volatile bool delay_shmem_resize = false;
/* Keeps track of the previous NBuffers value */
static int NBuffersOld = -1;
@@ -144,12 +145,11 @@ static int NBuffersPending = -1;
* makes sense to evaluate them more precise.
*/
static double SHMEM_RESIZE_RATIO[6] = {
- 0.1, /* MAIN_SHMEM_SEGMENT */
+ 0.15, /* MAIN_SHMEM_SEGMENT */
0.6, /* BUFFERS_SHMEM_SEGMENT */
0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
- 0.05, /* STRATEGY_SHMEM_SEGMENT */
};
/*
@@ -225,8 +225,6 @@ MappingName(int shmem_segment)
return "iocv";
case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
return "checkpoint";
- case STRATEGY_SHMEM_SEGMENT:
- return "strategy";
default:
return "unknown";
}
@@ -1125,13 +1123,17 @@ ProcessBarrierShmemResize(Barrier *barrier)
{
Assert(IsUnderPostmaster);
- elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d",
- NBuffersOld, NBuffersPending, pending_pm_shmem_resize);
+ elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d, %d",
+ NBuffersOld, NBuffersPending, pending_pm_shmem_resize, delay_shmem_resize);
/* Wait until we have seen the new NBuffers value */
if (!pending_pm_shmem_resize)
return false;
+ /* Wait till this process becomes ready to resize buffers. */
+ if (delay_shmem_resize)
+ return false;
+
/*
* First thing to do after attaching to the barrier is to wait for others.
* We can't simply use BarrierArriveAndWait, because backends might arrive
@@ -1182,6 +1184,15 @@ ProcessBarrierShmemResize(Barrier *barrier)
/* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * Before resuming regular background writer activity, adjust the
+ * statistics collected so far.
+ */
+ BgBufferSyncReset(NBuffersOld, NBuffers);
+ }
+
BarrierDetach(barrier);
return true;
}
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 0e72e373193..be64fa5a136 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -152,8 +152,15 @@ BufferManagerShmemInit(int FirstBufferToInit)
}
#endif
- /* Init other shared buffer-management stuff */
- StrategyInitialize(!foundDescs);
+ /*
+ * Init other shared buffer-management stuff from scratch configuring buffer
+ * pool the first time. If we are just resizing buffer pool adjust only the
+ * required structures.
+ */
+ if (FirstBufferToInit == 0)
+ StrategyInitialize(!foundDescs);
+ else
+ StrategyReInitialize(FirstBufferToInit);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
@@ -184,9 +191,6 @@ BufferManagerShmemSize(void)
size = add_size(size, mul_size(NBuffers, BLCKSZ));
Mappings[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
- /* size of stuff controlled by freelist.c */
- Mappings[STRATEGY_SHMEM_SEGMENT].shmem_req_size = StrategyShmemSize();
-
/* size of I/O condition variables, plus alignment padding */
size = add_size(0, mul_size(NBuffers,
sizeof(ConditionVariableMinimallyPadded)));
@@ -196,5 +200,10 @@ BufferManagerShmemSize(void)
/* size of checkpoint sort array in bufmgr.c */
Mappings[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
+ /* Allocations in the main memory segment, at the end. */
+
+ /* size of stuff controlled by freelist.c */
+ size = add_size(0, StrategyShmemSize());
+
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 18a78967138..e5a97e557d9 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -65,11 +65,18 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
+ /*
+ * The shared buffer look up table is set up only once with maximum possible
+ * entries considering maximum size of the buffer pool. It is not resized
+ * after that even if the buffer pool is resized. Hence it is allocated in
+ * the main shared memory segment and not in a resizeable shared memory
+ * segment.
+ */
SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE,
- STRATEGY_SHMEM_SEGMENT);
+ MAIN_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 467d9880f7b..fdcb5556235 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3614,6 +3614,32 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_DONE(NBuffers, num_written, num_to_scan);
}
+/*
+ * Information saved between BgBufferSync() calls so we can determine the
+ * strategy point's advance rate and avoid scanning already-cleaned buffers. The
+ * variables are global instead of static local so that BgBufferSyncReset() can
+ * adjust it when resizing shared buffers.
+ */
+static bool saved_info_valid = false;
+static int prev_strategy_buf_id;
+static uint32 prev_strategy_passes;
+static int next_to_clean;
+static uint32 next_passes;
+
+/* Moving averages of allocation rate and clean-buffer density */
+static float smoothed_alloc = 0;
+static float smoothed_density = 10.0;
+
+void
+BgBufferSyncReset(int NBuffersOld, int NBuffersNew)
+{
+ saved_info_valid = false;
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
+ NBuffersOld, NBuffersNew);
+#endif
+}
+
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
@@ -3633,20 +3659,6 @@ BgBufferSync(WritebackContext *wb_context)
uint32 strategy_passes;
uint32 recent_alloc;
- /*
- * Information saved between calls so we can determine the strategy
- * point's advance rate and avoid scanning already-cleaned buffers.
- */
- static bool saved_info_valid = false;
- static int prev_strategy_buf_id;
- static uint32 prev_strategy_passes;
- static int next_to_clean;
- static uint32 next_passes;
-
- /* Moving averages of allocation rate and clean-buffer density */
- static float smoothed_alloc = 0;
- static float smoothed_density = 10.0;
-
/* Potentially these could be tunables, but for now, not */
float smoothing_samples = 16;
float scan_whole_pool_milliseconds = 120000.0;
@@ -3669,6 +3681,22 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ /*
+ * If buffer pool is being shrunk the buffer being written out may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing out any. Hence wait till
+ * buffer resizing finishes i.e. go into hibernation mode.
+ */
+ if (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ return true;
+
+ /*
+ * Resizing shared buffers while this function is performing an LRU scan on
+ * them may lead to wrong results. Indicate that the resizing should wait for
+ * the LRU scan to complete.
+ */
+ delay_shmem_resize = true;
+
/*
* Find out where the clock-sweep currently is, and how many buffer
* allocations have happened since our last call.
@@ -3845,8 +3873,17 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
- while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
+ /*
+ * Execute the LRU scan.
+ *
+ * If buffer pool is being shrunk, the buffer being written may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing any. Hence stop what we are doing. This
+ * also unblocks other processes that are waiting for buffer resizing to
+ * finish.
+ */
+ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
+ pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) == NBuffers)
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -3905,6 +3942,9 @@ BgBufferSync(WritebackContext *wb_context)
#endif
}
+ /* Let the resizing commence. */
+ delay_shmem_resize = false;
+
/* Return true if OK to hibernate */
return (bufs_to_lap == 0 && recent_alloc == 0);
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 0da8fbb580e..55be5eebe0a 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -408,12 +408,21 @@ StrategyInitialize(bool init)
*
* Since we can't tolerate running out of lookup table entries, we must be
* sure to specify an adequate table size here. The maximum steady-state
- * usage is of course NBuffers entries, but BufferAlloc() tries to insert
- * a new entry before deleting the old. In principle this could be
- * happening in each partition concurrently, so we could need as many as
- * NBuffers + NUM_BUFFER_PARTITIONS entries.
+ * usage is of course is as many number of entries as the number of buffers
+ * in the buffer pool. Right now there is no way to free shared memory. Even
+ * if we shrink the buffer lookup table when shrinking the buffer pool the
+ * unused hash table entries can not be freed. When we expand the buffer
+ * pool, more entries can be allocated but we can not resize the hash table
+ * directory without rehashing all the entries. Just allocating more entries
+ * will lead to more contention. Hence we setup the buffer lookup table
+ * considering the maximum possible size of the buffer pool which is
+ * MaxAvailableMemory.
+ *
+ * Additionally BufferAlloc() tries to insert a new entry before deleting the
+ * old. In principle this could be happening in each partition concurrently,
+ * so we need extra NUM_BUFFER_PARTITIONS entries.
*/
- InitBufTable(NBuffers + NUM_BUFFER_PARTITIONS);
+ InitBufTable(MaxAvailableMemory + NUM_BUFFER_PARTITIONS);
/*
* Get or create the shared strategy control block
@@ -421,7 +430,7 @@ StrategyInitialize(bool init)
StrategyControl = (BufferStrategyControl *)
ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found, STRATEGY_SHMEM_SEGMENT);
+ &found, MAIN_SHMEM_SEGMENT);
if (!found)
{
@@ -446,6 +455,62 @@ StrategyInitialize(bool init)
Assert(!init);
}
+/*
+ * StrategyReInitialize -- re-initialize the buffer cache replacement
+ * strategy.
+ *
+ * To be called when resizing buffer manager and only from the coordinator.
+ * TODO: Assess the differences between this function and StrategyInitialize().
+ */
+void
+StrategyReInitialize(int FirstBufferIdToInit)
+{
+ bool found;
+
+ /*
+ * Resizing memory for buffer pools should not affect the address of
+ * StrategyControl.
+ */
+ if (StrategyControl != (BufferStrategyControl *)
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, MAIN_SHMEM_SEGMENT))
+ elog(FATAL, "something went wrong while re-initializing the buffer strategy");
+
+ Assert(found);
+
+ /* TODO: Buffer lookup table adjustment: There are two options:
+ *
+ * 1. Resize the buffer lookup table to match the new number of buffers. But
+ * this requires rehashing all the entries in the buffer lookup table with
+ * the new table size.
+ *
+ * 2. Allocate maximum size of the buffer lookup table at the beginning and
+ * never resize it. This leaves sparse buffer lookup table which is
+ * inefficient from both memory and time perspective. According to David
+ * Rowley, the sparse entries in the buffer look up table cause frequent
+ * cacheline reload which affect performance. If the impact of that
+ * inefficiency in a benchmark is significant, we will need to consider first
+ * option.
+ */
+ /*
+ * The clock sweep tick pointer might have got invalidated. Reset it as if
+ * starting a fresh server.
+ */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The old statistics is viewed in the context of the number of shared
+ * buffers. It does not make sense now that the number of shared buffers
+ * itself has changed.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_init_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* No pending notification */
+ StrategyControl->bgwprocno = -1;
+}
+
/* ----------------------------------------------------------------
* Backend-private buffer ring management
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index c1206a46aba..20bea8132fd 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -447,6 +447,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyReInitialize(int FirstBufferToInit);
/* buf_table.c */
extern Size BufTableShmemSize(int size);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index e7c973adca8..74e226269af 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -300,6 +300,7 @@ extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
extern bool BgBufferSync(WritebackContext *wb_context);
+extern void BgBufferSyncReset(int NBuffersOld, int NBuffersNew);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 847f56a36dc..6e7b0abb625 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -65,6 +65,7 @@ typedef void (*shmem_startup_hook_type) (void);
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
+extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 0a59746b472..704b065f9e9 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -65,7 +65,7 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define ANON_MAPPINGS 6
+#define ANON_MAPPINGS 5
extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS];
extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
@@ -172,7 +172,4 @@ extern void ShmemControlInit(void);
/* Checkpoint BufferIds */
#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
-/* Buffer strategy status */
-#define STRATEGY_SHMEM_SEGMENT 5
-
#endif /* PG_SHMEM_H */
--
2.34.1
[application/x-patch] 0016-Tests-for-dynamic-shared_buffers-resizing-20251013.patch (19.6K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/17-0016-Tests-for-dynamic-shared_buffers-resizing-20251013.patch)
download | inline diff:
From 603178a5c369456014cbd0fb6c171074f66aa163 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 3 Sep 2025 10:59:20 +0530
Subject: [PATCH 16/19] Tests for dynamic shared_buffers resizing
The commit adds two tests:
1. TAP test to stress test buffer pool resizing under concurrent load.
2. SQL test to test sanity of shared memory allocations and mappings
after buffer pool resizing operation.
Author: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
src/test/Makefile | 2 +-
src/test/README | 3 +
src/test/buffermgr/Makefile | 27 ++
src/test/buffermgr/README | 26 ++
src/test/buffermgr/expected/buffer_resize.out | 237 ++++++++++++++++++
src/test/buffermgr/meson.build | 17 ++
src/test/buffermgr/sql/buffer_resize.sql | 73 ++++++
src/test/buffermgr/t/001_resize_buffer.pl | 126 ++++++++++
src/test/meson.build | 1 +
9 files changed, 511 insertions(+), 1 deletion(-)
create mode 100644 src/test/buffermgr/Makefile
create mode 100644 src/test/buffermgr/README
create mode 100644 src/test/buffermgr/expected/buffer_resize.out
create mode 100644 src/test/buffermgr/meson.build
create mode 100644 src/test/buffermgr/sql/buffer_resize.sql
create mode 100644 src/test/buffermgr/t/001_resize_buffer.pl
diff --git a/src/test/Makefile b/src/test/Makefile
index 511a72e6238..95f8858a818 100644
--- a/src/test/Makefile
+++ b/src/test/Makefile
@@ -12,7 +12,7 @@ subdir = src/test
top_builddir = ../..
include $(top_builddir)/src/Makefile.global
-SUBDIRS = perl postmaster regress isolation modules authentication recovery subscription
+SUBDIRS = perl postmaster regress isolation modules authentication recovery subscription buffermgr
ifeq ($(with_icu),yes)
SUBDIRS += icu
diff --git a/src/test/README b/src/test/README
index afdc7676519..77f11607ff7 100644
--- a/src/test/README
+++ b/src/test/README
@@ -15,6 +15,9 @@ examples/
Demonstration programs for libpq that double as regression tests via
"make check"
+buffermgr/
+ Tests for resizing buffer pool without restarting the server
+
isolation/
Tests for concurrent behavior at the SQL level
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
new file mode 100644
index 00000000000..97c3da9e20a
--- /dev/null
+++ b/src/test/buffermgr/Makefile
@@ -0,0 +1,27 @@
+#-------------------------------------------------------------------------
+#
+# Makefile for src/test/buffermgr
+#
+# Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+# Portions Copyright (c) 1994, Regents of the University of California
+#
+# src/test/buffermgr/Makefile
+#
+#-------------------------------------------------------------------------
+
+EXTRA_INSTALL = contrib/pg_buffercache
+
+REGRESS = buffer_resize
+
+subdir = src/test/buffermgr
+top_builddir = ../../..
+include $(top_builddir)/src/Makefile.global
+
+check:
+ $(prove_check)
+
+installcheck:
+ $(prove_installcheck)
+
+clean distclean:
+ rm -rf tmp_check
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
new file mode 100644
index 00000000000..c375ad80989
--- /dev/null
+++ b/src/test/buffermgr/README
@@ -0,0 +1,26 @@
+src/test/buffermgr/README
+
+Regression tests for buffer manager
+===================================
+
+This directory contains a test suite for resizing buffer manager without restarting the server.
+
+
+Running the tests
+=================
+
+NOTE: You must have given the --enable-tap-tests argument to configure.
+
+Run
+ make check
+or
+ make installcheck
+You can use "make installcheck" if you previously did "make install".
+In that case, the code in the installation tree is tested. With
+"make check", a temporary installation tree is built from the current
+sources and then tested.
+
+Either way, this test initializes, starts, and stops a test Postgres
+cluster.
+
+See src/test/perl/README for more info about running these tests.
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
new file mode 100644
index 00000000000..a986be9a5da
--- /dev/null
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -0,0 +1,237 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+-- Create a separate schema for this test
+CREATE SCHEMA buffer_resize_test;
+SET search_path TO buffer_resize_test, public;
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, mapping_size, mapping_reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | descriptors | 1048576 | 1048576
+ Buffer IO Condition Variables | iocv | 262144 | 262144
+ Checkpoint BufferIds | checkpoint | 327680 | 327680
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 134225920 | 134225920 | 2576982016
+ checkpoint | 335872 | 335872 | 214753280
+ descriptors | 1056768 | 1056768 | 429498368
+ iocv | 270336 | 270336 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+SELECT pg_sleep(1);
+ pg_sleep
+----------
+
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 64MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+----------+----------------
+ Buffer Blocks | buffers | 67112960 | 67112960
+ Buffer Descriptors | descriptors | 524288 | 524288
+ Buffer IO Condition Variables | iocv | 131072 | 131072
+ Checkpoint BufferIds | checkpoint | 163840 | 163840
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+----------+--------------+-----------------------
+ buffers | 67117056 | 67117056 | 2576982016
+ checkpoint | 172032 | 172032 | 214753280
+ descriptors | 532480 | 532480 | 429498368
+ iocv | 139264 | 139264 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 8192
+(1 row)
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+SELECT pg_sleep(1);
+ pg_sleep
+----------
+
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 268439552 | 268439552
+ Buffer Descriptors | descriptors | 2097152 | 2097152
+ Buffer IO Condition Variables | iocv | 524288 | 524288
+ Checkpoint BufferIds | checkpoint | 655360 | 655360
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 268443648 | 268443648 | 2576982016
+ checkpoint | 663552 | 663552 | 214753280
+ descriptors | 2105344 | 2105344 | 429498368
+ iocv | 532480 | 532480 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 32768
+(1 row)
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+SELECT pg_sleep(1);
+ pg_sleep
+----------
+
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 100MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 104861696 | 104861696
+ Buffer Descriptors | descriptors | 819200 | 819200
+ Buffer IO Condition Variables | iocv | 204800 | 204800
+ Checkpoint BufferIds | checkpoint | 256000 | 256000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 104865792 | 104865792 | 2576982016
+ checkpoint | 262144 | 262144 | 214753280
+ descriptors | 827392 | 827392 | 429498368
+ iocv | 212992 | 212992 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 12800
+(1 row)
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+SELECT pg_sleep(1);
+ pg_sleep
+----------
+
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+--------+----------------
+ Buffer Blocks | buffers | 135168 | 135168
+ Buffer Descriptors | descriptors | 1024 | 1024
+ Buffer IO Condition Variables | iocv | 256 | 256
+ Checkpoint BufferIds | checkpoint | 320 | 320
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+--------+--------------+-----------------------
+ buffers | 139264 | 139264 | 2576982016
+ checkpoint | 8192 | 8192 | 214753280
+ descriptors | 8192 | 8192 | 429498368
+ iocv | 8192 | 8192 | 429498368
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16
+(1 row)
+
+-- Clean up the schema and all its objects
+RESET search_path;
+DROP SCHEMA buffer_resize_test CASCADE;
+NOTICE: drop cascades to 3 other objects
+DETAIL: drop cascades to view buffer_resize_test.buffer_allocations
+drop cascades to view buffer_resize_test.buffer_segments
+drop cascades to extension pg_buffercache
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
new file mode 100644
index 00000000000..e71dcdea685
--- /dev/null
+++ b/src/test/buffermgr/meson.build
@@ -0,0 +1,17 @@
+# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+
+tests += {
+ 'name': 'buffermgr',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'buffer_resize',
+ ],
+ },
+ 'tap': {
+ 'tests': [
+ 't/001_resize_buffer.pl',
+ ],
+ },
+}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
new file mode 100644
index 00000000000..45f5bb6d78b
--- /dev/null
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -0,0 +1,73 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+
+-- Create a separate schema for this test
+CREATE SCHEMA buffer_resize_test;
+SET search_path TO buffer_resize_test, public;
+
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, mapping_size, mapping_reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+SELECT pg_sleep(1);
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+SELECT pg_sleep(1);
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+SELECT pg_sleep(1);
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+SELECT pg_sleep(1);
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Clean up the schema and all its objects
+RESET search_path;
+DROP SCHEMA buffer_resize_test CASCADE;
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
new file mode 100644
index 00000000000..8cf9e4539ab
--- /dev/null
+++ b/src/test/buffermgr/t/001_resize_buffer.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Minimal test testing shared_buffer resizing under load
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Function to resize buffer pool and verify the change.
+sub apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ # Use a single background_psql session for consistency
+ my $psql_session = $node->background_psql('postgres');
+ $psql_session->query_safe("ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $psql_session->query_safe("SELECT pg_reload_conf()");
+
+ # Wait till the resizing finishes using the same session
+ #
+ # TODO: Right now there is no way to know when the resize has finished and
+ # all the backends are using new value of shared_buffers. Hence we poll
+ # manually until we get the expected value in the same session.
+ my $current_size;
+ my $attempts = 0;
+ my $max_attempts = 60; # 60 seconds timeout
+ do {
+ $current_size = $psql_session->query_safe("SHOW shared_buffers");
+ $attempts++;
+
+ # Only sleep if we didn't get the expected result and haven't timed out yet
+ if ($current_size ne $new_size && $attempts < $max_attempts) {
+ sleep(1);
+ }
+ } while ($current_size ne $new_size && $attempts < $max_attempts);
+
+ $psql_session->quit;
+
+ # Check if we succeeded or timed out
+ if ($current_size ne $new_size) {
+ die "Timeout waiting for shared_buffers to change to $new_size (got $current_size after ${attempts}s)";
+ }
+}
+
+# Initialize a cluster and start pgbench in the background for concurrent load.
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->start;
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+my $pgb_scale = 10;
+my $pgb_duration = 120;
+my $pgb_num_clients = 10;
+$node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$pgb_scale --quiet",
+ 0,
+ [qr{^$}],
+ [ # stderr patterns to verify initialization stages
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$pgb_scale)"
+);
+my ($pgbench_stdin, $pgbench_stdout, $pgbench_stderr) = ('', '', '');
+my $pgbench_process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-T', $pgb_duration,
+ '-c', $pgb_num_clients,
+ 'postgres'
+ ],
+ '<' => \$pgbench_stdin,
+ '>' => \$pgbench_stdout,
+ '2>' => \$pgbench_stderr
+);
+
+ok($pgbench_process, "pgbench started successfully");
+
+# Allow pgbench to establish connections and start generating load.
+#
+# TODO: When creating new backends is known to work well with buffer pool
+# resizing, this wait should be removed.
+sleep(1);
+
+# Resize buffer pool to various sizes while pgbench is running in the
+# background.
+#
+# TODO: These are pseudo-randomly picked sizes, but we can do better.
+my $tests_completed = 0;
+my @buffer_sizes = ('900MB', '500MB', '250MB', '400MB', '120MB', '600MB');
+for my $target_size (@buffer_sizes)
+{
+ # Verify workload generator is still running
+ if (!$pgbench_process->pumpable) {
+ ok(0, "pgbench is still running");
+ last;
+ }
+
+ apply_and_verify_buffer_change($node, $target_size);
+ $tests_completed++;
+
+ # Wait for the resized buffer pool to stabilize. If the resized buffer pool
+ # is utilized fully, it might hit any wrongly initialized areas of shared
+ # memory.
+ sleep(2);
+}
+is($tests_completed, scalar(@buffer_sizes), "All buffer sizes were tested");
+
+# Make sure that pgbench can end normally.
+$pgbench_process->signal('TERM');
+IPC::Run::finish $pgbench_process;
+ok(grep { $pgbench_process->result == $_ } (0, 15), "pgbench exited gracefully");
+
+# Log any error output from pgbench for debugging
+diag("pgbench stderr:\n$pgbench_stderr");
+diag("pgbench stdout:\n$pgbench_stdout");
+
+# Ensure database is still functional after all the buffer changes
+$node->connect_ok("dbname=postgres",
+ "Database remains accessible after $tests_completed buffer resize operations");
+
+done_testing();
\ No newline at end of file
diff --git a/src/test/meson.build b/src/test/meson.build
index ccc31d6a86a..2a5ba1dec39 100644
--- a/src/test/meson.build
+++ b/src/test/meson.build
@@ -4,6 +4,7 @@ subdir('regress')
subdir('isolation')
subdir('authentication')
+subdir('buffermgr')
subdir('postmaster')
subdir('recovery')
subdir('subscription')
--
2.34.1
[application/x-patch] 0017-Revert-Introduce-pending-flag-for-GUC-assig-20251013.patch (12.1K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/18-0017-Revert-Introduce-pending-flag-for-GUC-assig-20251013.patch)
download | inline diff:
From aeadb3079216f970505605217a2bd9fabeff584a Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Fri, 26 Sep 2025 16:47:24 +0530
Subject: [PATCH 17/19] Revert "Introduce pending flag for GUC assign hooks"
This reverts commit 0a13e56dceea8cc7a2685df7ee8cea434588681b.
---
src/backend/access/transam/xlog.c | 2 +-
src/backend/commands/variable.c | 6 +--
src/backend/libpq/pqcomm.c | 8 ++--
src/backend/tcop/postgres.c | 2 +-
src/backend/utils/misc/guc.c | 59 +++++++++-------------------
src/backend/utils/misc/stack_depth.c | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 20 +++++-----
8 files changed, 40 insertions(+), 61 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index cc48b253bc8..eceab341255 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -2197,7 +2197,7 @@ CalculateCheckpointSegments(void)
}
void
-assign_max_wal_size(int newval, void *extra, bool *pending)
+assign_max_wal_size(int newval, void *extra)
{
max_wal_size_mb = newval;
CalculateCheckpointSegments();
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index e40dae2ddf2..608f10d9412 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source)
* GUC assign_hook for maintenance_io_concurrency
*/
void
-assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
+assign_maintenance_io_concurrency(int newval, void *extra)
{
/*
* Reconfigure recovery prefetching, because a setting it depends on
@@ -1161,12 +1161,12 @@ assign_maintenance_io_concurrency(int newval, void *extra, bool *pending)
* they may be assigned in either order.
*/
void
-assign_io_max_combine_limit(int newval, void *extra, bool *pending)
+assign_io_max_combine_limit(int newval, void *extra)
{
io_combine_limit = Min(newval, io_combine_limit_guc);
}
void
-assign_io_combine_limit(int newval, void *extra, bool *pending)
+assign_io_combine_limit(int newval, void *extra)
{
io_combine_limit = Min(io_max_combine_limit, newval);
}
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index 1726a7c0993..25f739a6a17 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1951,7 +1951,7 @@ pq_settcpusertimeout(int timeout, Port *port)
* GUC assign_hook for tcp_keepalives_idle
*/
void
-assign_tcp_keepalives_idle(int newval, void *extra, bool *pending)
+assign_tcp_keepalives_idle(int newval, void *extra)
{
/*
* The kernel API provides no way to test a value without setting it; and
@@ -1984,7 +1984,7 @@ show_tcp_keepalives_idle(void)
* GUC assign_hook for tcp_keepalives_interval
*/
void
-assign_tcp_keepalives_interval(int newval, void *extra, bool *pending)
+assign_tcp_keepalives_interval(int newval, void *extra)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivesinterval(newval, MyProcPort);
@@ -2007,7 +2007,7 @@ show_tcp_keepalives_interval(void)
* GUC assign_hook for tcp_keepalives_count
*/
void
-assign_tcp_keepalives_count(int newval, void *extra, bool *pending)
+assign_tcp_keepalives_count(int newval, void *extra)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_setkeepalivescount(newval, MyProcPort);
@@ -2030,7 +2030,7 @@ show_tcp_keepalives_count(void)
* GUC assign_hook for tcp_user_timeout
*/
void
-assign_tcp_user_timeout(int newval, void *extra, bool *pending)
+assign_tcp_user_timeout(int newval, void *extra)
{
/* See comments in assign_tcp_keepalives_idle */
(void) pq_settcpusertimeout(newval, MyProcPort);
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 81881ef56c1..ee9f308379c 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3598,7 +3598,7 @@ check_log_stats(bool *newval, void **extra, GucSource source)
/* GUC assign hook for transaction_timeout */
void
-assign_transaction_timeout(int newval, void *extra, bool *pending)
+assign_transaction_timeout(int newval, void *extra)
{
if (IsTransactionState())
{
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index c9361a0e423..8794e26ef1d 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -1681,7 +1681,6 @@ InitializeOneGUCOption(struct config_generic *gconf)
struct config_int *conf = (struct config_int *) gconf;
int newval = conf->boot_val;
void *extra = NULL;
- bool pending = false;
Assert(newval >= conf->min);
Assert(newval <= conf->max);
@@ -1690,13 +1689,9 @@ InitializeOneGUCOption(struct config_generic *gconf)
elog(FATAL, "failed to initialize %s to %d",
conf->gen.name, newval);
if (conf->assign_hook)
- conf->assign_hook(newval, extra, &pending);
-
- if (!pending)
- {
- *conf->variable = conf->reset_val = newval;
- conf->gen.extra = conf->reset_extra = extra;
- }
+ conf->assign_hook(newval, extra);
+ *conf->variable = conf->reset_val = newval;
+ conf->gen.extra = conf->reset_extra = extra;
break;
}
case PGC_REAL:
@@ -2052,18 +2047,13 @@ ResetAllOptions(void)
case PGC_INT:
{
struct config_int *conf = (struct config_int *) gconf;
- bool pending = false;
if (conf->assign_hook)
conf->assign_hook(conf->reset_val,
- conf->reset_extra,
- &pending);
- if (!pending)
- {
- *conf->variable = conf->reset_val;
- set_extra_field(&conf->gen, &conf->gen.extra,
- conf->reset_extra);
- }
+ conf->reset_extra);
+ *conf->variable = conf->reset_val;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ conf->reset_extra);
break;
}
case PGC_REAL:
@@ -2440,21 +2430,16 @@ AtEOXact_GUC(bool isCommit, int nestLevel)
struct config_int *conf = (struct config_int *) gconf;
int newval = newvalue.val.intval;
void *newextra = newvalue.extra;
- bool pending = false;
if (*conf->variable != newval ||
conf->gen.extra != newextra)
{
if (conf->assign_hook)
- conf->assign_hook(newval, newextra, &pending);
-
- if (!pending)
- {
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- changed = true;
- }
+ conf->assign_hook(newval, newextra);
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ changed = true;
}
break;
}
@@ -3871,24 +3856,18 @@ set_config_with_handle(const char *name, config_handle *handle,
if (changeVal)
{
- bool pending = false;
-
/* Save old value to support transaction abort */
if (!makeDefault)
push_old_value(&conf->gen, action);
if (conf->assign_hook)
- conf->assign_hook(newval, newextra, &pending);
-
- if (!pending)
- {
- *conf->variable = newval;
- set_extra_field(&conf->gen, &conf->gen.extra,
- newextra);
- set_guc_source(&conf->gen, source);
- conf->gen.scontext = context;
- conf->gen.srole = srole;
- }
+ conf->assign_hook(newval, newextra);
+ *conf->variable = newval;
+ set_extra_field(&conf->gen, &conf->gen.extra,
+ newextra);
+ set_guc_source(&conf->gen, source);
+ conf->gen.scontext = context;
+ conf->gen.srole = srole;
}
if (makeDefault)
{
diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c
index ef59ae62008..8f7cf531fbc 100644
--- a/src/backend/utils/misc/stack_depth.c
+++ b/src/backend/utils/misc/stack_depth.c
@@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source)
/* GUC assign hook for max_stack_depth */
void
-assign_max_stack_depth(int newval, void *extra, bool *pending)
+assign_max_stack_depth(int newval, void *extra)
{
ssize_t newval_bytes = newval * (ssize_t) 1024;
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index c3056cd2da8..f21ec37da89 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc
typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source);
typedef void (*GucBoolAssignHook) (bool newval, void *extra);
-typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending);
+typedef void (*GucIntAssignHook) (int newval, void *extra);
typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 658c799419e..82ac8646a8d 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -81,12 +81,12 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
extern const char *show_log_timezone(void);
-extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending);
-extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending);
-extern void assign_io_combine_limit(int newval, void *extra, bool *pending);
-extern void assign_max_wal_size(int newval, void *extra, bool *pending);
+extern void assign_maintenance_io_concurrency(int newval, void *extra);
+extern void assign_io_max_combine_limit(int newval, void *extra);
+extern void assign_io_combine_limit(int newval, void *extra);
+extern void assign_max_wal_size(int newval, void *extra);
extern bool check_max_stack_depth(int *newval, void **extra, GucSource source);
-extern void assign_max_stack_depth(int newval, void *extra, bool *pending);
+extern void assign_max_stack_depth(int newval, void *extra);
extern bool check_multixact_member_buffers(int *newval, void **extra,
GucSource source);
extern bool check_multixact_offset_buffers(int *newval, void **extra,
@@ -141,13 +141,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra);
extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
-extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending);
+extern void assign_tcp_keepalives_count(int newval, void *extra);
extern const char *show_tcp_keepalives_count(void);
-extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending);
+extern void assign_tcp_keepalives_idle(int newval, void *extra);
extern const char *show_tcp_keepalives_idle(void);
-extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending);
+extern void assign_tcp_keepalives_interval(int newval, void *extra);
extern const char *show_tcp_keepalives_interval(void);
-extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending);
+extern void assign_tcp_user_timeout(int newval, void *extra);
extern const char *show_tcp_user_timeout(void);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
@@ -163,7 +163,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
-extern void assign_transaction_timeout(int newval, void *extra, bool *pending);
+extern void assign_transaction_timeout(int newval, void *extra);
extern const char *show_unix_socket_permissions(void);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
--
2.34.1
[application/x-patch] 0018-Re-implement-UI-and-synchronization-for-res-20251013.patch (117.5K, ../../CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail.com/19-0018-Re-implement-UI-and-synchronization-for-res-20251013.patch)
download | inline diff:
From 9f804f7a003d00771304af6f0f4f96a9839571f3 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Fri, 26 Sep 2025 19:12:45 +0530
Subject: [PATCH 18/19] Re-implement UI and synchronization for resizing buffer
pool
shared_buffers is not PGC_SIGHUP instead of PGC_POSTMASTER. The value
of this GUC is saved in NBuffersPending instead of NBuffers which now
shows the size of buffer pool in-effect. When the server starts, the
shared memory size is estimated and the memory is allocated using
NBuffersPending followed by setting NBuffers = NBuffersPending.
When a server is running, the new value of GUC (set using ALTER SYSTEM
... SET shared_buffers = ...; followed by SELECT pg_reload_conf()) does
not come into effect immediately. Instead a function
pg_resize_shared_buffers() is used to resize the buffer pool. The
function uses the current value of GUC in the backends where it is
executed. The function also coordinates the buffer resizing
synchronization across backends.
SHOW shared_buffers now shows the current size of the shared buffers but
it also shows pending size of shared buffers, if any.
A new GUC max_shared_buffers is introduced to control the maximum value
of shared_buffers that can be set. By default it is 0 and it is set to
shared_buffers' value. When explicitly set it needs to be higher than
'shared_buffers'. This GUC determines the size of address space reserved
for future buffer pool sizes and the size of buffer look up table.
TODO: In case the backend executing pg_resize_shared_buffers() exits
before the operation finishes, we will need somebody to clean up or
complete the half-finished resizing operation. Best possibility is to
use a background worker (mostly background writer) to do that. But then
I think making that background worker the coordinator itself might be a
better option since it will be restarted by the postmaster upon
premature exit.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
doc/src/sgml/config.sgml | 44 +-
doc/src/sgml/func/func-admin.sgml | 57 ++
src/backend/access/transam/slru.c | 2 +-
src/backend/access/transam/xlog.c | 2 +-
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/port/sysv_shmem.c | 433 ++--------------
src/backend/postmaster/postmaster.c | 40 +-
src/backend/storage/buffer/buf_init.c | 264 ++++++++--
src/backend/storage/buffer/bufmgr.c | 126 +++--
src/backend/storage/buffer/freelist.c | 118 ++---
src/backend/storage/ipc/ipci.c | 8 +-
src/backend/storage/ipc/procsignal.c | 14 +-
src/backend/storage/ipc/shmem.c | 485 +++++++++++++++++-
src/backend/tcop/postgres.c | 13 +-
.../utils/activity/wait_event_names.txt | 4 +-
src/backend/utils/init/globals.c | 6 +-
src/backend/utils/init/postinit.c | 32 ++
src/backend/utils/misc/guc.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 16 +-
src/include/catalog/pg_proc.dat | 6 +
src/include/miscadmin.h | 6 +-
src/include/storage/buf_internals.h | 2 +-
src/include/storage/bufmgr.h | 17 +-
src/include/storage/ipc.h | 1 -
src/include/storage/pg_shmem.h | 24 +-
src/include/storage/procsignal.h | 5 +-
src/include/storage/shmem.h | 8 +
src/include/utils/guc.h | 2 +
src/test/buffermgr/Makefile | 3 +
src/test/buffermgr/buffermgr_test.conf | 9 +
src/test/buffermgr/expected/buffer_resize.out | 184 +++++--
src/test/buffermgr/meson.build | 5 +
src/test/buffermgr/sql/buffer_resize.sql | 44 +-
src/test/buffermgr/t/001_resize_buffer.pl | 44 +-
.../buffermgr/t/003_parallel_resize_buffer.pl | 71 +++
35 files changed, 1387 insertions(+), 712 deletions(-)
create mode 100644 src/test/buffermgr/buffermgr_test.conf
create mode 100644 src/test/buffermgr/t/003_parallel_resize_buffer.pl
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 39e658b7808..732f9636857 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -1724,7 +1724,6 @@ include_dir 'conf.d'
that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
(Non-default values of <symbol>BLCKSZ</symbol> change the minimum
value.)
- This parameter can only be set at server start.
</para>
<para>
@@ -1747,6 +1746,49 @@ include_dir 'conf.d'
appropriate, so as to leave adequate space for the operating system.
</para>
+ <para>
+ The shared memory consumed by the buffer pool is allocated and
+ initialized according to the value of the GUC at the time of starting
+ the server. A desired new value of GUC can be loaded while the server is
+ running using <systemitem>SIGHUP</systemitem>. But the buffer pool will
+ not be resized immediately. Use
+ <function>pg_resize_shared_buffers()</function> to dynamically resize
+ the shared buffer pool (see <xref linkend="functions-admin"/> for details).
+ <command>SHOW shared_buffers</command> shows the current number of
+ shared buffers and pending number, if any. Please note that when the GUC
+ is changed, the other GUCS which use this GUCs value to set their
+ defaults will not be changed. They may still require a server restart to
+ consider new value.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-max-shared-buffers" xreflabel="max_shared_buffers">
+ <term><varname>max_shared_buffers</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>max_shared_buffers</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Sets the upper limit for the <varname>shared_buffers</varname> value.
+ The default value is <literal>0</literal>,
+ which means no explicit limit is set and <varname>max_shared_buffers</varname>
+ will be automatically set to the value of <varname>shared_buffers</varname>
+ at server startup.
+ If this value is specified without units, it is taken as blocks,
+ that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
+ This parameter can only be set at server start.
+ </para>
+
+ <para>
+ This parameter determines the amount of memory address space to reserve
+ in each backend for expanding the buffer pool in future. While the
+ memory for buffer pool is allocated on demand as it is resized, the
+ memory required to hold the buffer manager metadata is allocated
+ statically at the server start accounting for the largest buffer pool
+ size allowed by this parameter.
+ </para>
</listitem>
</varlistentry>
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 1b465bc8ba7..0dc89b07c76 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -99,6 +99,63 @@
<returnvalue>off</returnvalue>
</para></entry>
</row>
+
+ <row>
+ <entry role="func_table_entry"><para role="func_signature">
+ <indexterm>
+ <primary>pg_resize_shared_buffers</primary>
+ </indexterm>
+ <function>pg_resize_shared_buffers</function> ()
+ <returnvalue>boolean</returnvalue>
+ </para>
+ <para>
+ Dynamically resizes the shared buffer pool to match the current
+ value of the <varname>shared_buffers</varname> parameter. This
+ function implements a coordinated resize process that ensures all
+ backend processes acknowledge the change before completing the
+ operation. The resize happens in multiple phases to maintain
+ data consistency and system stability. Returns <literal>true</literal>
+ if the resize was successful, or raises an error if the operation
+ fails. This function can only be called by superusers.
+ </para>
+ <para>
+ To resize shared buffers, first update the <varname>shared_buffers</varname>
+ setting and reload the configuration, then verify the new value is loaded
+ before calling this function. For example:
+<programlisting>
+postgres=# ALTER SYSTEM SET shared_buffers = '256MB';
+ALTER SYSTEM
+postgres=# SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers;
+ shared_buffers
+-------------------------
+ 128MB (pending: 256MB)
+(1 row)
+
+postgres=# SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+</programlisting>
+ The <command>SHOW shared_buffers</command> step is important to verify
+ that the configuration reload was successful and the new value is
+ available to the current session before attempting the resize. The
+ output shows both the current and pending values when a change is waiting
+ to be applied.
+ </para></entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/access/transam/slru.c b/src/backend/access/transam/slru.c
index 5d3fcd62c94..3eae1d0c7e9 100644
--- a/src/backend/access/transam/slru.c
+++ b/src/backend/access/transam/slru.c
@@ -232,7 +232,7 @@ SimpleLruAutotuneBuffers(int divisor, int max)
{
return Min(max - (max % SLRU_BANK_SIZE),
Max(SLRU_BANK_SIZE,
- NBuffers / divisor - (NBuffers / divisor) % SLRU_BANK_SIZE));
+ NBuffersPending / divisor - (NBuffersPending / divisor) % SLRU_BANK_SIZE));
}
/*
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index eceab341255..ea01befe15c 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4662,7 +4662,7 @@ XLOGChooseNumBuffers(void)
{
int xbuffers;
- xbuffers = NBuffers / 32;
+ xbuffers = NBuffersPending / 32;
if (xbuffers > (wal_segment_size / XLOG_BLCKSZ))
xbuffers = (wal_segment_size / XLOG_BLCKSZ);
if (xbuffers < 8)
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index fc8638c1b61..226944e4588 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -335,6 +335,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ InitializeMaxNBuffers();
+
CreateSharedMemoryAndSemaphores();
/*
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 3be28e228ae..380ecbc9751 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -102,14 +102,8 @@ void *UsedShmemSegAddr = NULL;
AnonymousMapping Mappings[ANON_MAPPINGS];
-/* Flag telling postmaster that resize is needed */
-volatile bool pending_pm_shmem_resize = false;
volatile bool delay_shmem_resize = false;
-/* Keeps track of the previous NBuffers value */
-static int NBuffersOld = -1;
-static int NBuffersPending = -1;
-
/*
* Anonymous mapping layout we use looks like this:
*
@@ -137,20 +131,9 @@ static int NBuffersPending = -1;
* reservation, into which shared memory segment can be extended and is
* represented by the second /memfd:main with no permissions.
*
- * The reserved space for each segment is calculated as a fraction of the total
- * reserved space (MaxAvailableMemory), as specified in the SHMEM_RESIZE_RATIO
- * array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take up to 60% of the whole
- * space when resizing, based on the fact that it most likely will be the main
- * consumer of this memory. Those numbers are pulled out of thin air for now,
- * makes sense to evaluate them more precise.
+ * The reserved space for buffer manager related segments is calculated based on
+ * MaxNBuffers.
*/
-static double SHMEM_RESIZE_RATIO[6] = {
- 0.15, /* MAIN_SHMEM_SEGMENT */
- 0.6, /* BUFFERS_SHMEM_SEGMENT */
- 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */
- 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */
- 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */
-};
/*
* Flag telling that we have decided to use huge pages.
@@ -160,13 +143,6 @@ static double SHMEM_RESIZE_RATIO[6] = {
*/
static bool huge_pages_on = false;
-/*
- * Flag telling that we have prepared the memory layout to be resizable. If
- * false after all shared memory segments creation, it means we failed to setup
- * needed layout and falled back to the regular non-resizable approach.
- */
-static bool shmem_resizable = false;
-
/*
* Currently broadcasted value of NBuffers in shared memory.
*
@@ -791,8 +767,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping)
if (mapping->shmem_reserved < mapping->shmem_size)
ereport(ERROR,
(errcode(ERRCODE_INSUFFICIENT_RESOURCES),
- errmsg("not enough shared memory is reserved"),
- errhint("You may need to increase \"max_available_memory\".")));
+ errmsg("not enough shared memory is reserved")));
mmap_flags = PG_MMAP_FLAGS | mmap_flags;
}
@@ -961,36 +936,27 @@ AnonymousShmemDetach(int status, Datum arg)
}
/*
- * Resize all shared memory segments based on the current NBuffers value, which
- * is is applied from NBuffersPending. The actual segment resizing is done via
- * ftruncate, which will fail if is not sufficient space to expand the anon
- * file. When finished, based on the new and old values initialize new buffer
- * blocks if any.
- *
- * If reinitializing took place, as the last step this function does buffers
- * reinitialization as well and broadcasts the new value of NSharedBuffers. All
- * of that needs to be done only by one backend, the first one that managed to
- * grab the ShmemResizeLock.
+ * Resize all shared memory segments based on the new shared_buffers value (saved
+ * in ShmemCtrl area). The actual segment resizing is done via ftruncate, which
+ * will fail if there is not sufficient space to expand the anon file.
+ *
+ * TODO: Rename this to BufferShmemResize() or something. Only buffer manager's
+ * memory should be resized in this function.
*/
bool
AnonymousShmemResize(void)
{
- int numSemas;
- bool reinit = false;
int mmap_flags = PG_MMAP_FLAGS;
Size hugepagesize;
- NBuffers = NBuffersPending;
-
- elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers);
-
- /*
- * XXX: Where to reset the flag is still an open question. E.g. do we
- * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize
- * and reset the flag?
- */
- pending_pm_shmem_resize = false;
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+ /* TODO: This is a hack. NBuffersPending should never be written by anything
+ * other than GUC system. Find a way to pass new NBuffers value to
+ * BufferManagerShmemSize(). */
+ NBuffersPending = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffers, NBuffersPending);
+
#ifndef MAP_HUGETLB
/* PrepareHugePages should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
@@ -1005,8 +971,8 @@ AnonymousShmemResize(void)
}
#endif
- /* Note that CalculateShmemSize indirectly depends on NBuffers */
- CalculateShmemSize(&numSemas);
+ /* Note that BufferManagerShmemSize() indirectly depends on NBuffersPending. */
+ BufferManagerShmemSize(false);
for(int i = 0; i < ANON_MAPPINGS; i++)
{
@@ -1014,10 +980,18 @@ AnonymousShmemResize(void)
ShmemSegment *segment = &Segments[i];
PGShmemHeader *shmem_hdr = segment->ShmemSegHdr;
+ /* Main shared memory segment is always static. Ignore it. */
+ if (i == MAIN_SHMEM_SEGMENT)
+ continue;
+
+ m->shmem_req_size = add_size(m->shmem_req_size, 8192 - (m->shmem_req_size % 8192));
#ifdef MAP_HUGETLB
if (huge_pages_on && (m->shmem_req_size % hugepagesize != 0))
m->shmem_req_size += hugepagesize - (m->shmem_req_size % hugepagesize);
#endif
+ elog(DEBUG1, "segment[%s]: requested size %zu, current size %zu, reserved %zu",
+ MappingName(m->shmem_segment), m->shmem_req_size, m->shmem_size,
+ m->shmem_reserved);
if (m->shmem == NULL)
continue;
@@ -1025,26 +999,28 @@ AnonymousShmemResize(void)
if (m->shmem_size == m->shmem_req_size)
continue;
+ /* We should have reserved enough address space. Also made sure that the
+ * new size can fit in the existing mapping. PANIC if that's not the
+ * case. */
if (m->shmem_reserved < m->shmem_req_size)
- ereport(ERROR,
+ ereport(PANIC,
(errcode(ERRCODE_INSUFFICIENT_RESOURCES),
- errmsg("not enough shared memory is reserved"),
- errhint("You may need to increase \"max_available_memory\".")));
+ errmsg("not enough shared memory is reserved")));
elog(DEBUG1, "segment[%s]: resize from %zu to %zu at address %p",
MappingName(m->shmem_segment), m->shmem_size,
m->shmem_req_size, m->shmem);
- /* Resize the backing anon file. */
+ /* Resize the backing anon file. If the operation fails, in one backend and we do not know the status in other backends, it will lead to inconsistent buffer manager structures across backends. PANIC. */
if(ftruncate(m->segment_fd, m->shmem_req_size) == -1)
- ereport(FATAL,
+ ereport(PANIC,
(errcode(ERRCODE_SYSTEM_ERROR),
errmsg("could not truncase anonymous file for \"%s\": %m",
MappingName(m->shmem_segment))));
/* Adjust memory accessibility */
if(mprotect(m->shmem, m->shmem_req_size, PROT_READ | PROT_WRITE) == -1)
- ereport(FATAL,
+ ereport(PANIC,
(errcode(ERRCODE_SYSTEM_ERROR),
errmsg("could not mprotect anonymous shared memory for \"%s\": %m",
MappingName(m->shmem_segment))));
@@ -1052,308 +1028,19 @@ AnonymousShmemResize(void)
/* If shrinking, make reserved space unavailable again */
if(m->shmem_req_size < m->shmem_size &&
mprotect(m->shmem + m->shmem_req_size, m->shmem_size - m->shmem_req_size, PROT_NONE) == -1)
- ereport(FATAL,
+ ereport(PANIC,
(errcode(ERRCODE_SYSTEM_ERROR),
errmsg("could not mprotect reserved shared memory for \"%s\": %m",
MappingName(m->shmem_segment))));
- reinit = true;
m->shmem_size = m->shmem_req_size;
shmem_hdr->totalsize = m->shmem_size;
segment->ShmemEnd = m->shmem + m->shmem_size;
}
- if (reinit)
- {
- if(IsUnderPostmaster &&
- LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
- {
- /*
- * If the new NBuffers was already broadcasted, the buffer pool was
- * already initialized before.
- *
- * Since we're not on a hot path, we use lwlocks and do not need to
- * involve memory barrier.
- */
- if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
- {
- /*
- * Allow the first backend that managed to get the lock to
- * reinitialize the new portion of buffer pool. Every other
- * process will wait on the shared barrier for that to finish,
- * since it's a part of the SHMEM_RESIZE_DONE phase.
- *
- * Note that it's enough when only one backend will do that,
- * even the ShmemInitStruct part. The reason is that resized
- * shared memory will maintain the same addresses, meaning that
- * all the pointers are still valid, and we only need to update
- * structures size in the ShmemIndex once -- any other backend
- * will pick up this shared structure from the index.
- */
- BufferManagerShmemInit(NBuffersOld);
-
- /*
- * Wipe out the evictor PID so that it can be used for the next
- * buffer resizing operation.
- */
- ShmemCtrl->evictor_pid = 0;
- /* If all fine, broadcast the new value */
- pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
- }
-
- LWLockRelease(ShmemResizeLock);
- }
- }
-
- return true;
-}
-
-/*
- * We are asked to resize shared memory. Wait for all ProcSignal participants
- * to join the barrier, then do the resize and wait on the barrier until all
- * participating finish resizing as well -- otherwise we face danger of
- * inconsistency between backends.
- *
- * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not
- * proceed with AnonymousShmemResize after receiving SIGHUP, until something
- * will be sent.
- */
-bool
-ProcessBarrierShmemResize(Barrier *barrier)
-{
- Assert(IsUnderPostmaster);
-
- elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d, %d",
- NBuffersOld, NBuffersPending, pending_pm_shmem_resize, delay_shmem_resize);
-
- /* Wait until we have seen the new NBuffers value */
- if (!pending_pm_shmem_resize)
- return false;
-
- /* Wait till this process becomes ready to resize buffers. */
- if (delay_shmem_resize)
- return false;
-
- /*
- * First thing to do after attaching to the barrier is to wait for others.
- * We can't simply use BarrierArriveAndWait, because backends might arrive
- * here in disjoint groups, e.g. first two backends, pause, then second two
- * backends. If the resize is quick enough that can lead to a situation
- * when the first group is already finished before the second has appeared,
- * and the barrier will only synchonize withing those groups.
- */
- if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED)
- WaitForProcSignalBarrierReceived(
- pg_atomic_read_u64(&ShmemCtrl->Generation));
-
- /*
- * Now start the procedure, and elect one backend to ping postmaster to do
- * the same.
- *
- * XXX: If we need to be able to abort resizing, this has to be done later,
- * after the SHMEM_RESIZE_DONE.
- */
-
- /*
- * Evict extra buffers when shrinking shared buffers. We need to do this
- * while the memory for extra buffers is still mapped i.e. before remapping
- * the shared memory segments to a smaller memory area.
- */
- if (NBuffersOld > NBuffersPending)
- {
- BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START);
-
- /*
- * TODO: If the buffer eviction fails for any reason, we should
- * gracefully rollback the shared buffer resizing and try again. But the
- * infrastructure to do so is not available right now. Hence just raise
- * a FATAL so that the system restarts.
- */
- if (!EvictExtraBuffers(NBuffersPending, NBuffersOld))
- elog(FATAL, "buffer eviction failed");
-
- if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_EVICT))
- SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
- }
- else
- if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START))
- SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
-
- AnonymousShmemResize();
-
- /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */
- BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE);
-
- if (MyBackendType == B_BG_WRITER)
- {
- /*
- * Before resuming regular background writer activity, adjust the
- * statistics collected so far.
- */
- BgBufferSyncReset(NBuffersOld, NBuffers);
- }
-
- BarrierDetach(barrier);
return true;
}
-/*
- * GUC assign hook for shared_buffers.
- *
- * When setting the GUC first time after starting the server, the GUC value is
- * changed immediately since there is not shared memory setup yet.
- *
- * After the shared memory is setup, changing the GUC value requires resizing and
- * reiniatializing (at least parts of) the shared memory structures related to
- * shared buffers. That's a long and complicated process. It's recommended for
- * an assign hook to be as minimal as possible, thus we just request shared
- * memory resize and remember the previous value.
- */
-void
-assign_shared_buffers(int newval, void *extra, bool *pending)
-{
- /*
- * TODO: If a backend joins while the buffer resizing is in progress or it
- * reads a value of shared_buffers from configuration which is different from
- * the value being used by existing backends, this method may not work. Need
- * to think of a better solution.
- */
- if (BufferBlocks)
- {
- elog(DEBUG1, "bufferpool is already initialized with size = %d, reinitializing it with size = %d",
- NBuffers, newval);
- pending_pm_shmem_resize = true;
- *pending = true;
- NBuffersPending = newval;
- NBuffersOld = NBuffers;
- }
- else
- {
- elog(DEBUG1, "initializing buffer pool with size = %d", newval);
- NBuffers = newval;
- *pending = false;
- pending_pm_shmem_resize = false;
- }
-}
-
-/*
- * Test if we have somehow missed a shmem resize signal and NBuffers value
- * differs from NSharedBuffers. If yes, catchup and do resize.
- */
-void
-AdjustShmemSize(void)
-{
- uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers);
-
- if (NSharedBuffers != NBuffers)
- {
- /*
- * If the broadcasted shared_buffers is different from the one we see,
- * it could be that the backend has missed a resize signal. To avoid
- * any inconsistency, adjust the shared mappings, before having a
- * chance to access the buffer pool.
- */
- ereport(LOG,
- (errmsg("shared_buffers has been changed from %d to %d, "
- "resize shared memory",
- NBuffers, NSharedBuffers)));
- NBuffers = NSharedBuffers;
- AnonymousShmemResize();
- }
-}
-
-/*
- * Start resizing procedure, making sure all existing processes will have
- * consistent view of shared memory size. Must be called only in postmaster.
- */
-void
-CoordinateShmemResize(void)
-{
- elog(DEBUG1, "Coordinating shmem resize from %d to %d",
- NBuffersOld, NBuffers);
- Assert(!IsUnderPostmaster);
-
- /*
- * We use dynamic barrier to help dealing with backends that were spawned
- * during the resize.
- */
- BarrierInit(&ShmemCtrl->Barrier, 0);
-
- /*
- * If the value did not change, or shared memory segments are not
- * initialized yet, skip the resize.
- */
- if (NBuffersPending == NBuffersOld)
- {
- elog(DEBUG1, "Skip resizing, new %d, old %d",
- NBuffers, NBuffersOld);
- return;
- }
-
- /*
- * Shared memory resize requires some coordination done by postmaster,
- * and consists of three phases:
- *
- * - Before the resize all existing backends have the same old NBuffers.
- * - When resize is in progress, backends are expected to have a
- * mixture of old a new values. They're not allowed to touch buffer
- * pool during this time frame.
- * - After resize has been finished, all existing backends, that can access
- * the buffer pool, are expected to have the same new value of NBuffers.
- *
- * Those phases are ensured by joining the shared barrier associated with
- * the procedure. Since resizing takes time, we need to take into account
- * that during that time:
- *
- * - New backends can be spawned. They will check status of the barrier
- * early during the bootstrap, and wait until everything is over to work
- * with the new NBuffers value.
- *
- * - Old backends can exit before attempting to resize. Synchronization
- * used between backends relies on ProcSignalBarrier and waits for all
- * participants received the message at the beginning to gather all
- * existing backends.
- *
- * - Some backends might be blocked and not responsing either before or
- * after receiving the message. In the first case such backend still
- * have ProcSignalSlot and should be waited for, in the second case
- * shared barrier will make sure we still waiting for those backends. In
- * any case there is an unbounded wait.
- *
- * - Backends might join barrier in disjoint groups with some time in
- * between. That means that relying only on the shared dynamic barrier is
- * not enough -- it will only synchronize resize procedure withing those
- * groups. That's why we wait first for all participants of ProcSignal
- * mechanism who received the message.
- */
- elog(DEBUG1, "Emit a barrier for shmem resizing");
- pg_atomic_init_u64(&ShmemCtrl->Generation,
- EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE));
-
- /* To order everything after setting Generation value */
- pg_memory_barrier();
-
- /*
- * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all
- * the rest of the pack has started the procedure and it can resize shared
- * memory as well.
- *
- * Normally we would call WaitForProcSignalBarrier here to wait until every
- * backend has reported on the ProcSignalBarrier. But for shared memory
- * resize we don't need this, as every participating backend will
- * synchronize on the ProcSignal barrier. In fact even if we would like to
- * wait here, it wouldn't be possible -- we're in the postmaster, without
- * any waiting infrastructure available.
- *
- * If at some point it will turn out that waiting is essential, we would
- * need to consider some alternatives. E.g. it could be a designated
- * coordination process, which is not a postmaster. Another option would be
- * to introduce a CoordinateShmemResize lock and allow only one process to
- * take it (this probably would have to be something different than
- * LWLocks, since they block interrupts, and coordination relies on them).
- */
-}
-
/*
* PGSharedMemoryCreate
*
@@ -1374,7 +1061,7 @@ PGSharedMemoryCreate(AnonymousMapping *mapping,
void *memAddress;
PGShmemHeader *hdr;
struct stat statbuf;
- Size sysvsize, total_reserved;
+ Size sysvsize;
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -1398,12 +1085,6 @@ PGSharedMemoryCreate(AnonymousMapping *mapping,
/* Prepare the mapping information */
mapping->shmem_size = mapping->shmem_req_size;
- total_reserved = (Size) MaxAvailableMemory * BLCKSZ;
- mapping->shmem_reserved = total_reserved * SHMEM_RESIZE_RATIO[mapping->shmem_segment];
-
- /* Round up to be a multiple of BLCKSZ */
- mapping->shmem_reserved = mapping->shmem_reserved + BLCKSZ -
- (mapping->shmem_reserved % BLCKSZ);
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
@@ -1666,29 +1347,6 @@ PGSharedMemoryDetach(void)
}
}
-void
-WaitOnShmemBarrier()
-{
- Barrier *barrier = &ShmemCtrl->Barrier;
-
- /* Nothing to do if resizing is not started */
- if (BarrierPhase(barrier) < SHMEM_RESIZE_START)
- return;
-
- BarrierAttach(barrier);
-
- /* Otherwise wait through all available phases */
- while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE)
- {
- ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting",
- BarrierPhase(barrier))));
-
- BarrierArriveAndWait(barrier, 0);
- }
-
- BarrierDetach(barrier);
-}
-
void
ShmemControlInit(void)
{
@@ -1700,16 +1358,13 @@ ShmemControlInit(void)
if (!foundShmemCtrl)
{
- /*
- * The barrier is missing here, it will be initialized right before
- * starting the resizing process as a convenient way to reset it.
- */
-
- /* Initialize with the currently known value */
- pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers);
-
- /* shmem_resizable should be initialized by now */
- ShmemCtrl->Resizable = shmem_resizable;
- ShmemCtrl->evictor_pid = 0;
+ pg_atomic_init_u32(&ShmemCtrl->targetNBuffers, 0);
+ pg_atomic_init_u32(&ShmemCtrl->activeNBuffers, 0);
+ pg_atomic_init_u32(&ShmemCtrl->transitNBuffers, 0);
+ pg_atomic_init_flag(&ShmemCtrl->resize_in_progress);
+
+ ShmemCtrl->coordinator = 0;
+ ShmemCtrl->pmwork_done = false;
+ ConditionVariableInit(&ShmemCtrl->pm_cv);
}
}
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index ba9528d5dfa..3be146abac2 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -110,9 +110,11 @@
#include "replication/slotsync.h"
#include "replication/walsender.h"
#include "storage/aio_subsys.h"
+#include "storage/bufmgr.h"
#include "storage/fd.h"
#include "storage/io_worker.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "tcop/backend_startup.h"
@@ -125,7 +127,6 @@
#ifdef EXEC_BACKEND
#include "common/file_utils.h"
-#include "storage/pg_shmem.h"
#endif
@@ -959,6 +960,11 @@ PostmasterMain(int argc, char *argv[])
*/
InitializeFastPathLocks();
+ /*
+ * Calculate MaxNBuffers for buffer pool resizing.
+ */
+ InitializeMaxNBuffers();
+
/*
* Give preloaded libraries a chance to request additional shared memory.
*/
@@ -1698,9 +1704,6 @@ ServerLoop(void)
if (pending_pm_pmsignal)
process_pm_pmsignal();
- if (pending_pm_shmem_resize)
- process_pm_shmem_resize();
-
if (events[i].events & WL_SOCKET_ACCEPT)
{
ClientSocket s;
@@ -2046,15 +2049,34 @@ process_pm_reload_request(void)
}
}
+/*
+ * Handle requests from the coordinator to resize shared memory maps so that the
+ * new backends can inherit those.
+ */
static void
process_pm_shmem_resize(void)
{
+ elog(LOG, "postmaster received PMSIGNAL_SHMEM_RESIZE, coordinating memory remapping");
+
/*
- * Failure to resize is considered to be fatal and will not be
- * retried, which means we can disable pending flag right here.
+ * Perform the memory remapping in postmaster process. This should never fail
+ * since the address map is always reserved. If it fails the address maps in
+ * the backends will becomes inconsistent which is a fundamental assumption
+ * in PostgreSQL architecture. Hence PANIC.
*/
- pending_pm_shmem_resize = false;
- CoordinateShmemResize();
+ if (!AnonymousShmemResize())
+ elog(PANIC, "postmaster failed to resize anonymous shared memory");
+ else
+ {
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+ elog(LOG, "postmaster successfully completed shared memory remapping");
+
+ BufferManagerShmemValidate(targetNBuffers);
+ elog(LOG, "postmaster successfully validated buffer manager shared memory");
+ ShmemCtrl->pmwork_done = true;
+ ConditionVariableBroadcast(&ShmemCtrl->pm_cv);
+ NBuffers = targetNBuffers;
+ }
}
/*
@@ -3878,7 +3900,7 @@ process_pm_pmsignal(void)
}
if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE))
- AnonymousShmemResize();
+ process_pm_shmem_resize();
/*
* Try to advance postmaster's state machine, if a child requests it.
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index be64fa5a136..80a168ec2ce 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -19,6 +19,7 @@
#include "storage/pg_shmem.h"
#include "storage/bufmgr.h"
#include "storage/pg_shmem.h"
+#include "utils/guc.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -63,47 +64,39 @@ CkptSortItem *CkptBufferIds;
/*
* Initialize shared buffer pool
*
- * This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend) or during shared-memory resize. Size
- * of data structures initialized here depends on NBuffers, and to be able to
- * change NBuffers without a restart we store each structure into a separate
- * shared memory segment, which could be resized on demand.
- *
- * FirstBufferToInit tells where to start initializing buffers. For
- * initialization it always will be zero, but when resizing shared-memory it
- * indicates the number of already initialized buffers.
- *
+ * This is called once during shared-memory initialization.
+ * TODO: Restore this function to it's initial form. This function should see no
+ * change in buffer resize patches, except may be use of NBuffersPending.
+ *
* No locks are taking in this function, it is the caller responsibility to
* make sure only one backend can work with new buffers.
*/
void
-BufferManagerShmemInit(int FirstBufferToInit)
+BufferManagerShmemInit(void)
{
bool foundBufs,
foundDescs,
foundIOCV,
foundBufCkpt;
int i;
- elog(DEBUG1, "BufferManagerShmemInit from %d to %d",
- FirstBufferToInit, NBuffers);
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
ShmemInitStructInSegment("Buffer Descriptors",
- NBuffers * sizeof(BufferDescPadded),
+ NBuffersPending * sizeof(BufferDescPadded),
&foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
ShmemInitStructInSegment("Buffer Blocks",
- NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ NBuffersPending * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
&foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
ShmemInitStructInSegment("Buffer IO Condition Variables",
- NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ NBuffersPending * sizeof(ConditionVariableMinimallyPadded),
&foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
@@ -115,7 +108,7 @@ BufferManagerShmemInit(int FirstBufferToInit)
*/
CkptBufferIds = (CkptSortItem *)
ShmemInitStructInSegment("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ NBuffersPending * sizeof(CkptSortItem), &foundBufCkpt,
CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
@@ -124,15 +117,14 @@ BufferManagerShmemInit(int FirstBufferToInit)
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
/*
* note: this path is only taken in EXEC_BACKEND case when initializing
- * shared memory, or in all cases when resizing shared memory.
+ * shared memory.
*/
}
-#ifndef EXEC_BACKEND
/*
* Initialize all the buffer headers.
*/
- for (i = FirstBufferToInit; i < NBuffers; i++)
+ for (i = 0; i < NBuffersPending; i++)
{
BufferDesc *buf = GetBufferDescriptor(i);
@@ -150,21 +142,18 @@ BufferManagerShmemInit(int FirstBufferToInit)
ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
-#endif
/*
- * Init other shared buffer-management stuff from scratch configuring buffer
- * pool the first time. If we are just resizing buffer pool adjust only the
- * required structures.
+ * Init other shared buffer-management stuff.
*/
- if (FirstBufferToInit == 0)
- StrategyInitialize(!foundDescs);
- else
- StrategyReInitialize(FirstBufferToInit);
+ StrategyInitialize(!foundDescs);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
&backend_flush_after);
+
+ /* Declare the size of current buffer pool. */
+ NBuffers = NBuffersPending;
}
/*
@@ -175,30 +164,61 @@ BufferManagerShmemInit(int FirstBufferToInit)
* shared memory segment. The main segment must not allocate anything
* related to buffers, every other segment will receive part of the
* data.
+ *
+ * If set_reserved is true, also sets the shmem_reserved field for each
+ * segment based on MaxNBuffers. This should be true during server startup
+ * but false during buffer pool resizing.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(bool set_reserved)
{
size_t size;
/* size of buffer descriptors, plus alignment padding */
- size = add_size(0, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ size = add_size(0, mul_size(NBuffersPending, sizeof(BufferDescPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
Mappings[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_req_size = size;
+ if (set_reserved)
+ {
+ /* reserved size based on MaxNBuffers */
+ size = add_size(0, mul_size(MaxNBuffers, sizeof(BufferDescPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ Mappings[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_reserved = size;
+ }
/* size of data pages, plus alignment padding */
size = add_size(0, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ size = add_size(size, mul_size(NBuffersPending, BLCKSZ));
Mappings[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
+ if (set_reserved)
+ {
+ /* reserved size based on MaxNBuffers */
+ size = add_size(0, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(MaxNBuffers, BLCKSZ));
+ Mappings[BUFFERS_SHMEM_SEGMENT].shmem_reserved = size;
+ }
/* size of I/O condition variables, plus alignment padding */
- size = add_size(0, mul_size(NBuffers,
+ size = add_size(0, mul_size(NBuffersPending,
sizeof(ConditionVariableMinimallyPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
Mappings[BUFFER_IOCV_SHMEM_SEGMENT].shmem_req_size = size;
+ if (set_reserved)
+ {
+ /* reserved size based on MaxNBuffers */
+ size = add_size(0, mul_size(MaxNBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
+ Mappings[BUFFER_IOCV_SHMEM_SEGMENT].shmem_reserved = size;
+ }
/* size of checkpoint sort array in bufmgr.c */
- Mappings[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
+ Mappings[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffersPending, sizeof(CkptSortItem));
+ if (set_reserved)
+ {
+ /* reserved size based on MaxNBuffers */
+ Mappings[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_reserved = mul_size(MaxNBuffers, sizeof(CkptSortItem));
+ }
/* Allocations in the main memory segment, at the end. */
@@ -207,3 +227,181 @@ BufferManagerShmemSize(void)
return size;
}
+
+/*
+ * Reinitialize shared buffer manager structures when resizing the buffer pool.
+ *
+ * This function is called in the backend which coordinates buffer resizing
+ * operation.
+ *
+ * TODO: Avoid code duplication with BufferManagerShmemInit() and also assess
+ * which functionality in the latter is required in this function.
+ */
+void
+BufferManagerShmemResize(int currentNBuffers, int targetNBuffers)
+{
+ bool found;
+ int i;
+ void *tmpPtr;
+
+ tmpPtr = (BufferDescPadded *)
+ ShmemUpdateStructInSegment("Buffer Descriptors",
+ targetNBuffers * sizeof(BufferDescPadded),
+ &found, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+ if (BufferDescriptors != tmpPtr || !found)
+ elog(FATAL, "resizing buffer descriptors failed: expected pointer %p, got %p, found=%d",
+ BufferDescriptors, tmpPtr, found);
+
+ tmpPtr = (ConditionVariableMinimallyPadded *)
+ ShmemUpdateStructInSegment("Buffer IO Condition Variables",
+ targetNBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &found, BUFFER_IOCV_SHMEM_SEGMENT);
+ if (BufferIOCVArray != tmpPtr || !found)
+ elog(FATAL, "resizing buffer IO condition variables failed: expected pointer %p, got %p, found=%d",
+ BufferIOCVArray, tmpPtr, found);
+
+ tmpPtr = (CkptSortItem *)
+ ShmemUpdateStructInSegment("Checkpoint BufferIds",
+ targetNBuffers * sizeof(CkptSortItem), &found,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+ if (CkptBufferIds != tmpPtr || !found)
+ elog(FATAL, "resizing checkpoint buffer IDs failed: expected pointer %p, got %p, found=%d",
+ CkptBufferIds, tmpPtr, found);
+
+ tmpPtr = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemUpdateStructInSegment("Buffer Blocks",
+ targetNBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &found, BUFFERS_SHMEM_SEGMENT));
+ if (BufferBlocks != tmpPtr || !found)
+ elog(FATAL, "resizing buffer blocks failed: expected pointer %p, got %p, found=%d",
+ BufferBlocks, tmpPtr, found);
+
+ /*
+ * Initialize the headers for new buffers. If we are shrinking the
+ * buffers, currentNBuffers >= targetNBuffers, thus this loop doesn't execute.
+ */
+ for (i = currentNBuffers; i < targetNBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
+
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+
+ buf->buf_id = i;
+
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
+
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ }
+
+ StrategyReset(targetNBuffers);
+}
+
+/*
+ * BufferManagerShmemValidate
+ * Validate that buffer manager shared memory structures have correct
+ * pointers and sizes after a resize operation.
+ *
+ * This function is called by backends during ProcessBarrierShmemResizeStruct
+ * to ensure their view of the buffer structures is consistent after memory
+ * remapping.
+ */
+void
+BufferManagerShmemValidate(int targetNBuffers)
+{
+ bool found;
+ void *tmpPtr;
+
+ /* Validate Buffer Descriptors */
+ tmpPtr = (BufferDescPadded *)
+ ShmemInitStructInSegment("Buffer Descriptors",
+ targetNBuffers * sizeof(BufferDescPadded),
+ &found, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+ if (!found || BufferDescriptors != tmpPtr)
+ elog(FATAL, "validating buffer descriptors failed: expected pointer %p, got %p, found=%d",
+ BufferDescriptors, tmpPtr, found);
+
+ /* Validate Buffer IO Condition Variables */
+ tmpPtr = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ targetNBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &found, BUFFER_IOCV_SHMEM_SEGMENT);
+ if (!found || BufferIOCVArray != tmpPtr)
+ elog(FATAL, "validating buffer IO condition variables failed: expected pointer %p, got %p, found=%d",
+ BufferIOCVArray, tmpPtr, found);
+
+ /* Validate Checkpoint BufferIds */
+ tmpPtr = (CkptSortItem *)
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ targetNBuffers * sizeof(CkptSortItem), &found,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+ if (!found || CkptBufferIds != tmpPtr)
+ elog(FATAL, "validating checkpoint buffer IDs failed: expected pointer %p, got %p, found=%d",
+ CkptBufferIds, tmpPtr, found);
+
+ /* Validate Buffer Blocks */
+ tmpPtr = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSegment("Buffer Blocks",
+ targetNBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &found, BUFFERS_SHMEM_SEGMENT));
+ if (!found || BufferBlocks != tmpPtr)
+ elog(FATAL, "validating buffer blocks failed: expected pointer %p, got %p, found=%d",
+ BufferBlocks, tmpPtr, found);
+}
+
+/*
+ * check_shared_buffers
+ * GUC check_hook for shared_buffers
+ *
+ * When reloading the configuration, shared_buffers should not be set to a value
+ * higher than max_shared_buffers fixed at the boot time.
+ */
+bool
+check_shared_buffers(int *newval, void **extra, GucSource source)
+{
+ if (finalMaxNBuffers && *newval > MaxNBuffers)
+ {
+ GUC_check_errdetail("\"shared_buffers\" must be less than \"max_shared_buffers\".");
+ return false;
+ }
+ return true;
+}
+
+/*
+ * show_shared_buffers
+ * GUC show_hook for shared_buffers
+ *
+ * Shows both current and pending buffer counts with proper unit formatting.
+ */
+const char *
+show_shared_buffers(void)
+{
+ static char buffer[128];
+ int64 current_value, pending_value;
+ const char *current_unit, *pending_unit;
+
+ if (NBuffers == NBuffersPending)
+ {
+ /* No buffer pool resizing pending. */
+ convert_int_from_base_unit(NBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s", current_value, current_unit);
+ }
+ else
+ {
+ /*
+ * New value for NBuffers is loaded but not applied yet, show both
+ * current and pending.
+ */
+ convert_int_from_base_unit(NBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ convert_int_from_base_unit(NBuffersPending, GUC_UNIT_BLOCKS, &pending_value, &pending_unit);
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s (pending: " INT64_FORMAT "%s)",
+ current_value, current_unit, pending_value, pending_unit);
+ }
+
+ return buffer;
+}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index fdcb5556235..14200a38a0f 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3631,12 +3631,12 @@ static float smoothed_alloc = 0;
static float smoothed_density = 10.0;
void
-BgBufferSyncReset(int NBuffersOld, int NBuffersNew)
+BgBufferSyncReset(int currentNBuffers, int targetNBuffers)
{
saved_info_valid = false;
#ifdef BGW_DEBUG
elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
- NBuffersOld, NBuffersNew);
+ currentNBuffers, targetNBuffers);
#endif
}
@@ -3686,8 +3686,11 @@ BgBufferSync(WritebackContext *wb_context)
* valid. If the buffer pool is being expanded, more buffers will become
* available without even this function writing out any. Hence wait till
* buffer resizing finishes i.e. go into hibernation mode.
+ *
+ * TODO: We may not need this synchronization if background worker itself
+ * becomes the coordinator.
*/
- if (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers)
+ if (!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress))
return true;
/*
@@ -3883,7 +3886,7 @@ BgBufferSync(WritebackContext *wb_context)
* finish.
*/
while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
- pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) == NBuffers)
+ !pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress))
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -4255,7 +4258,23 @@ DebugPrintBufferRefcount(Buffer buffer)
void
CheckPointBuffers(int flags)
{
+ /* Mark that buffer sync is in progress - delay any shared memory resizing. */
+ /*
+ * TODO: We need to assess whether we should allow checkpoint and buffer
+ * resizing to run in parallel. When expanding buffers it may be fine to let
+ * the checkpointer run in RESIZE_MAP_AND_MEM phase but delay phase EXPAND
+ * phase till the checkpoint finishes, at the same time not allow checkpoint
+ * to run during expansion phase. When shrinking the buffers, we should
+ * delay SHRINK phase till checkpoint finishes and not allow to start
+ * checkpoint till SHRINK phase is done, but allow it to run in
+ * RESIZE_MAP_AND_MEM phase. This needs careful analysis and testing.
+ */
+ delay_shmem_resize = true;
+
BufferSync(flags);
+
+ /* Mark that buffer sync is no longer in progress - allow shared memory resizing */
+ delay_shmem_resize = false;
}
/*
@@ -7504,10 +7523,12 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
* of the shrunk buffer pool.
*/
bool
-EvictExtraBuffers(int newBufSize, int oldBufSize)
+EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
{
bool result = true;
+ Assert(targetNBuffers < currentNBuffers);
+
/*
* If the buffer being evicated is locked, this function will need to wait.
* This function should not be called from a Postmaster since it can not wait on a lock.
@@ -7515,77 +7536,50 @@ EvictExtraBuffers(int newBufSize, int oldBufSize)
Assert(IsUnderPostmaster);
/*
- * Let only one backend perform eviction. We could split the work across all
- * the backends but that doesn't seem necessary.
- *
- * The first backend to acquire ShmemResizeLock, sets its own PID as the
- * evictor PID for other backends to know that the eviction is in progress or
- * has already been performed. The evictor backend releases the lock when it
- * finishes eviction. While the eviction is in progress, backends other than
- * evictor backend won't be able to take the lock. They won't perform
- * eviction. A backend may acquire the lock after eviction has completed, but
- * it will not perform eviction since the evictor PID is already set. Evictor
- * PID is reset only when the buffer resizing finishes. Thus only one backend
- * will perform eviction in a given instance of shared buffers resizing.
- *
- * Any backend which acquires this lock will release it before the eviction
- * phase finishes, hence the same lock can be reused for the next phase of
- * resizing buffers.
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
*/
- if (LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE))
+ for (Buffer buf = targetNBuffers + 1; buf <= currentNBuffers; buf++)
{
- if (ShmemCtrl->evictor_pid == 0)
- {
- ShmemCtrl->evictor_pid = MyProcPid;
-
- /*
- * TODO: Before evicting any buffer, we should check whether any of the
- * buffers are pinned. If we find that a buffer is pinned after evicting
- * most of them, that will impact performance since all those evicted
- * buffers might need to be read again.
- */
- for (Buffer buf = newBufSize + 1; buf <= oldBufSize; buf++)
- {
- BufferDesc *desc = GetBufferDescriptor(buf - 1);
- uint32 buf_state;
- bool buffer_flushed;
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
+ bool buffer_flushed;
- buf_state = pg_atomic_read_u32(&desc->state);
+ buf_state = pg_atomic_read_u32(&desc->state);
- /*
- * Nobody is expected to touch the buffers while resizing is
- * going one hence unlocked precheck should be safe and saves
- * some cycles.
- */
- if (!(buf_state & BM_VALID))
- continue;
+ /*
+ * Nobody is expected to touch the buffers while resizing is
+ * going one hence unlocked precheck should be safe and saves
+ * some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
- /*
- * XXX: Looks like CurrentResourceOwner can be NULL here, find
- * another one in that case?
- * */
- if (CurrentResourceOwner)
- ResourceOwnerEnlarge(CurrentResourceOwner);
+ /*
+ * XXX: Looks like CurrentResourceOwner can be NULL here, find
+ * another one in that case?
+ * */
+ if (CurrentResourceOwner)
+ ResourceOwnerEnlarge(CurrentResourceOwner);
- ReservePrivateRefCountEntry();
+ ReservePrivateRefCountEntry();
- LockBufHdr(desc);
+ LockBufHdr(desc);
- /*
- * Now that we have locked buffer descriptor, make sure that the
- * buffer without valid data has been skipped above.
- */
- Assert(buf_state & BM_VALID);
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
- if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
- {
- elog(WARNING, "could not remove buffer %u, it is pinned", buf);
- result = false;
- break;
- }
- }
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
}
- LWLockRelease(ShmemResizeLock);
}
return result;
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 55be5eebe0a..c09875934d4 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -33,10 +33,16 @@ typedef struct
/* Spinlock: protects the values below */
slock_t buffer_strategy_lock;
+ /*
+ * Number of active buffers that can be allocated. During buffer resizing,
+ * this may be different from NBuffers which tracks the global buffer count.
+ */
+ pg_atomic_uint32 activeNBuffers;
+
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo NBuffers.
+ * get an actual buffer, it needs to be used modulo activeNBuffers.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -101,6 +107,7 @@ static inline uint32
ClockSweepTick(void)
{
uint32 victim;
+ int activeBuffers;
/*
* Atomically move hand ahead one buffer - if there's several processes
@@ -110,12 +117,15 @@ ClockSweepTick(void)
victim =
pg_atomic_fetch_add_u32(&StrategyControl->nextVictimBuffer, 1);
- if (victim >= NBuffers)
+ /* Read the current active buffer count atomically */
+ activeBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+
+ if (victim >= activeBuffers)
{
uint32 originalVictim = victim;
/* always wrap what we look up in BufferDescriptors */
- victim = victim % NBuffers;
+ victim = victim % activeBuffers;
/*
* If we're the one that just caused a wraparound, force
@@ -143,7 +153,7 @@ ClockSweepTick(void)
*/
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
- wrapped = expected % NBuffers;
+ wrapped = expected % activeBuffers;
success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
&expected, wrapped);
@@ -177,6 +187,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
BufferDesc *buf;
int bgwprocno;
int trycounter;
+ int activeNBuffers;
*from_ring = false;
@@ -228,7 +239,9 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
pg_atomic_fetch_add_u32(&StrategyControl->numBufferAllocs, 1);
/* Use the "clock sweep" algorithm to find a free buffer */
- trycounter = NBuffers;
+ activeNBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+ trycounter = activeNBuffers;
+
for (;;)
{
uint32 old_buf_state;
@@ -280,7 +293,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
local_buf_state))
{
- trycounter = NBuffers;
+ trycounter = activeNBuffers;
break;
}
}
@@ -323,10 +336,12 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
{
uint32 nextVictimBuffer;
int result;
+ uint32 activeNBuffers;
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
- result = nextVictimBuffer % NBuffers;
+ activeNBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+ result = nextVictimBuffer % activeNBuffers;
if (complete_passes)
{
@@ -336,7 +351,7 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
* Additionally add the number of wraparounds that happened before
* completePasses could be incremented. C.f. ClockSweepTick().
*/
- *complete_passes += nextVictimBuffer / NBuffers;
+ *complete_passes += nextVictimBuffer / activeNBuffers;
}
if (num_buf_alloc)
@@ -391,6 +406,31 @@ StrategyShmemSize(void)
return size;
}
+void
+StrategyReset(int activeNBuffers)
+{
+ Assert(StrategyControl);
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ /* Update the active buffer count for the strategy */
+ pg_atomic_write_u32(&StrategyControl->activeNBuffers, activeNBuffers);
+
+ /* Reset the clock-sweep pointer to start from beginning */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The statistics is viewed in the context of the number of shared buffers.
+ * Reset it as the size of active number of shared buffers changes.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_write_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* TODO: Do we need to seset background writer notifications? */
+ StrategyControl->bgwprocno = -1;
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+}
+
/*
* StrategyInitialize -- initialize the buffer cache replacement
* strategy.
@@ -416,13 +456,13 @@ StrategyInitialize(bool init)
* directory without rehashing all the entries. Just allocating more entries
* will lead to more contention. Hence we setup the buffer lookup table
* considering the maximum possible size of the buffer pool which is
- * MaxAvailableMemory.
+ * MaxNBuffers.
*
* Additionally BufferAlloc() tries to insert a new entry before deleting the
* old. In principle this could be happening in each partition concurrently,
* so we need extra NUM_BUFFER_PARTITIONS entries.
*/
- InitBufTable(MaxAvailableMemory + NUM_BUFFER_PARTITIONS);
+ InitBufTable(MaxNBuffers + NUM_BUFFER_PARTITIONS);
/*
* Get or create the shared strategy control block
@@ -441,6 +481,8 @@ StrategyInitialize(bool init)
SpinLockInit(&StrategyControl->buffer_strategy_lock);
+ /* Initialize the active buffer count */
+ pg_atomic_init_u32(&StrategyControl->activeNBuffers, NBuffersPending);
/* Initialize the clock-sweep pointer */
pg_atomic_init_u32(&StrategyControl->nextVictimBuffer, 0);
@@ -455,62 +497,6 @@ StrategyInitialize(bool init)
Assert(!init);
}
-/*
- * StrategyReInitialize -- re-initialize the buffer cache replacement
- * strategy.
- *
- * To be called when resizing buffer manager and only from the coordinator.
- * TODO: Assess the differences between this function and StrategyInitialize().
- */
-void
-StrategyReInitialize(int FirstBufferIdToInit)
-{
- bool found;
-
- /*
- * Resizing memory for buffer pools should not affect the address of
- * StrategyControl.
- */
- if (StrategyControl != (BufferStrategyControl *)
- ShmemInitStructInSegment("Buffer Strategy Status",
- sizeof(BufferStrategyControl),
- &found, MAIN_SHMEM_SEGMENT))
- elog(FATAL, "something went wrong while re-initializing the buffer strategy");
-
- Assert(found);
-
- /* TODO: Buffer lookup table adjustment: There are two options:
- *
- * 1. Resize the buffer lookup table to match the new number of buffers. But
- * this requires rehashing all the entries in the buffer lookup table with
- * the new table size.
- *
- * 2. Allocate maximum size of the buffer lookup table at the beginning and
- * never resize it. This leaves sparse buffer lookup table which is
- * inefficient from both memory and time perspective. According to David
- * Rowley, the sparse entries in the buffer look up table cause frequent
- * cacheline reload which affect performance. If the impact of that
- * inefficiency in a benchmark is significant, we will need to consider first
- * option.
- */
- /*
- * The clock sweep tick pointer might have got invalidated. Reset it as if
- * starting a fresh server.
- */
- pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
-
- /*
- * The old statistics is viewed in the context of the number of shared
- * buffers. It does not make sense now that the number of shared buffers
- * itself has changed.
- */
- StrategyControl->completePasses = 0;
- pg_atomic_init_u32(&StrategyControl->numBufferAllocs, 0);
-
- /* No pending notification */
- StrategyControl->bgwprocno = -1;
-}
-
/* ----------------------------------------------------------------
* Backend-private buffer ring management
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index bd75f06047e..cfd952e621e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -133,7 +133,7 @@ CalculateShmemSize(int *num_semaphores)
* memory segment that it uses in the corresponding AnonymousMappings.
* Consider size required from only the main shared memory segment here.
*/
- size = add_size(size, BufferManagerShmemSize());
+ size = add_size(size, BufferManagerShmemSize(true));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
@@ -187,10 +187,14 @@ CalculateShmemSize(int *num_semaphores)
* shared memory segment.
*/
Mappings[MAIN_SHMEM_SEGMENT].shmem_req_size = size;
+ Mappings[MAIN_SHMEM_SEGMENT].shmem_reserved = size;
/* might as well round it off to a multiple of a typical page size */
for (int segment = 0; segment < ANON_MAPPINGS; segment++)
+ {
Mappings[segment].shmem_req_size = add_size(Mappings[segment].shmem_req_size, 8192 - (Mappings[segment].shmem_req_size % 8192));
+ Mappings[segment].shmem_reserved = add_size(Mappings[segment].shmem_reserved, 8192 - (Mappings[segment].shmem_reserved % 8192));
+ }
return size;
}
@@ -341,7 +345,7 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
- BufferManagerShmemInit(0);
+ BufferManagerShmemInit();
/*
* Set up lock manager
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 2160d258fa7..0a173f038a3 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -657,9 +657,17 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
- case PROCSIGNAL_BARRIER_SHMEM_RESIZE:
- processed = ProcessBarrierShmemResize(
- &ShmemCtrl->Barrier);
+ case PROCSIGNAL_BARRIER_SHBUF_SHRINK:
+ processed = ProcessBarrierShmemShrink();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM:
+ processed = ProcessBarrierShmemResizeMapAndMem();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_EXPAND:
+ processed = ProcessBarrierShmemExpand();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED:
+ processed = ProcessBarrierShmemResizeFailed();
break;
}
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 0f9abf69fd5..9793d27042a 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -69,11 +69,19 @@
#include "funcapi.h"
#include "miscadmin.h"
#include "port/pg_numa.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
#include "storage/shmem.h"
#include "storage/spin.h"
#include "utils/builtins.h"
+#include "utils/injection_point.h"
+#include "utils/wait_event.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
@@ -498,28 +506,15 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. Verify the structure's size:
- * - If it's the same, we've found the expected structure.
- * - If it's different, we're resizing the expected structure.
- *
- * XXX: There is an implicit assumption this can only happen in
- * "resizable" segments, where only one shared structure is allowed.
- * This has to be implemented more cleanly. Probably we should implement
- * ShmemReallocRawInSegment functionality just to adjust the size
- * according to alignment, return the allocated size and update the
- * mapping offset.
+ * already. The size better be the same as the size we are trying to
*/
if (result->size != size)
{
- Size delta = size - result->size;
-
- result->size = size;
- result->allocated_size = size;
-
- /* Reflect size change in the shared segment */
- SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
- SpinLockRelease(Segments[shmem_segment].ShmemLock);
+ LWLockRelease(ShmemIndexLock);
+ ereport(ERROR,
+ (errmsg("ShmemIndex entry size is wrong for data structure"
+ " \"%s\": expected %zu, actual %zu",
+ name, size, result->size)));
}
structPtr = result->location;
@@ -556,6 +551,59 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
return structPtr;
}
+/*
+ * ShmemUpdateStructInSegment -- Update the size of a structure in shared memory.
+ *
+ * This function updates the size of an existing shared memory structure. It
+ * finds the structure in the shmem index and updates its size information while
+ * preserving the existing memory location.
+ *
+ * Returns: pointer to the existing structure location.
+ */
+void *
+ShmemUpdateStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
+{
+ ShmemIndexEnt *result;
+ void *structPtr;
+ Size delta;
+
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+
+ Assert(ShmemIndex);
+
+ /* Look up the structure in the shmem index */
+ result = (ShmemIndexEnt *)
+ hash_search(ShmemIndex, name, HASH_FIND, foundPtr);
+
+ Assert(*foundPtr);
+ Assert(result);
+ Assert(result->shmem_segment == shmem_segment);
+
+ delta = size - result->size;
+ /* Store the existing structure pointer */
+ structPtr = result->location;
+
+ /* Update the size information.
+ TODO: Ideally we should implement repalloc kind of functionality for shared memory which will return allocated size. */
+ result->size = size;
+ result->allocated_size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
+ LWLockRelease(ShmemIndexLock);
+
+ /* Verify the structure is still in the correct segment */
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
+ Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
+
+ return structPtr;
+}
+
+
+
/*
* Add two Size values, checking for overflow
*/
@@ -892,3 +940,402 @@ pg_get_shmem_segments(PG_FUNCTION_ARGS)
return (Datum) 0;
}
+/*
+ * TODO: The function henceforth are related to buffer manager and better be
+ * placed in buffer manager related file.
+ */
+
+/*
+ * Prepare ShmemCtrl for resizing the shared buffer pool.
+ */
+static void
+MarkBufferResizingStart(int targetNBuffers, int currentNBuffers)
+{
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, currentNBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, targetNBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->activeNBuffers, Min(targetNBuffers, currentNBuffers));
+ pg_atomic_write_u32(&ShmemCtrl->transitNBuffers, currentNBuffers);
+ ShmemCtrl->coordinator = MyProcPid;
+ ShmemCtrl->pmwork_done = false;
+}
+
+/*
+ * Reset ShmemCtrl after resizing the shared buffer pool is done.
+ */
+static void
+MarkBufferResizingEnd(int NBuffers)
+{
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, NBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, NBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->activeNBuffers, NBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->transitNBuffers, NBuffers);
+ ShmemCtrl->coordinator = -1;
+ ShmemCtrl->pmwork_done = false;
+}
+
+/*
+ * Function which updates the shared buffers according to the current values of
+ * shared_buffers GUCs.
+ *
+ * When resizing the buffer pool is divided into two portions
+ *
+ * - active buffer pool, which is the part of buffer pool which remains active
+ * even during resizing. Its size is given by activeNBuffers. Newly allocated
+ * buffers will have their buffer ids less than activeNBuffers.
+ *
+ * - in-transit buffer pool, which is the part of buffer pool which may be
+ * accessible to some backends but not others. When shrinking the buffer pool
+ * this is the part of buffer pool which will be evicted. When expanding the
+ * buffer pool this is the expanded portion. Its size is given by
+ * transitNBuffers. The backends may see buffer ids upto transitNBuffers.
+ *
+ * Before starting resizing, activeNBuffers = transitNBuffers = NBuffers. And
+ * NewNBuffers is the new size of shared buffer pool.
+ *
+ * In order to synchronize with other running backends, the coordinator sends
+ * following ProcSignalBarriers in the order given below:
+ *
+ * 1. When shrinking the shared buffer pool (with size NBuffers), the coordinator
+ * sends SHBUF_SHRINK ProcSignalBarrier. Every backend sets activeNBuffers =
+ * NewNBuffers to restrict its buffer pool allocations to the new size of the
+ * buffer pool and acknowledges the ProcSignalBarrrier. Once every backend has
+ * acknowledged, the coordinator evicts the buffers in the area being shrunk.
+ * Note that tansitNBuffers is still NBuffers, so the backends may see buffer ids
+ * upto NBuffers from earlier allocations.
+ *
+ * 2. In both cases, when expanding the buffer pool or shrinking the buffer pool,
+ * the coordinator sends SHBUF_RESIZE_MAP_AND_MEM ProcSignalBarrier. Every
+ * backend is expected to adjust their shared memory segment maps (by calling
+ * AnonymousShmemResize()) and validate that their pointers to the shared buffers
+ * structure are valid and have the right size. When shrinking shared buffer pool
+ * transitNBuffers is set to NewNBuffers and the backends should no more see
+ * buffer ids beyond NewNBuffers. When expanding they should also set
+ * transitNBuffers to NewNBuffers to accomodate backends which may accept the
+ * next barrier earlier than the others. After this the backends should
+ * acknowledge the ProcSignalBarrier.
+ *
+ * 3. When expanding the buffer pool, the coordinator sends SHBUF_EXPAND
+ * ProcSignalBarrier. The backends are expected to set activeNBuffers =
+ * NewNBuffers and start allocating buffers from the expanded range.
+ *
+ * Find a better place for this function, also a name if we find this interface
+ * viable.
+ *
+ * TODO: Should this function be in bufmgr.c?
+ *
+ * TODO: Handle the case when the backend executing this function dies or the
+ * query is cancelled.
+ */
+Datum
+pg_resize_shared_buffers(PG_FUNCTION_ARGS)
+{
+ bool result = true;
+ int currentNBuffers = NBuffers;
+ int targetNBuffers = NBuffersPending;
+
+ if (currentNBuffers == targetNBuffers)
+ {
+ elog(LOG, "shared buffers are already at %d, no need to resize", currentNBuffers);
+ PG_RETURN_BOOL(true);
+ }
+
+ if (!pg_atomic_test_set_flag(&ShmemCtrl->resize_in_progress))
+ {
+ elog(LOG, "shared buffer resizing already in progress");
+ PG_RETURN_BOOL(false);
+ }
+
+ MarkBufferResizingStart(targetNBuffers, currentNBuffers);
+ elog(LOG, "resizing shared buffers from %d to %d", currentNBuffers, targetNBuffers);
+
+ INJECTION_POINT("pg-resize-shared-buffers-flag-set", NULL);
+
+ /* Phase 1: SHBUF_SHRINK - Only for shrinking buffer pool */
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * Phase 1: Shrinking - send SHBUF_SHRINK barrier
+ * Every backend sets activeNBuffers = NewNBuffers to restrict
+ * buffer pool allocations to the new size
+ */
+ elog(LOG, "Phase 1: Shrinking buffer pool, restricting allocations to %d buffers", targetNBuffers);
+
+ WaitForProcSignalBarrier(EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHBUF_SHRINK));
+ elog(LOG, "all backends acknowledged shrink phase");
+
+ /* Evict buffers in the area being shrunk */
+ elog(LOG, "evicting buffers %u..%u", targetNBuffers + 1, currentNBuffers);
+ if (!EvictExtraBuffers(targetNBuffers, currentNBuffers))
+ {
+ elog(ERROR, "failed to evict extra buffers during shrinking");
+ WaitForProcSignalBarrier(EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED));
+ MarkBufferResizingEnd(currentNBuffers);
+ pg_atomic_clear_flag(&ShmemCtrl->resize_in_progress);
+ Assert(NBuffers == currentNBuffers);
+ NBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+ PG_RETURN_BOOL(false);
+ }
+
+ /* This backend handles NBuffers itself instead of relying on the barrier
+ * handler, so that barrier handlers do not interfere with its
+ * operations. */
+ NBuffers = targetNBuffers;
+ }
+
+ /* Phase 2: SHBUF_RESIZE_MAP_AND_MEM - Both expanding and shrinking */
+ elog(LOG, "Phase 2: Remapping shared memory segments and updating structures");
+ if (!AnonymousShmemResize())
+ {
+ /*
+ * This should never fail since address map should already be reserved.
+ * So the failure should be treated as PANIC.
+ */
+ elog(PANIC, "failed to resize anonymous shared memory");
+ }
+
+ /* When shrinking no backends should see buffers beyond active portion of the
+ * buffer pool. When expanding, update transitNBuffers so backends can see
+ * the new range. */
+ pg_atomic_write_u32(&ShmemCtrl->transitNBuffers, targetNBuffers);
+
+ /* Update structure pointers and sizes */
+ BufferManagerShmemResize(currentNBuffers, targetNBuffers);
+
+ /* Request Postmaster to remap and resize. TODO: Handle the case when Postmaster is not able to remap and resize the shared memory structures. */
+ SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE);
+ elog(LOG, "waiting for the postmaster to finish remapping and resizing the shared buffers");
+ while (!ShmemCtrl->pmwork_done)
+ {
+ if (ConditionVariableTimedSleep(&ShmemCtrl->pm_cv,
+ 5000,
+ WAIT_EVENT_PM_BUFFER_RESIZE_WAIT))
+ ereport(LOG,
+ (errmsg("still waiting for the postmaster PID %d to finish resizing buffers",
+ (int) PostmasterPid)));
+ }
+ ConditionVariableCancelSleep();
+ elog(LOG, "postmaster remapped and resized the shared memory");
+
+ WaitForProcSignalBarrier(EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM));
+ elog(LOG, "all backends acknowledged memory remapping and structure updates");
+
+ /* Phase 3: SHBUF_EXPAND - Only for expanding buffer pool */
+ if (targetNBuffers > currentNBuffers)
+ {
+ /*
+ * Phase 3: Expanding - send SHBUF_EXPAND barrier
+ * Backends set activeNBuffers = NewNBuffers and start allocating
+ * buffers from the expanded range
+ */
+ elog(LOG, "Phase 3: Expanding buffer pool, enabling allocations up to %d buffers", targetNBuffers);
+
+ WaitForProcSignalBarrier(EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHBUF_EXPAND));
+ elog(LOG, "all backends acknowledged expand phase");
+
+ /* This backend handles NBuffers itself instead of relying on the barrier
+ * handler, so that barrier handlers do not interfere with its
+ * operations. */
+ NBuffers = targetNBuffers;
+ }
+
+ /*
+ * Reset buffer resize control area.
+ */
+ MarkBufferResizingEnd(targetNBuffers);
+
+ pg_atomic_clear_flag(&ShmemCtrl->resize_in_progress);
+
+ elog(LOG, "successfully resized shared buffers to %d", targetNBuffers);
+
+ PG_RETURN_BOOL(result);
+}
+
+bool
+ProcessBarrierShmemShrink(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+ int activeNBuffers = pg_atomic_read_u32(&ShmemCtrl->activeNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /* The work to be done by the coordinator is done in the function which sends the barriers. Hence acknowledge immediately. */
+ if (ShmemCtrl->coordinator == MyProcPid)
+ {
+ elog(LOG, "Phase 1: Coordinator backend %d acknowledging SHBUF_SHRINK barrier immediately", MyProcPid);
+ return true;
+ }
+
+ /*
+ * Delay adjusting the new active size of buffer pool till this process
+ * becomes ready to resize buffers.
+ */
+ if (delay_shmem_resize)
+ {
+ elog(LOG, "Phase 1: Delaying SHBUF_SHRINK barrier - restricting allocations from %d to %d buffers, coordinator is %d",
+ NBuffers, targetNBuffers, ShmemCtrl->coordinator);
+
+ return false;
+ }
+
+ elog(LOG, "Phase 1: Processing SHBUF_SHRINK barrier - restricting allocations from %d to %d buffers, coordinator is %d",
+ NBuffers, targetNBuffers, ShmemCtrl->coordinator);
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * Before resuming regular background writer activity, adjust the
+ * statistics collected so far.
+ */
+ BgBufferSyncReset(NBuffers, targetNBuffers);
+ /* Reset strategy control to new size */
+ StrategyReset(targetNBuffers);
+ }
+
+ /* Update local knowledge of activeNBuffers */
+ NBuffers = activeNBuffers;
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemResizeMapAndMem(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+#ifdef USE_ASSERT_CHECKING
+ int activeNBuffers = pg_atomic_read_u32(&ShmemCtrl->activeNBuffers);
+ int transitNBuffers = pg_atomic_read_u32(&ShmemCtrl->transitNBuffers);
+#endif /* USE_ASSERT_CHECKING */
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /* The work to be done by the coordinator is done in the function which sends the barriers. Hence acknowledge immediately. */
+ if (ShmemCtrl->coordinator == MyProcPid)
+ {
+ elog(LOG, "Phase 2: Coordinator backend %d acknowledging SHBUF_RESIZE_MAP_AND_MEM barrier immediately", MyProcPid);
+ return true;
+ }
+
+ /*
+ * If buffer pool is being shrunk, we are already working with a smaller
+ * buffer pool, so shrinking address space and shared structures should not
+ * be a problem. When expanding, expanding the address space and shared
+ * structures beyond the current boundaries is not going to be a problem
+ * since we are not accessing that memory yet. So there is no reason to
+ * delay processing this barrier.
+ */
+
+ elog(LOG, "Phase 2: Processing SHBUF_RESIZE_MAP_AND_MEM barrier - adjusting memory maps and validating structures, coordinator is %d",
+ ShmemCtrl->coordinator);
+
+ /*
+ * NBuffers should already be set to activeNBuffers from Phase 1.
+ * When shrinking, NBuffers should also be same as transitNBuffers in this phase.
+ */
+ Assert(NBuffers == activeNBuffers);
+ if (targetNBuffers < pg_atomic_read_u32(&ShmemCtrl->currentNBuffers))
+ {
+ /* Shrinking case - verify NBuffers equals transitNBuffers */
+ Assert(NBuffers == transitNBuffers);
+ }
+
+ /*
+ * Address space should already be reserved so resizing should not fail. If
+ * it fails, the address map of this backend may go out of sync with other
+ * backends. Hence PANIC.
+ */
+ if (!AnonymousShmemResize())
+ elog(PANIC, "failed to resize anonymous shared memory in backend %d", MyProcPid);
+
+ elog(LOG, "Backend %d successfully remapped shared memory segments for buffer resize", MyProcPid);
+
+ /*
+ * Backends validate that their pointers to shared buffer structures are
+ * still valid and have the correct size after memory remapping.
+ */
+ BufferManagerShmemValidate(targetNBuffers);
+
+ /*
+ * TODO: Save new transitNBuffers value in process local memory, if
+ * necessary.
+ */
+ elog(LOG, "Backend %d successfully validated structure pointers after resize", MyProcPid);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemExpand(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+#ifdef USE_ASSERT_CHECKING
+ int transitNBuffers = pg_atomic_read_u32(&ShmemCtrl->transitNBuffers);
+#endif /* USE_ASSERT_CHECKING */
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /* The work to be done by the coordinator is done in the function which sends the barriers. Hence acknowledge immediately. */
+ if (ShmemCtrl->coordinator == MyProcPid)
+ {
+ elog(LOG, "Phase 3: Coordinator backend %d acknowledging SHBUF_EXPAND barrier immediately", MyProcPid);
+ return true;
+ }
+
+ /*
+ * Delay adjusting the new active size of buffer pool till this process
+ * becomes ready to resize buffers.
+ */
+ if (delay_shmem_resize)
+ {
+ elog(LOG, "Phase 3: delaying SHBUF_EXPAND barrier - enabling allocations up to %d buffers, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+ return false;
+ }
+
+ elog(LOG, "Phase 3: Processing SHBUF_EXPAND barrier - enabling allocations up to %d buffers, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * Adjust background writer statistics for the expanded buffer pool
+ */
+ BgBufferSyncReset(NBuffers, targetNBuffers);
+ StrategyReset(targetNBuffers);
+ }
+
+ /* Update local knowledge about the size of active buffer pool. */
+ NBuffers = targetNBuffers;
+
+ /* When expanding, NBuffers should be same as transitNBuffers previous phase. */
+ Assert(NBuffers == transitNBuffers);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemResizeFailed(void)
+{
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /* The work to be done by the coordinator is done in the function which sends the barriers. Hence acknowledge immediately. */
+ if (ShmemCtrl->coordinator == MyProcPid)
+ {
+ elog(LOG, "Coordinator backend %d acknowledging SHBUF_RESIZE_FAILED barrier immediately", MyProcPid);
+ return true;
+ }
+
+ elog(LOG, "received proc signal indicating failure to resize shared buffers from %d to %d, restoring to %d, coordinator is %d",
+ NBuffers, targetNBuffers, currentNBuffers, ShmemCtrl->coordinator);
+
+ /* Restore NBuffers to the original value */
+ NBuffers = currentNBuffers;
+
+ return true;
+}
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index ee9f308379c..b43f1408855 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4129,6 +4129,9 @@ PostgresSingleUserMain(int argc, char *argv[],
/* Initialize size of fast-path lock cache. */
InitializeFastPathLocks();
+ /* Initialize MaxNBuffers for buffer pool resizing. */
+ InitializeMaxNBuffers();
+
/*
* Give preloaded libraries a chance to request additional shared memory.
*/
@@ -4319,14 +4322,12 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
- /* Verify the shared barrier, if it's still active: join and wait. */
- WaitOnShmemBarrier();
-
/*
- * After waiting on the barrier above we guaranteed to have NSharedBuffers
- * broadcasted, so we can use it in the function below.
+ * TODO: The new backend should fetch the shared buffers status. If the
+ * resizing is going on, it should bring itself upto speed with it. If not,
+ * simply fetch the latest pointers are sizes. Is this the right place to do
+ * that?
*/
- AdjustShmemSize();
/*
* Also set up handler to log session end; we have to wait till now to be
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 9a6a6275305..5794d9522d7 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -155,14 +155,12 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so
REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped."
RESTORE_COMMAND "Waiting for <xref linkend="guc-restore-command"/> to complete."
SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a <literal>READ ONLY DEFERRABLE</literal> transaction."
-SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory."
-SHMEM_RESIZE_EVICT "Waiting for other backends to finish buffer evication phase."
-SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory."
SYNC_REP "Waiting for confirmation from a remote server during synchronous replication."
WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
WAL_RECEIVER_WAIT_START "Waiting for startup process to send initial data for streaming replication."
WAL_SUMMARY_READY "Waiting for a new WAL summary to be generated."
XACT_GROUP_UPDATE "Waiting for the group leader to update transaction status at transaction end."
+PM_BUFFER_RESIZE_WAIT "Waiting for the postmaster to complete shared buffer pool resize operations."
ABI_compatibility:
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index 90d3feb547c..894a04caf0f 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -139,8 +139,10 @@ int max_parallel_maintenance_workers = 2;
* MaxBackends is computed by PostmasterMain after modules have had a chance to
* register background workers.
*/
-int NBuffers = 16384;
-int MaxAvailableMemory = 524288;
+int NBuffers = 0;
+int NBuffersPending = 16384;
+bool finalMaxNBuffers = false;
+int MaxNBuffers = 0;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 641e535a73c..e0401fb6477 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -599,6 +599,38 @@ InitializeFastPathLocks(void)
pg_nextpower2_32(FastPathLockGroupsPerBackend));
}
+/*
+ * Initialize MaxNBuffers variable with validation.
+ *
+ * This must be called after GUCs have been loaded but before shared memory size
+ * is determined.
+ *
+ * Since MaxNBuffers limits the size of the buffer pool, it must be at least as
+ * much as NBuffersPending. If MaxNBuffers is 0 (default), set it to
+ * NBuffersPending. Otherwise, validate that MaxNBuffers is not less than
+ * NBuffersPending.
+ */
+void
+InitializeMaxNBuffers(void)
+{
+ if (MaxNBuffers == 0) /* default/boot value */
+ MaxNBuffers = NBuffersPending;
+ else
+ {
+ if (MaxNBuffers < NBuffersPending)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("max_shared_buffers (%d) cannot be less than current shared_buffers (%d)",
+ MaxNBuffers, NBuffersPending),
+ errhint("Increase max_shared_buffers or decrease shared_buffers.")));
+ }
+ }
+
+ Assert(!finalMaxNBuffers);
+ finalMaxNBuffers = true;
+}
+
/*
* Early initialization of a backend (either standalone or under postmaster).
* This happens even before InitPostgres.
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 8794e26ef1d..71a09a65182 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -2731,7 +2731,7 @@ convert_to_base_unit(double value, const char *unit,
* the value without loss. For example, if the base unit is GUC_UNIT_KB, 1024
* is converted to 1 MB, but 1025 is represented as 1025 kB.
*/
-static void
+void
convert_int_from_base_unit(int64 base_value, int base_unit,
int64 *value, const char **unit)
{
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index 7b3ac5f3716..262f42c06c3 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -1108,25 +1108,23 @@
{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
- variable => 'NBuffers',
+ variable => 'NBuffersPending',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
- assign_hook => 'assign_shared_buffers'
+ check_hook => 'check_shared_buffers',
+ show_hook => 'show_shared_buffers',
},
-# TODO: should this be PGC_POSTMASTER?
-{ name => "max_available_memory", type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
+{ name => "max_shared_buffers", type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
short_desc => 'Sets the upper limit for the shared_buffers value.',
- long_desc => 'Shared memory could be resized at runtime, this parameters sets the upper limit for it, beyond which resizing would not be supported. Normally this value would be the same as the total available memory.',
flags => 'GUC_UNIT_BLOCKS',
- variable => 'MaxAvailableMemory',
- boot_val => '524288',
- min => '16',
+ variable => 'MaxNBuffers',
+ boot_val => '0',
+ min => '0',
max => 'INT_MAX / 2',
},
-
{ name => 'vacuum_buffer_usage_limit', type => 'int', context => 'PGC_USERSET', group => 'RESOURCES_MEM',
short_desc => 'Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum.',
flags => 'GUC_UNIT_KB',
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 8f1d0b7c031..ce5110b8636 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12620,4 +12620,10 @@
proargnames => '{pid,io_id,io_generation,state,operation,off,length,target,handle_data_len,raw_result,result,target_desc,f_sync,f_localmem,f_buffered}',
prosrc => 'pg_get_aios' },
+{ oid => '9999', descr => 'resize shared buffers according to the value of GUC `shared_buffers`',
+ proname => 'pg_resize_shared_buffers',
+ provolatile => 'v',
+ prorettype => 'bool',
+ proargtypes => '',
+ prosrc => 'pg_resize_shared_buffers'},
]
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index a0c37a7749e..efe3d3c73ff 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -172,8 +172,11 @@ extern PGDLLIMPORT bool ExitOnAnyError;
extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
+/* TODO: This is no more a GUC variable; should be moved somewhere else. */
extern PGDLLIMPORT int NBuffers;
-extern PGDLLIMPORT int MaxAvailableMemory;
+extern PGDLLIMPORT int NBuffersPending;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
@@ -502,6 +505,7 @@ extern PGDLLIMPORT ProcessingMode Mode;
extern void pg_split_opts(char **argv, int *argcp, const char *optstr);
extern void InitializeMaxBackends(void);
extern void InitializeFastPathLocks(void);
+extern void InitializeMaxNBuffers(void);
extern void InitPostgres(const char *in_dbname, Oid dboid,
const char *username, Oid useroid,
bits32 flags,
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 20bea8132fd..bbb7a225216 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -447,7 +447,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
-extern void StrategyReInitialize(int FirstBufferToInit);
+extern void StrategyReset(int activeNBuffers);
/* buf_table.c */
extern Size BufTableShmemSize(int size);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 74e226269af..6866d09dc22 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -20,6 +20,7 @@
#include "storage/buf.h"
#include "storage/bufpage.h"
#include "storage/relfilelocator.h"
+#include "utils/guc.h"
#include "utils/relcache.h"
#include "utils/snapmgr.h"
@@ -151,6 +152,7 @@ typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersPending;
/* in bufmgr.c */
extern PGDLLIMPORT bool zero_damaged_pages;
@@ -197,6 +199,11 @@ extern PGDLLIMPORT int32 *LocalRefCount;
#define BUFFER_LOCK_SHARE 1
#define BUFFER_LOCK_EXCLUSIVE 2
+/*
+ * prototypes for functions in buf_init.c
+ */
+extern const char *show_shared_buffers(void);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
/*
* prototypes for functions in bufmgr.c
@@ -300,7 +307,7 @@ extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
extern bool BgBufferSync(WritebackContext *wb_context);
-extern void BgBufferSyncReset(int NBuffersOld, int NBuffersNew);
+extern void BgBufferSyncReset(int currentNBuffers, int targetNBuffers);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
@@ -317,11 +324,13 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_evicted,
int32 *buffers_flushed,
int32 *buffers_skipped);
-extern bool EvictExtraBuffers(int fromBuf, int toBuf);
+extern bool EvictExtraBuffers(int targetNBuffers, int currentNBuffers);
/* in buf_init.c */
-extern void BufferManagerShmemInit(int);
-extern Size BufferManagerShmemSize(void);
+extern void BufferManagerShmemInit(void);
+extern Size BufferManagerShmemSize(bool set_reserved);
+extern void BufferManagerShmemResize(int currentNBuffers, int targetNBuffers);
+extern void BufferManagerShmemValidate(int targetNBuffers);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 6e7b0abb625..10e74b34813 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -64,7 +64,6 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
-extern PGDLLIMPORT volatile bool pending_pm_shmem_resize;
extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 704b065f9e9..34b5e6c48ca 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,9 +24,11 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "port/atomics.h"
#include "storage/barrier.h"
#include "storage/dsm_impl.h"
#include "storage/spin.h"
+#include "utils/guc.h"
typedef struct AnonymousMapping
{
@@ -73,14 +75,20 @@ extern PGDLLIMPORT AnonymousMapping Mappings[ANON_MAPPINGS];
/*
* ShmemControl is shared between backends and helps to coordinate shared
* memory resize.
+ *
+ * TODO: I think we need a lock to protect this structure. If we do so, do we
+ * need to use atomic integers?
*/
typedef struct
{
- pg_atomic_uint32 NSharedBuffers;
- pid_t evictor_pid;
- Barrier Barrier;
- pg_atomic_uint64 Generation;
- bool Resizable;
+ pg_atomic_flag resize_in_progress; /* true if resizing is in progress. false otherwise. */
+ pg_atomic_uint32 currentNBuffers; /* Original NBuffers value before resize started */
+ pg_atomic_uint32 targetNBuffers;
+ pg_atomic_uint32 activeNBuffers; /* Active portion of buffer pool during resizing. */
+ pg_atomic_uint32 transitNBuffers; /* Part of the buffer pool beyond activeNBuffers which may remain accessible during resizing. */
+ pid_t coordinator;
+ ConditionVariable pm_cv; /* Coordinator waits for PM to complete its work using this CV. */
+ bool pmwork_done; /* PM has completed its work of resizing buffers. */
} ShmemControl;
extern PGDLLIMPORT ShmemControl *ShmemCtrl;
@@ -95,7 +103,8 @@ extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
-extern PGDLLIMPORT int MaxAvailableMemory;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -145,7 +154,8 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
void PrepareHugePages(void);
bool ProcessBarrierShmemResize(Barrier *barrier);
-void assign_shared_buffers(int newval, void *extra, bool *pending);
+const char *show_shared_buffers(void);
+bool check_shared_buffers(int *newval, void **extra, GucSource source);
void AdjustShmemSize(void);
extern void WaitOnShmemBarrier(void);
extern void ShmemControlInit(void);
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index 97033f84dce..b80b05f2804 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,7 +54,10 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
- PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */
+ PROCSIGNAL_BARRIER_SHBUF_SHRINK, /* shrink buffer pool - restrict allocations to new size */
+ PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM, /* remap shared memory segments and update structure pointers */
+ PROCSIGNAL_BARRIER_SHBUF_EXPAND, /* expand buffer pool - enable allocations in new range */
+ PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED, /* signal backends that the shared buffer resizing failed. */
} ProcSignalBarrierType;
/*
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 64ff5a286ba..6944560d485 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -50,11 +50,19 @@ extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
extern void *ShmemInitStructInSegment(const char *name, Size size,
bool *foundPtr, int shmem_segment);
+extern void *ShmemUpdateStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
extern PGDLLIMPORT Size pg_get_shmem_pagesize(void);
+extern bool ProcessBarrierShmemShrink(void);
+extern bool ProcessBarrierShmemResizeMapAndMem(void);
+extern bool ProcessBarrierShmemExpand(void);
+extern bool ProcessBarrierShmemResizeFailed(void);
+
+
/* ipci.c */
extern void RequestAddinShmemSpace(Size size);
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f21ec37da89..08a84373fb7 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -459,6 +459,8 @@ extern config_handle *get_config_handle(const char *name);
extern void AlterSystemSetConfigFile(AlterSystemStmt *altersysstmt);
extern char *GetConfigOptionByName(const char *name, const char **varname,
bool missing_ok);
+extern void convert_int_from_base_unit(int64 base_value, int base_unit,
+ int64 *value, const char **unit);
extern void TransformGUCArray(ArrayType *array, List **names,
List **values);
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
index 97c3da9e20a..eb275027fa6 100644
--- a/src/test/buffermgr/Makefile
+++ b/src/test/buffermgr/Makefile
@@ -13,6 +13,9 @@ EXTRA_INSTALL = contrib/pg_buffercache
REGRESS = buffer_resize
+# Custom configuration for buffer manager tests
+TEMP_CONFIG = $(srcdir)/buffermgr_test.conf
+
subdir = src/test/buffermgr
top_builddir = ../../..
include $(top_builddir)/src/Makefile.global
diff --git a/src/test/buffermgr/buffermgr_test.conf b/src/test/buffermgr/buffermgr_test.conf
new file mode 100644
index 00000000000..21ccf66d9c7
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.conf
@@ -0,0 +1,9 @@
+# Configuration for buffer manager regression tests
+
+# Even if max_shared_buffers is set multiple times only the last one is used to
+# as the limit on shared_buffers.
+max_shared_buffers = 128kB
+# Set initial shared_buffers as expected by test
+shared_buffers = 128MB
+# Set a larger value for max_shared_buffers to allow testing resize operations
+max_shared_buffers = 300MB
\ No newline at end of file
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
index a986be9a5da..d5cb9d78437 100644
--- a/src/test/buffermgr/expected/buffer_resize.out
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -1,9 +1,8 @@
-- Test buffer pool resizing and shared memory allocation tracking
-- This test resizes the buffer pool multiple times and monitors
-- shared memory allocations related to buffer management
--- Create a separate schema for this test
-CREATE SCHEMA buffer_resize_test;
-SET search_path TO buffer_resize_test, public;
+-- TODO: The test sets shared_buffers values in MBs. Instead it could use values
+-- in kBs so that the test runs on very small machines.
-- Create a view for buffer-related shared memory allocations
CREATE VIEW buffer_allocations AS
SELECT name, segment, size, allocated_size
@@ -28,6 +27,49 @@ SHOW shared_buffers;
128MB
(1 row)
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | descriptors | 1048576 | 1048576
+ Buffer IO Condition Variables | iocv | 262144 | 262144
+ Checkpoint BufferIds | checkpoint | 327680 | 327680
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 134225920 | 134225920 | 314580992
+ checkpoint | 335872 | 335872 | 770048
+ descriptors | 1056768 | 1056768 | 2465792
+ iocv | 270336 | 270336 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
SELECT * FROM buffer_allocations;
name | segment | size | allocated_size
-------------------------------+-------------+-----------+----------------
@@ -40,10 +82,10 @@ SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
name | size | mapping_size | mapping_reserved_size
-------------+-----------+--------------+-----------------------
- buffers | 134225920 | 134225920 | 2576982016
- checkpoint | 335872 | 335872 | 214753280
- descriptors | 1056768 | 1056768 | 429498368
- iocv | 270336 | 270336 | 429498368
+ buffers | 134225920 | 134225920 | 314580992
+ checkpoint | 335872 | 335872 | 770048
+ descriptors | 1056768 | 1056768 | 2465792
+ iocv | 270336 | 270336 | 622592
(4 rows)
SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
@@ -60,10 +102,18 @@ SELECT pg_reload_conf();
t
(1 row)
-SELECT pg_sleep(1);
- pg_sleep
-----------
-
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 128MB (pending: 64MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
(1 row)
SHOW shared_buffers;
@@ -84,10 +134,10 @@ SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
name | size | mapping_size | mapping_reserved_size
-------------+----------+--------------+-----------------------
- buffers | 67117056 | 67117056 | 2576982016
- checkpoint | 172032 | 172032 | 214753280
- descriptors | 532480 | 532480 | 429498368
- iocv | 139264 | 139264 | 429498368
+ buffers | 67117056 | 67117056 | 314580992
+ checkpoint | 172032 | 172032 | 770048
+ descriptors | 532480 | 532480 | 2465792
+ iocv | 139264 | 139264 | 622592
(4 rows)
SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
@@ -104,10 +154,18 @@ SELECT pg_reload_conf();
t
(1 row)
-SELECT pg_sleep(1);
- pg_sleep
-----------
-
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 64MB (pending: 256MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
(1 row)
SHOW shared_buffers;
@@ -128,10 +186,10 @@ SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
name | size | mapping_size | mapping_reserved_size
-------------+-----------+--------------+-----------------------
- buffers | 268443648 | 268443648 | 2576982016
- checkpoint | 663552 | 663552 | 214753280
- descriptors | 2105344 | 2105344 | 429498368
- iocv | 532480 | 532480 | 429498368
+ buffers | 268443648 | 268443648 | 314580992
+ checkpoint | 663552 | 663552 | 770048
+ descriptors | 2105344 | 2105344 | 2465792
+ iocv | 532480 | 532480 | 622592
(4 rows)
SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
@@ -148,10 +206,18 @@ SELECT pg_reload_conf();
t
(1 row)
-SELECT pg_sleep(1);
- pg_sleep
-----------
-
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 256MB (pending: 100MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
(1 row)
SHOW shared_buffers;
@@ -172,10 +238,10 @@ SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
name | size | mapping_size | mapping_reserved_size
-------------+-----------+--------------+-----------------------
- buffers | 104865792 | 104865792 | 2576982016
- checkpoint | 262144 | 262144 | 214753280
- descriptors | 827392 | 827392 | 429498368
- iocv | 212992 | 212992 | 429498368
+ buffers | 104865792 | 104865792 | 314580992
+ checkpoint | 262144 | 262144 | 770048
+ descriptors | 827392 | 827392 | 2465792
+ iocv | 212992 | 212992 | 622592
(4 rows)
SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
@@ -192,10 +258,18 @@ SELECT pg_reload_conf();
t
(1 row)
-SELECT pg_sleep(1);
- pg_sleep
-----------
-
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 100MB (pending: 128kB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
(1 row)
SHOW shared_buffers;
@@ -216,10 +290,10 @@ SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
name | size | mapping_size | mapping_reserved_size
-------------+--------+--------------+-----------------------
- buffers | 139264 | 139264 | 2576982016
- checkpoint | 8192 | 8192 | 214753280
- descriptors | 8192 | 8192 | 429498368
- iocv | 8192 | 8192 | 429498368
+ buffers | 139264 | 139264 | 314580992
+ checkpoint | 8192 | 8192 | 770048
+ descriptors | 8192 | 8192 | 2465792
+ iocv | 8192 | 8192 | 622592
(4 rows)
SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
@@ -228,10 +302,28 @@ SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
16
(1 row)
--- Clean up the schema and all its objects
-RESET search_path;
-DROP SCHEMA buffer_resize_test CASCADE;
-NOTICE: drop cascades to 3 other objects
-DETAIL: drop cascades to view buffer_resize_test.buffer_allocations
-drop cascades to view buffer_resize_test.buffer_segments
-drop cascades to extension pg_buffercache
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+ERROR: invalid value for parameter "shared_buffers": 51200
+DETAIL: "shared_buffers" must be less than "max_shared_buffers".
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
index e71dcdea685..561630e846f 100644
--- a/src/test/buffermgr/meson.build
+++ b/src/test/buffermgr/meson.build
@@ -8,10 +8,15 @@ tests += {
'sql': [
'buffer_resize',
],
+ 'regress_args': ['--temp-config', files('buffermgr_test.conf')],
},
'tap': {
+ 'env': {
+ 'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
+ },
'tests': [
't/001_resize_buffer.pl',
+ 't/003_parallel_resize_buffer.pl',
],
},
}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
index 45f5bb6d78b..dfaaeabfcbb 100644
--- a/src/test/buffermgr/sql/buffer_resize.sql
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -1,10 +1,8 @@
-- Test buffer pool resizing and shared memory allocation tracking
-- This test resizes the buffer pool multiple times and monitors
-- shared memory allocations related to buffer management
-
--- Create a separate schema for this test
-CREATE SCHEMA buffer_resize_test;
-SET search_path TO buffer_resize_test, public;
+-- TODO: The test sets shared_buffers values in MBs. Instead it could use values
+-- in kBs so that the test runs on very small machines.
-- Create a view for buffer-related shared memory allocations
CREATE VIEW buffer_allocations AS
@@ -28,6 +26,13 @@ CREATE EXTENSION IF NOT EXISTS pg_buffercache;
-- Test 1: Default shared_buffers
SHOW shared_buffers;
+SHOW max_shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
@@ -35,7 +40,10 @@ SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
-- Test 2: Set to 64MB
ALTER SYSTEM SET shared_buffers = '64MB';
SELECT pg_reload_conf();
-SELECT pg_sleep(1);
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
SHOW shared_buffers;
SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
@@ -44,7 +52,10 @@ SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
-- Test 3: Set to 256MB
ALTER SYSTEM SET shared_buffers = '256MB';
SELECT pg_reload_conf();
-SELECT pg_sleep(1);
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
SHOW shared_buffers;
SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
@@ -53,7 +64,10 @@ SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
-- Test 4: Set to 100MB (non-power-of-two)
ALTER SYSTEM SET shared_buffers = '100MB';
SELECT pg_reload_conf();
-SELECT pg_sleep(1);
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
SHOW shared_buffers;
SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
@@ -62,12 +76,20 @@ SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
-- Test 5: Set to minimum 128kB
ALTER SYSTEM SET shared_buffers = '128kB';
SELECT pg_reload_conf();
-SELECT pg_sleep(1);
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
SHOW shared_buffers;
SELECT * FROM buffer_allocations;
SELECT * FROM buffer_segments;
SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
--- Clean up the schema and all its objects
-RESET search_path;
-DROP SCHEMA buffer_resize_test CASCADE;
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+SHOW max_shared_buffers;
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
index 8cf9e4539ab..7b0a78ebe5b 100644
--- a/src/test/buffermgr/t/001_resize_buffer.pl
+++ b/src/test/buffermgr/t/001_resize_buffer.pl
@@ -14,40 +14,26 @@ sub apply_and_verify_buffer_change
{
my ($node, $new_size) = @_;
- # Use a single background_psql session for consistency
- my $psql_session = $node->background_psql('postgres');
- $psql_session->query_safe("ALTER SYSTEM SET shared_buffers = '$new_size'");
- $psql_session->query_safe("SELECT pg_reload_conf()");
-
- # Wait till the resizing finishes using the same session
- #
- # TODO: Right now there is no way to know when the resize has finished and
- # all the backends are using new value of shared_buffers. Hence we poll
- # manually until we get the expected value in the same session.
- my $current_size;
- my $attempts = 0;
- my $max_attempts = 60; # 60 seconds timeout
- do {
- $current_size = $psql_session->query_safe("SHOW shared_buffers");
- $attempts++;
-
- # Only sleep if we didn't get the expected result and haven't timed out yet
- if ($current_size ne $new_size && $attempts < $max_attempts) {
- sleep(1);
- }
- } while ($current_size ne $new_size && $attempts < $max_attempts);
-
- $psql_session->quit;
-
- # Check if we succeeded or timed out
- if ($current_size ne $new_size) {
- die "Timeout waiting for shared_buffers to change to $new_size (got $current_size after ${attempts}s)";
- }
+ # Use the new pg_resize_shared_buffers() interface which handles everything synchronously
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ # Call the resize function - it returns when the operation is complete
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"), 't',
+ 'resizing to ' . $new_size . ' succeeded');
+ is($node->safe_psql('postgres', "SHOW shared_buffers"), $new_size,
+ 'SHOW after resizing to '. $new_size . ' succeeded');
}
# Initialize a cluster and start pgbench in the background for concurrent load.
my $node = PostgreSQL::Test::Cluster->new('main');
$node->init;
+
+# Permit resizing up to 1GB for this test and let the server start with 128MB.
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 1GB
+shared_buffers = 128MB
+});
+
$node->start;
$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
my $pgb_scale = 10;
diff --git a/src/test/buffermgr/t/003_parallel_resize_buffer.pl b/src/test/buffermgr/t/003_parallel_resize_buffer.pl
new file mode 100644
index 00000000000..9cbb5452fd2
--- /dev/null
+++ b/src/test/buffermgr/t/003_parallel_resize_buffer.pl
@@ -0,0 +1,71 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test that only one pg_resize_shared_buffers() call succeeds when multiple
+# sessions attempt to resize buffers concurrently
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Initialize a cluster
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', 'shared_buffers = 128kB');
+$node->append_conf('postgresql.conf', 'max_shared_buffers = 256kB');
+$node->start;
+
+# Load injection points extension for test coordination
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Test 1: Two concurrent pg_resize_shared_buffers() calls
+# Set up injection point to pause the first resize call
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('pg-resize-shared-buffers-flag-set', 'wait')");
+
+# Change shared_buffers for the resize operation
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '144kB'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+# Start first resize session (will pause at injection point)
+my $session1 = $node->background_psql('postgres');
+$session1->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+);
+
+# Wait until session actually reaches the injection point
+$node->wait_for_event('client backend', 'pg-resize-shared-buffers-flag-set');
+
+# Start second resize session (should fail immediately since resize is in progress)
+my $result2 = $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()");
+
+# The second call should return false (already in progress)
+is($result2, 'f', 'Second concurrent resize call returns false');
+
+# Wake up the first session
+$node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('pg-resize-shared-buffers-flag-set')");
+
+# The pg_resize_shared_buffers() in session1 should now complete successfully
+# We can't easily capture the return value from query_until, but we can
+# verify the session completes without error and the resize actually happened
+$session1->quit;
+
+# Detach injection point
+$node->safe_psql('postgres',
+ "SELECT injection_points_detach('pg-resize-shared-buffers-flag-set')");
+
+done_testing();
\ No newline at end of file
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-10-14 08:35 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-10-16 16:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-10-14 08:35 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
As I've mentioned in our off list communication, I'm working on the new
design and was planning to post some intermediate results in a couple of
weeks. Thus I'm surprised that instead of aligning on plans you've
decided to post you own version earlier. It most certainly doesn't make
things easier for me, so what's your plan anyway? Are you trying to
hijack the thread with your own patches? It doesn't strike me as
particularly constructive thing to do.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-14 08:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-10-16 16:25 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-17 10:33 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-10-16 16:25 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
Hi Dmitry,
On Tue, Oct 14, 2025 at 2:05 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> As I've mentioned in our off list communication, I'm working on the new
> design and was planning to post some intermediate results in a couple of
> weeks. Thus I'm surprised that instead of aligning on plans you've
> decided to post you own version earlier. It most certainly doesn't make
> things easier for me, so what's your plan anyway? Are you trying to
> hijack the thread with your own patches? It doesn't strike me as
> particularly constructive thing to do.
I am sorry if you have felt that way. That wasn't the intention.
Please allow me to explain ...
Tomas and Andres have pointed out some serious faults in the patchset
posted at [2], including problems in my patches. There are many of
those. If we work in parallel we can make good progress. So I
continued working based on your last patchset [2], which was posted
more than 3 months ago. Knowing that you are working on a design, I
tried not to touch the synchronization and UI part and yet find
solutions to some of the open problems (my patchset in [3] is a recent
example). As I mentioned in my email at [1], every open question I
tried to solve next was blocked because of a single problem, which I
have described in my previous email - A problem in synchronization in
the patchset at [2].
Instead of just doing nothing, I thought I would try to implement the
UI and synchronization that I had in mind. Once I implemented it and
saw that it could address a few serious concerns raised by Andres, I
thought I would share it with hackers to get some early feedback.
Early feedback from people like Andres and Tomas is important to avoid
going down the wrong path (and wasting time). Is there something wrong
with that? BTW, this idea isn't new and it's certainly not only mine.
It's a combination of an implementation shared by Thomas Munro [4] and
an implementation I had shared with you offlist on 30th January 2025.
I never saw any comments from you on the specific changes in those
implementations and neither anything from those patchsets was absorbed
in your patchsets.
If I would have posted my alternate solution in January itself, that
might have been considered hijacking (that's a serious accusation,
btw). But instead I worked with your patches, improving them as long
as I could. Even the patchset I shared is still
on top of your patchset in [2].
I don't know your solution. But if it's similar to my proposal, we are
in agreement and can work further in parallel on subproblems. If it's
different, let's discuss pros and cons of both - maybe there is some
value in letting those evolve parallely and let the community choose
the best, or choose best of both solutions giving rise to a new
solution. My patchset might give you solutions/code for the problems
you are trying to solve. It has tests which you can adapt to your
solution. Many exciting possibilities lie ahead with multiple working
solutions. Knowing nothing about the solution you are attempting, it's
hard to know which of these apply and help you.
[1] https://www.postgresql.org/message-id/CAExHW5sOu8%2B9h6t7jsA5jVcQ--N-LCtjkPnCw%2BrpoN0ovT6PHg%40mail...
[2] https://www.postgresql.org/message-id/my4hukmejato53ef465ev7lk3sqiqvneh7436rz64wmtc7rbfj%40hmuxsf2ng...
[3] https://www.postgresql.org/message-id/CAExHW5vB8sAmDtkEN5dcYYeBok3D8eAzMFCOH1k%2Bkrxht1yFjA%40mail.g...
[4] https://www.postgresql.org/message-id/CA%2BhUKGL5hW3i_pk5y_gcbF_C5kP-pWFjCuM8bAyCeHo3xUaH8g%40mail.g...
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-14 08:35 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-10-16 16:25 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-10-17 10:33 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-10-17 10:33 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
> On Thu, Oct 16, 2025 at 09:55:05PM +0530, Ashutosh Bapat wrote:
>
> BTW, this idea isn't new and it's certainly not only mine.
> It's a combination of an implementation shared by Thomas Munro [4] and
> an implementation I had shared with you offlist on 30th January 2025.
> I never saw any comments from you on the specific changes in those
> implementations and neither anything from those patchsets was absorbed
> in your patchsets.
Well, this is imply not true. We had an extensive discussion for long
time off-list and even a few video calls to talk through various design
options and agree about next steps.
> I don't know your solution. But if it's similar to my proposal, we are
> in agreement and can work further in parallel on subproblems. If it's
> different, let's discuss pros and cons of both - maybe there is some
> value in letting those evolve parallely and let the community choose
> the best, or choose best of both solutions giving rise to a new
> solution. My patchset might give you solutions/code for the problems
> you are trying to solve. It has tests which you can adapt to your
> solution. Many exciting possibilities lie ahead with multiple working
> solutions. Knowing nothing about the solution you are attempting, it's
> hard to know which of these apply and help you.
I've shared many times on- and off-list the general directions I'm
working in and even the expected timeline, so it's strange to state you
don't know it.
In the end you're free to do whatever you want, fortunately it's open
source. But posting an alternative patch series and "let the community
choose" does sound like hijacking to me, and a direct way to split and
reduce already scarse review attention.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-11-14 11:53 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-11-14 11:53 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
Hi,
PFA new patchset with some TODOs from previous email addressed:
On Mon, Oct 13, 2025 at 9:28 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
> 1. New backends join while the synchronization is going on.
Done. Explained the solution below in detail.
> An existing backend exiting.
Not tested specifically, but should work.
> 2. Failure or crash in the backend which is executing pg_resize_buffer_pool()
still a TODO
> 3. Fix crashes in the tests.
core regression passes, pg_buffercache regression tests pass and the
tests for buffer resizing pass most of the time. So far I have seen
two issues
1. An assertion from AIO worker - which happened only once and I
couldn't reproduce again. Need to study interaction of AIO worker with
buffer resizing.
2. checkpointer crashes - which is one of the TODOs listed below.
3. Also there's an shared memory id related failure, which I don't
understand but happen more frequently than the first one. Need to look
into that.
> go through Tomas's detailed comments and address those
> which still apply.
Still a TODO. But since many of those patches are revised heavily, I
think many of the comments may have been addressed, some may not apply
anymore.
> And the patches are still WIP, with many TODOs. But I wanted to get some feedback on the proposed UI and synchronization
This is still a request.
> Patches 0001 to 0016 are the same as the previous patchset. I haven't
> touched them in case someone would like to see an incremental change.
> However, it's getting unwieldy at this point, so I will squash
> relevant patches together and provide a patchset with fewer patches
> next.
I have squashed the patches into 3 so that it's easy to review, read
and work with those patches. The work is still WIP and there are many
TODOs in the patches.
Patch 0001: SQL interface to read contents of buffer lookup table. It
was there in the previous patchset as 0001 but in this patchset I have
moved the SQL function to the pg_buffercache module and renamed it
accordingly. I added this change because I found it useful to debug
issues I found while testing buffer resizing patches. The issues were
related to page->buffer mappings which existed in the buffer look up
table but were not present in the buffer descriptor array or buffer
blocks. pg_buffercache, which traverses just the buffer descriptor
array, isn't enough. Even without the resizing functionality this will
help us catch situations where buffer descriptor array and buffer
lookup table goes out of sync. I plan to keep it in this patchset as a
debugging tool. If other developers feel that it could be useful, I
will propose it in a separate thread.
Patch 0002: This is a single patch squashing all patches (0005, 0006,
0007, 0008, 0009 and 0010) related to shared memory management and
address space reservation together. This patch allows the creation of
multiple shared memory segments and also lays them out so as to make
those resizable. The actual code to resize the segments is in the next
patch. The APIs used for memory management and address space
reservation are described later. Prominent changes from the previous
patches are:
1. modifies CalculateShmemSize() so that it can work with multiple
shared memory segments.
2. It also combines AnonymousMapping and ShmemSegment structures
together as suggested by Tomas upthread. The merger is still going on.
There are some old comments or variable names referring to memory
mapping when they should be mentioning shared memory segments. I will
work on that when I start polishing this patch.
4. GUC to specify the maximum size of buffer pool has been renamed and
moved to the next patch which deals with actual resizing.
5. Changes to process config reload in AIO workers are removed. Those
are not needed after 55b454d0e14084c841a034073abbf1a0ea937a45.
Patch 0003: Implements the UI and synchronization described in the
previous email [1] with additional improvements to support a new
backend joining while resizing is in progress. This patch squashes
other patches 0002 - 0004 and 0011 onward patches from the previous
patchset, but it also gets rid of a lot of code related to the old
synchronization method and the old UI. The code related to resizing
including implementation of pg_resize_shared_buffers() is moved to
storage/buffer/buf_resize.c, a new file. There is no change to the UI.
The buffer resizing still looks like as described in the previous
email.
> SHOW shared_buffers; -- default
> shared_buffers
> ----------------
> 128MB
> (1 row)
>
> ALTER SYSTEM SET shared_buffers = '64MB';
> SELECT pg_reload_conf();
> pg_reload_conf
> ----------------
> t
> (1 row)
>
> SHOW shared_buffers;
> shared_buffers
> -----------------------
> 128MB (pending: 64MB)
> (1 row)
>
> SELECT pg_resize_shared_buffers();
> pg_resize_shared_buffers
> --------------------------
> t
> (1 row)
>
> SHOW shared_buffers;
> shared_buffers
> ----------------
> 64MB
> (1 row)
>
> ALTER SYSTEM SET shared_buffers = '256MB';
> SELECT pg_reload_conf();
> pg_reload_conf
> ----------------
> t
> (1 row)
>
> SHOW shared_buffers;
> shared_buffers
> -----------------------
> 64MB (pending: 256MB)
> (1 row)
>
> SELECT pg_resize_shared_buffers();
> pg_resize_shared_buffers
> --------------------------
> t
> (1 row)
>
> SHOW shared_buffers;
> shared_buffers
> ----------------
> 256MB
> (1 row)
>
The implementation uses a similar strategy as described in the
previous email with changes described below.
A new backend inherits the address space of shared memory segments and
the local variable NBuffers through Postmaster. These are changed when
resizing the buffer pool. And the same changes need to be applied to
the Postmaster so that a new backend inherits them. Since Postmaster
is not part of the ProcSignalBarrier mechanism, the coordinator has to
send signals to the Postmaster separately. This has the following
drawbacks
1. Additional code to signal Postmaster
2. coordinator has to wait for Postmaster to apply the changes
separately, thus adding extra delays
3. platforms which use fork() + exec(), will add more complexity to
transfer the state to new child
4. If the postmaster is signaled after sending a barrier to other
backends, the newly joined backend will miss the state update as well
as the barrier. If the postmaster is signaled before sending a barrier
to other backends, a newly joining backend will receive the barrier as
well as state update from Postmaster. This means the barrier handling
code is required to be idempotent. This will make the barrier handling
code more complex and also constrained.
Instead the approach taken by Thomas Munro in [2] does not require
updating the address space. It uses shared memory variables instead of
process local memory variables to save the state of the shared buffer
pool. This patchset uses a similar approach and
1. avoids involving Postmaster in the resizing process
2. additionally making barrier handling code super thin.
Shared Memory and address space management
========================================
An fd is created using memfd_create to manage the size of the shared
memory segment using ftruncate and fallocate(). That fd is passed to
mmap() which reserves the maximum required address space and maps the
anonymous file (and the backing memory) in that address space. mmap
uses MAP_NORESERVE so as not to allocate memory against mapping. The
size of the anonymous file controls the amount of memory allocated.
For the main shared memory segment, the size of the reserved space is
the same as the amount of memory required. But for shared buffer pool
related segments the size of the reserved space is decided by GUC
max_shared_buffers (mentioned in the previous email and quoted below).
When resizing shared buffers only the anonymous file is resized and
not the address space. I tested this protocol with an attached small
program (mfdtruncate.c). Sharing it in case somebody finds it useful.
Saving shared buffer pool sizes in the shared memory
=========================================
When resizing, we need to track two ranges of buffers 1. active
buffers, which is the range of buffers from which the new allocations
happen at a given time and 2. valid buffers which is the range of
buffers which are valid at a given time. When shrinking, the active
buffers is set to the new size while the valid buffers remains same as
the old size till all the buffers outside the new size are evicted.
When expanding, valid buffers and active buffers are both changed to
new size after memory is resized and expanded data structures are
initialized. Current global variable NBuffers is insufficient to track
these two numbers.
Instead we have a new member StrategyControl::activeNBuffers which
tracks the active buffer range. The shared memory structure
controlling the resizing operation (ShmemCtrl) has a member
currentNBuffers which gives the range of valid number of shared
buffers at a given point in time. (I am planning to merge ShmemCtrl
and StrategyControl, so that we have all the metadata about shared
buffers in one place in the shared memory). These two numbers are
saved in the shared memory for the reasons explained below and replace
current NBuffers. They are modified by the coordinator as the resizing
progresses. Some usages of NBuffers are replaced by one of the two
variables as appropriate but more work is required.
Next I will be working on
1. Background writer synchronization
2. Checkpoint synchronization
3. Make all the shared buffer pool structures, except buffer blocks,
static and maximally allocated as suggested by Andres earlier. [3]
4. Replace NBuffers usages as explained above
3. merge ShmemCtrl and StrategyControl as explained above
4. Handle failures in resizing
5. There have been concerns raised earlier that anonymous file backed
memory is not dumped with core. I am thinking of not using an
anonymous file for the main memory segment so that it gets dumped with
core. But shared buffers still will be dumped. However, I am skeptical
as to whether we need GBs (say) of shared buffers being dumped along
with core or should we leave that choice to users.
[1] https://www.postgresql.org/message-id/CAExHW5sOu8+9h6t7jsA5jVcQ--N-LCtjkPnCw+rpoN0ovT6PHg@mail.gmail...
[2] https://www.postgresql.org/message-id/CA%2BhUKGL5hW3i_pk5y_gcbF_C5kP-pWFjCuM8bAyCeHo3xUaH8g%40mail.g...
[3] https://www.postgresql.org/message-id/qltuzcdxapofdtb5mrd4em3bzu2qiwhp3cdwdsosmn7rhrtn4u%40yaogvphfw...
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] 0001-Add-a-view-to-read-contents-of-shared-buffe-20251114.patch (14.7K, ../../CAExHW5sVxEwQsuzkgjjJQP9-XVe0H2njEVw1HxeYFdT7u7J+eQ@mail.gmail.com/2-0001-Add-a-view-to-read-contents-of-shared-buffe-20251114.patch)
download | inline diff:
From a24f17114aa9119dbf899166128d48aaf4106ca7 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 25 Aug 2025 19:23:50 +0530
Subject: [PATCH 1/4] Add a view to read contents of shared buffer lookup table
The view exposes the contents of the shared buffer lookup table for
debugging, testing and investigation.
This helped me in debugging issues where the buffer descriptor array and
buffer lookup table were out of sync; either the buffer lookup table had
a mapping page->buffer which wasn't present in the buffer descriptor
array or a page in the buffer descriptor array didn't have corresponding
entry in the buffer lookup table. pg_buffercache doesn't help with those
kind of issues. Also doing that under the debugger in very painful.
I intend to keep this patch while the rest of the code matures. If it is
found useful as a debugging tool, we may consider make it committable
and commit it.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../expected/pg_buffercache.out | 39 ++++++++
.../pg_buffercache--1.5--1.6.sql | 24 +++++
contrib/pg_buffercache/pg_buffercache_pages.c | 18 ++++
contrib/pg_buffercache/sql/pg_buffercache.sql | 20 +++++
doc/src/sgml/system-views.sgml | 89 +++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 58 ++++++++++++
src/include/storage/buf_internals.h | 2 +
7 files changed, 250 insertions(+)
diff --git a/contrib/pg_buffercache/expected/pg_buffercache.out b/contrib/pg_buffercache/expected/pg_buffercache.out
index 9a9216dc7b1..2f27bf34637 100644
--- a/contrib/pg_buffercache/expected/pg_buffercache.out
+++ b/contrib/pg_buffercache/expected/pg_buffercache.out
@@ -23,6 +23,26 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
t
(1 row)
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -34,6 +54,10 @@ SELECT * FROM pg_buffercache_summary();
ERROR: permission denied for function pg_buffercache_summary
SELECT * FROM pg_buffercache_usage_counts();
ERROR: permission denied for function pg_buffercache_usage_counts
+SELECT * FROM pg_buffercache_lookup_table_entries();
+ERROR: permission denied for function pg_buffercache_lookup_table_entries
+SELECT * FROM pg_buffercache_lookup_table;
+ERROR: permission denied for view pg_buffercache_lookup_table
RESET role;
-- Check that pg_monitor is allowed to query view / function
SET ROLE pg_monitor;
@@ -55,6 +79,21 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts();
t
(1 row)
+RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
RESET role;
------
---- Test pg_buffercache_evict* functions
diff --git a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
index 458f054a691..9bf58567878 100644
--- a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
+++ b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
@@ -44,3 +44,27 @@ CREATE FUNCTION pg_buffercache_evict_all(
OUT buffers_skipped int4)
AS 'MODULE_PATHNAME', 'pg_buffercache_evict_all'
LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Add the buffer lookup table function
+CREATE FUNCTION pg_buffercache_lookup_table_entries(
+ OUT tablespace oid,
+ OUT database oid,
+ OUT relfilenode oid,
+ OUT forknum int2,
+ OUT blocknum int8,
+ OUT bufferid int4)
+RETURNS SETOF record
+AS 'MODULE_PATHNAME', 'pg_buffercache_lookup_table_entries'
+LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Create a view for convenient access.
+CREATE VIEW pg_buffercache_lookup_table AS
+ SELECT * FROM pg_buffercache_lookup_table_entries();
+
+-- Don't want these to be available to public.
+REVOKE ALL ON FUNCTION pg_buffercache_lookup_table_entries() FROM PUBLIC;
+REVOKE ALL ON pg_buffercache_lookup_table FROM PUBLIC;
+
+-- Grant access to monitoring role.
+GRANT EXECUTE ON FUNCTION pg_buffercache_lookup_table_entries() TO pg_read_all_stats;
+GRANT SELECT ON pg_buffercache_lookup_table TO pg_read_all_stats;
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index ab790533ff6..fe9af45febe 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -16,6 +16,7 @@
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "utils/rel.h"
+#include "utils/tuplestore.h"
#define NUM_BUFFERCACHE_PAGES_MIN_ELEM 8
@@ -100,6 +101,7 @@ PG_FUNCTION_INFO_V1(pg_buffercache_usage_counts);
PG_FUNCTION_INFO_V1(pg_buffercache_evict);
PG_FUNCTION_INFO_V1(pg_buffercache_evict_relation);
PG_FUNCTION_INFO_V1(pg_buffercache_evict_all);
+PG_FUNCTION_INFO_V1(pg_buffercache_lookup_table_entries);
/* Only need to touch memory once per backend process lifetime */
@@ -776,3 +778,19 @@ pg_buffercache_evict_all(PG_FUNCTION_ARGS)
PG_RETURN_DATUM(result);
}
+
+/*
+ * Return lookup table content as a set of records.
+ */
+Datum
+pg_buffercache_lookup_table_entries(PG_FUNCTION_ARGS)
+{
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* Fill the tuplestore */
+ BufTableGetContents(rsinfo->setResult, rsinfo->setDesc);
+
+ return (Datum) 0;
+}
diff --git a/contrib/pg_buffercache/sql/pg_buffercache.sql b/contrib/pg_buffercache/sql/pg_buffercache.sql
index 47cca1907c7..569b28aebb9 100644
--- a/contrib/pg_buffercache/sql/pg_buffercache.sql
+++ b/contrib/pg_buffercache/sql/pg_buffercache.sql
@@ -12,6 +12,18 @@ from pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -19,6 +31,8 @@ SELECT * FROM pg_buffercache;
SELECT * FROM pg_buffercache_pages() AS p (wrong int);
SELECT * FROM pg_buffercache_summary();
SELECT * FROM pg_buffercache_usage_counts();
+SELECT * FROM pg_buffercache_lookup_table_entries();
+SELECT * FROM pg_buffercache_lookup_table;
RESET role;
-- Check that pg_monitor is allowed to query view / function
@@ -28,6 +42,12 @@ SELECT buffers_used + buffers_unused > 0 FROM pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts();
RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+RESET role;
+
------
---- Test pg_buffercache_evict* functions
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 7971498fe75..8f3e2741051 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -71,6 +71,11 @@
<entry>backend memory contexts</entry>
</row>
+ <row>
+ <entry><link linkend="view-pg-buffer-lookup-table"><structname>pg_buffer_lookup_table</structname></link></entry>
+ <entry>shared buffer lookup table</entry>
+ </row>
+
<row>
<entry><link linkend="view-pg-config"><structname>pg_config</structname></link></entry>
<entry>compile-time configuration parameters</entry>
@@ -901,6 +906,90 @@ AND c1.path[c2.level] = c2.path[c2.level];
</para>
</sect1>
+ <sect1 id="view-pg-buffer-lookup-table">
+ <title><structname>pg_buffer_lookup_table</structname></title>
+ <indexterm>
+ <primary>pg_buffer_lookup_table</primary>
+ </indexterm>
+ <para>
+ The <structname>pg_buffer_lookup_table</structname> view exposes the current
+ contents of the shared buffer lookup table. Each row represents an entry in
+ the lookup table mapping a relation page to the ID of buffer in which it is
+ cached. The shared buffer lookup table is locked for a short duration while
+ reading so as to ensure consistency. This may affect performance if this view
+ is queried very frequently.
+ </para>
+ <table id="pg-buffer-lookup-table-view" xreflabel="pg_buffer_lookup_table">
+ <title><structname>pg_buffer_lookup_table</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>tablespace</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the tablespace containing the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>database</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database containing the relation (zero for shared relations)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relfilenode</structfield> <type>oid</type>
+ </para>
+ <para>
+ relfilenode identifying the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>forknum</structfield> <type>int2</type>
+ </para>
+ <para>
+ Fork number within the relation (see <xref linkend="storage-file-layout"/>)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>blocknum</structfield> <type>int8</type>
+ </para>
+ <para>
+ Block number within the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>bufferid</structfield> <type>int4</type>
+ </para>
+ <para>
+ ID of the buffer caching the page
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ <para>
+ Access to this view is restricted to members of the
+ <literal>pg_read_all_stats</literal> role by default.
+ </para>
+ </sect1>
+
<sect1 id="view-pg-config">
<title><structname>pg_config</structname></title>
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 9d256559bab..f0c39ec2822 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -21,7 +21,12 @@
*/
#include "postgres.h"
+#include "fmgr.h"
+#include "funcapi.h"
#include "storage/buf_internals.h"
+#include "storage/lwlock.h"
+#include "utils/rel.h"
+#include "utils/builtins.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -159,3 +164,56 @@ BufTableDelete(BufferTag *tagPtr, uint32 hashcode)
if (!result) /* shouldn't happen */
elog(ERROR, "shared buffer hash table corrupted");
}
+
+/*
+ * BufTableGetContents
+ * Fill the given tuplestore with contents of the shared buffer lookup table
+ *
+ * This function is used by pg_buffercache extension to expose buffer lookup
+ * table contents via SQL. The caller is responsible for setting up the
+ * tuplestore and result set info.
+ */
+void
+BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc)
+{
+/* Expected number of attributes of the buffer lookup table entry. */
+#define BUFTABLE_CONTENTS_COLS 6
+
+ HASH_SEQ_STATUS hstat;
+ BufferLookupEnt *ent;
+ Datum values[BUFTABLE_CONTENTS_COLS];
+ bool nulls[BUFTABLE_CONTENTS_COLS];
+ int i;
+
+ memset(nulls, 0, sizeof(nulls));
+
+ Assert(tupdesc->natts == BUFTABLE_CONTENTS_COLS);
+
+ /*
+ * Lock all buffer mapping partitions to ensure a consistent view of the
+ * hash table during the scan. Must grab LWLocks in partition-number order
+ * to avoid LWLock deadlock.
+ */
+ for (i = 0; i < NUM_BUFFER_PARTITIONS; i++)
+ LWLockAcquire(BufMappingPartitionLockByIndex(i), LW_SHARED);
+
+ hash_seq_init(&hstat, SharedBufHash);
+ while ((ent = (BufferLookupEnt *) hash_seq_search(&hstat)) != NULL)
+ {
+ values[0] = ObjectIdGetDatum(ent->key.spcOid);
+ values[1] = ObjectIdGetDatum(ent->key.dbOid);
+ values[2] = ObjectIdGetDatum(ent->key.relNumber);
+ values[3] = ObjectIdGetDatum(ent->key.forkNum);
+ values[4] = Int64GetDatum(ent->key.blockNum);
+ values[5] = Int32GetDatum(ent->id);
+
+ tuplestore_putvalues(tupstore, tupdesc, values, nulls);
+ }
+
+ /*
+ * Release all buffer mapping partition locks in the reverse order so as
+ * to avoid LWLock deadlock.
+ */
+ for (i = NUM_BUFFER_PARTITIONS - 1; i >= 0; i--)
+ LWLockRelease(BufMappingPartitionLockByIndex(i));
+}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 5400c56a965..519692702a0 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -28,6 +28,7 @@
#include "storage/spin.h"
#include "utils/relcache.h"
#include "utils/resowner.h"
+#include "utils/tuplestore.h"
/*
* Buffer state is a single 32-bit variable where following data is combined.
@@ -520,6 +521,7 @@ extern uint32 BufTableHashCode(BufferTag *tagPtr);
extern int BufTableLookup(BufferTag *tagPtr, uint32 hashcode);
extern int BufTableInsert(BufferTag *tagPtr, uint32 hashcode, int buf_id);
extern void BufTableDelete(BufferTag *tagPtr, uint32 hashcode);
+extern void BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc);
/* localbuf.c */
extern bool PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount);
base-commit: df53fa1c1ebf9bb3e8c17217f7cc1435107067fb
--
2.34.1
[text/x-patch] 0004-WIP-test-shared-buffers-resizing-and-checkp-20251114.patch (11.2K, ../../CAExHW5sVxEwQsuzkgjjJQP9-XVe0H2njEVw1HxeYFdT7u7J+eQ@mail.gmail.com/3-0004-WIP-test-shared-buffers-resizing-and-checkp-20251114.patch)
download | inline diff:
From 17b83eb9d1b5a825e1e2bfca9d360a738213bf01 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 1 Oct 2025 09:38:19 +0530
Subject: [PATCH 4/4] WIP test shared buffers resizing and checkpoint
A new test triggers an injection point in the BufferSync() after it has
collected buffers to flushed. Simultaneously it starts buffer shrinking. The
expectation is that the checkpointer would crash accessing a buffer (descriptor)
outside the new range of shared buffers. But that does not happen because of a
bug in synchronization. The checkpointer does not reload configuration when
checkpoint is going on. It does not load the new value of the configuration.
When the resizing is triggered by the PM, checkpointer receives the proc signal
barrier but it does not start it doesn't enter the barrier mechanism and doesn't
alter its address maps or memory sizes. Hence the test does not crash. But of
course it means that it won't consider the correct size of buffers next time it
performs a checkpoint.
The test was at least useful to detect this anomaly. Once we fix the
synchronization issue we should see the crash and then fix the crash.
Author: Ashutosh Bapat
Notes to reviewers
------------------
1. pg_buffercache used a query on pg_settings to fetch the value of the
number of buffers. That doesn't work anymore because of change in the
SHOW shared_buffers. Modified the test to convert the setting value
to the number of shared buffers, save it in a variable and use the
variable in queries which need the number of shared buffers. We could
instead fix ShowGUCOption() to pass use_units flag to show_hook and
let it output the number of shared buffers instead. But that seems a
larger change. There aren't other GUCs whose show_hook outputs their
values with units. So this local fix might be better.
---
.../expected/pg_buffercache.out | 19 ++-
contrib/pg_buffercache/sql/pg_buffercache.sql | 19 ++-
src/backend/storage/buffer/bufmgr.c | 4 +
src/test/buffermgr/meson.build | 1 +
.../t/002_checkpoint_buffer_resize.pl | 111 ++++++++++++++++++
5 files changed, 130 insertions(+), 24 deletions(-)
create mode 100644 src/test/buffermgr/t/002_checkpoint_buffer_resize.pl
diff --git a/contrib/pg_buffercache/expected/pg_buffercache.out b/contrib/pg_buffercache/expected/pg_buffercache.out
index 2f27bf34637..632b12abbf8 100644
--- a/contrib/pg_buffercache/expected/pg_buffercache.out
+++ b/contrib/pg_buffercache/expected/pg_buffercache.out
@@ -1,8 +1,9 @@
CREATE EXTENSION pg_buffercache;
-select count(*) = (select setting::bigint
- from pg_settings
- where name = 'shared_buffers')
-from pg_buffercache;
+select pg_size_bytes(setting)/(select setting::bigint from pg_settings where name = 'block_size') AS nbuffers
+ from pg_settings
+ where name = 'shared_buffers'
+\gset
+select count(*) = :nbuffers from pg_buffercache;
?column?
----------
t
@@ -24,20 +25,14 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
(1 row)
-- Test the buffer lookup table function and count is <= shared_buffers
-select count(*) <= (select setting::bigint
- from pg_settings
- where name = 'shared_buffers')
-from pg_buffercache_lookup_table_entries();
+select count(*) <= :nbuffers from pg_buffercache_lookup_table_entries();
?column?
----------
t
(1 row)
-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
-select count(*) <= (select setting::bigint
- from pg_settings
- where name = 'shared_buffers')
-from pg_buffercache_lookup_table;
+select count(*) <= :nbuffers from pg_buffercache_lookup_table;
?column?
----------
t
diff --git a/contrib/pg_buffercache/sql/pg_buffercache.sql b/contrib/pg_buffercache/sql/pg_buffercache.sql
index 569b28aebb9..11fe85ceb3b 100644
--- a/contrib/pg_buffercache/sql/pg_buffercache.sql
+++ b/contrib/pg_buffercache/sql/pg_buffercache.sql
@@ -1,9 +1,10 @@
CREATE EXTENSION pg_buffercache;
-select count(*) = (select setting::bigint
- from pg_settings
- where name = 'shared_buffers')
-from pg_buffercache;
+select pg_size_bytes(setting)/(select setting::bigint from pg_settings where name = 'block_size') AS nbuffers
+ from pg_settings
+ where name = 'shared_buffers'
+\gset
+select count(*) = :nbuffers from pg_buffercache;
select buffers_used + buffers_unused > 0,
buffers_dirty <= buffers_used,
@@ -13,16 +14,10 @@ from pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
-- Test the buffer lookup table function and count is <= shared_buffers
-select count(*) <= (select setting::bigint
- from pg_settings
- where name = 'shared_buffers')
-from pg_buffercache_lookup_table_entries();
+select count(*) <= :nbuffers from pg_buffercache_lookup_table_entries();
-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
-select count(*) <= (select setting::bigint
- from pg_settings
- where name = 'shared_buffers')
-from pg_buffercache_lookup_table;
+select count(*) <= :nbuffers from pg_buffercache_lookup_table;
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 6c8f8552a4c..f489ae2932f 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -67,6 +67,7 @@
#include "utils/rel.h"
#include "utils/resowner.h"
#include "utils/timestamp.h"
+#include "utils/injection_point.h"
/* Note: these two macros only work on shared buffers, not local ones! */
@@ -3416,6 +3417,9 @@ BufferSync(int flags)
ProcessProcSignalBarrier();
}
+ /* Injection point after scanning all buffers for dirty pages */
+ INJECTION_POINT("buffer-sync-dirty-buffer-scan", NULL);
+
if (num_to_scan == 0)
return; /* nothing to do */
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
index c24bff721e6..f33feb64a06 100644
--- a/src/test/buffermgr/meson.build
+++ b/src/test/buffermgr/meson.build
@@ -16,6 +16,7 @@ tests += {
},
'tests': [
't/001_resize_buffer.pl',
+ 't/002_checkpoint_buffer_resize.pl',
't/003_parallel_resize_buffer.pl',
't/004_client_join_buffer_resize.pl',
],
diff --git a/src/test/buffermgr/t/002_checkpoint_buffer_resize.pl b/src/test/buffermgr/t/002_checkpoint_buffer_resize.pl
new file mode 100644
index 00000000000..9ab615b6557
--- /dev/null
+++ b/src/test/buffermgr/t/002_checkpoint_buffer_resize.pl
@@ -0,0 +1,111 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test shared_buffer resizing coordination with checkpoint using injection points
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Initialize cluster with injection points enabled
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', 'shared_buffers = 256kB');
+# Disable background writer to prevent interference with dirty buffers
+$node->append_conf('postgresql.conf', 'bgwriter_lru_maxpages = 0');
+$node->start;
+
+# Load the injection points extension
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Create some data to make checkpoint meaningful and ensure many dirty buffers
+$node->safe_psql('postgres', "CREATE TABLE test_data (id int, data text)");
+# Insert enough data to fill more than 16 buffers (each row ~1KB, so 20+ rows per page)
+$node->safe_psql('postgres', "INSERT INTO test_data SELECT i, repeat('x', 1000) FROM generate_series(1, 5000) i");
+
+# Create additional tables to ensure we have plenty of dirty buffers
+$node->safe_psql('postgres', "CREATE TABLE test_data2 AS SELECT * FROM test_data WHERE id <= 2500");
+$node->safe_psql('postgres', "CREATE TABLE test_data3 AS SELECT * FROM test_data WHERE id > 2500");
+
+# Update data to create more dirty buffers
+$node->safe_psql('postgres', "UPDATE test_data SET data = repeat('y', 1000) WHERE id % 3 = 0");
+$node->safe_psql('postgres', "UPDATE test_data2 SET data = repeat('z', 1000) WHERE id % 2 = 0");
+
+# Prepare the new shared_buffers configuration before starting checkpoint
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '128kB'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+# Set up the injection point to make checkpoint wait
+$node->safe_psql('postgres', "SELECT injection_points_attach('buffer-sync-dirty-buffer-scan', 'wait')");
+
+# Start a checkpoint in the background that will trigger the injection point
+my $checkpoint_session = $node->background_psql('postgres');
+$checkpoint_session->query_until(
+ qr/starting_checkpoint/,
+ q(
+ \echo starting_checkpoint
+ CHECKPOINT;
+ \q
+ )
+);
+
+# Wait until checkpointer actually reaches the injection point
+$node->wait_for_event('checkpointer', 'buffer-sync-dirty-buffer-scan');
+
+# Verify checkpoint is waiting by checking if it hasn't completed
+my $checkpoint_running = $node->safe_psql('postgres',
+ "SELECT COUNT(*) FROM pg_stat_activity WHERE backend_type = 'checkpointer' AND wait_event = 'buffer-sync-dirty-buffer-scan'");
+is($checkpoint_running, '1', 'Checkpoint is waiting at injection point');
+
+# Start the resize operation in the background (don't wait for completion)
+my $resize_session = $node->background_psql('postgres');
+$resize_session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+);
+
+# Continue the checkpoint and wait for its completion
+my $log_offset = -s $node->logfile;
+$node->safe_psql('postgres', "SELECT injection_points_wakeup('buffer-sync-dirty-buffer-scan')");
+
+# Wait for both checkpoint and resize to complete
+$node->wait_for_log(qr/checkpoint complete/, $log_offset);
+
+# Wait for the resize operation to complete using the proper method
+$resize_session->query(q(\echo 'resize_complete'));
+
+pass('Checkpoint and buffer resize both completed after injection point was released');
+
+# Verify the resize actually worked
+is($node->safe_psql('postgres', "SHOW shared_buffers"), '128kB',
+ 'Buffer resize completed successfully after checkpoint coordination');
+
+# Cleanup the background session
+$resize_session->quit;
+
+# Clean up the injection point
+$node->safe_psql('postgres', "SELECT injection_points_detach('buffer-sync-dirty-buffer-scan')");
+
+# Verify system remains stable after coordinated operations
+
+# Perform a normal checkpoint to ensure everything is working
+$node->safe_psql('postgres', "CHECKPOINT");
+
+pass('System remains stable after injection point testing');
+
+# Cleanup
+$node->safe_psql('postgres', "DROP TABLE test_data, test_data2, test_data3");
+
+done_testing();
\ No newline at end of file
--
2.34.1
[text/x-patch] 0003-Allow-to-resize-shared-memory-without-resta-20251114.patch (138.9K, ../../CAExHW5sVxEwQsuzkgjjJQP9-XVe0H2njEVw1HxeYFdT7u7J+eQ@mail.gmail.com/4-0003-Allow-to-resize-shared-memory-without-resta-20251114.patch)
download | inline diff:
From e49b35344a9cbbb7e3229df2e7a2ff0a83290d59 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 14:16:55 +0200
Subject: [PATCH 3/4] Allow to resize shared memory without restart
shared_buffers is now PGC_SIGHUP instead of PGC_POSTMASTER. The value
of this GUC is saved in NBuffersPending instead of NBuffers. When the
server starts, the shared memory size is estimated and the memory is
allocated using NBuffersPending.
When a server is running, the new value of GUC (set using ALTER SYSTEM
... SET shared_buffers = ...; followed by SELECT pg_reload_conf()) does
not come into effect immediately. Instead a function
pg_resize_shared_buffers() is used to resize the buffer pool. The
function uses the current value of GUC in the backends where it is
executed. The function also coordinates the buffer resizing
synchronization across backends.
SHOW shared_buffers now shows the current size of the shared buffer pool
but it also shows pending size of shared buffers, if any.
A new GUC max_shared_buffers is introduced to control the maximum value
of shared_buffers that can be set. By default it is 0. When explicitly
set it needs to be higher than 'shared_buffers'. When max_shared_buffers
is set to 0, it assumes the same value as GUC shared_buffers. This GUC
determines the size of address space reserved for future buffer pool
sizes and the size of buffer look up table.
TBD: Describe the protocol used by pg_resize_shared_buffers() to
synchronize buffer resizing operation with other backends.
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
If a buffer being evicted is pinned, we abort the resizing operation.
There are other alternative which are not implemented in the current
patches 1. to wait for the pinned buffer to get unpinned, 2. the backend
is killed or it itself cancels the query or 3. rollback the operation.
Note that option 1 and 2 would require the pinning related local and
shared records to be accessed. But we need infrastructure to do either
of this right now.
So far the buffer pool metdata (NBuffers and the shared memory segment
address space) is saved in process local heap memory since it's static
for the life of a server. It is passed to a new backend through
Postmaster. But with buffer pool being resized while the server running,
we need Postmaster to update its buffer pool metadata as the resizing
progresses and pass it to the new backend. This has few complications:
1. Postmaster does not receive ProcSignalBarrier. So we need to signal
it separately.
2. Postmaster's local state is inherited by the new backend when
fork()ed. But we need more complex implementation to pass it to an
exec()ed backend.
3. A new backend may receive the updated state from Postmaster and also
the signal barrier which prompts the same update. Thus the proc
signal barrier code needs to be idempotent; adding further complexity
to it.
4. This task takes away Postmaster resources from it's core
functionality.
This can be avoided by following two changes:
1. The shared memory is resized only in a single backend without
requiring any changes to the memory address space.
2. Maintaining the buffer pool metadata in the shared memory instead of
process local memory. This change may affect performance so verify
that performance is not degraded.
TODO: In case the backend executing pg_resize_shared_buffers() exits
before the operation finishes, we need to make sure that the changes
made to the shared memory while resizing are cleaned up properly.
Removing the evicted buffers from buffer ring
=============================================
If the buffer pool has been shrunk, the buffers in the buffer ring may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition and masks any bug that
introduces an invalid buffer into the ring. The alternative fix is more
complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure thus accessible when processing buffer
resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work.
Author: Ashutosh Bapat
Author: Dmitrii Dolgov
Author of some tests: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
Reviewed-by: Tomas Vondra
More detailed note follow: Need to see which of those fit in the commit
message and which should be removed.
Reinitializing strategry control area
=====================================
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
2. &StrategyControl->buffer_strategy_lock need not be initialized again.
3. nextVictimBuffer, completePasses and numBufferAllocs are viewed in
the context of NBuffers. Now that NBuffers itself has changed, those
three do not make sense. Reset them as if the server has restarted
again.
Ability to delay resizing operation
===================================
This commit introduces a flag delay_shmem_resize, which postgresql
backends and workers can use to signal the coordinator to delay resizing
operation. Background writer sets this flag when its scanning buffers.
Background writer operation (needs a rethink)
===========================
Background writer is blocked when the actual resizing is in progress. It
stops a scan in progress when it sees that the resizing has begun or is
about to begin. Once the buffer resizing is finished, before resuming
the regular operation, bgwriter resets the information saved so far.
This information is viewed in the context of NBuffers and hence does not
make sense after resizing which chanegs NBuffers.
Buffer lookup table
===================
Right now there is no way to free shared memory. Even if we shrink the
buffer lookup table when shrinking the buffer pool the unused hash table
entries can not be freed. When we expand the buffer pool, more entries
can be allocated but we can not resize the hash table directory without
rehashing all the entries. Just allocating more entries will lead to
more contention. Hence we setup the buffer lookup table considering the
maximum possible size of the buffer pool which is MaxAvailableMemory
only once at the beginning. Shared buffer lookup table and
StrategyControl are not resized even if the buffer pool is resized hence
they are allocated in the main shared memory segment
BgWriter refactoring
====================
The way BgBufferSync is written today, it packs four functionalities:
setting up the buffer sync state, performing the buffer sync, resetting
the buffer sync state when bgwriter_lru_maxpages <= 0 and setting it up
again after bgwriter_lru_maxpages > 0. That makes the code hard to read.
It will be good to divide this function into 3/4 different functions
each performing one functionality. Then pack all the state (the local
variables from that function converted to static global) into a
structure, which is passed to these functions. Once that happens
BgBufferSyncReset() will call one of the functions to reset the state
when buffer pool is resized.
---
contrib/pg_buffercache/pg_buffercache_pages.c | 18 +-
doc/src/sgml/config.sgml | 44 +-
doc/src/sgml/func/func-admin.sgml | 57 +++
src/backend/access/transam/slru.c | 2 +-
src/backend/access/transam/xlog.c | 2 +-
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/port/sysv_shmem.c | 176 +++++++-
src/backend/postmaster/checkpointer.c | 12 +-
src/backend/postmaster/postmaster.c | 10 +-
src/backend/storage/buffer/Makefile | 3 +-
src/backend/storage/buffer/buf_init.c | 279 ++++++++++--
src/backend/storage/buffer/buf_resize.c | 399 ++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 9 +-
src/backend/storage/buffer/bufmgr.c | 160 ++++++-
src/backend/storage/buffer/freelist.c | 106 ++++-
src/backend/storage/buffer/meson.build | 1 +
src/backend/storage/ipc/ipci.c | 13 +-
src/backend/storage/ipc/procsignal.c | 55 +++
src/backend/storage/ipc/shmem.c | 66 ++-
src/backend/tcop/postgres.c | 11 +
.../utils/activity/wait_event_names.txt | 2 +
src/backend/utils/init/globals.c | 5 +-
src/backend/utils/init/postinit.c | 49 +++
src/backend/utils/misc/guc.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 15 +-
src/include/catalog/pg_proc.dat | 6 +
src/include/miscadmin.h | 5 +
src/include/storage/buf_internals.h | 1 +
src/include/storage/bufmgr.h | 20 +-
src/include/storage/ipc.h | 3 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 52 ++-
src/include/storage/pmsignal.h | 3 +-
src/include/storage/procsignal.h | 4 +
src/include/storage/shmem.h | 3 +
src/include/utils/guc.h | 2 +
src/test/Makefile | 2 +-
src/test/README | 3 +
src/test/buffermgr/Makefile | 30 ++
src/test/buffermgr/README | 26 ++
src/test/buffermgr/buffermgr_test.conf | 11 +
src/test/buffermgr/expected/buffer_resize.out | 329 +++++++++++++++
src/test/buffermgr/meson.build | 23 +
src/test/buffermgr/sql/buffer_resize.sql | 95 +++++
src/test/buffermgr/t/001_resize_buffer.pl | 135 ++++++
.../buffermgr/t/003_parallel_resize_buffer.pl | 71 ++++
.../t/004_client_join_buffer_resize.pl | 241 +++++++++++
src/test/meson.build | 1 +
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 76 ++++
src/tools/pgindent/typedefs.list | 1 +
50 files changed, 2535 insertions(+), 107 deletions(-)
create mode 100644 src/backend/storage/buffer/buf_resize.c
create mode 100644 src/test/buffermgr/Makefile
create mode 100644 src/test/buffermgr/README
create mode 100644 src/test/buffermgr/buffermgr_test.conf
create mode 100644 src/test/buffermgr/expected/buffer_resize.out
create mode 100644 src/test/buffermgr/meson.build
create mode 100644 src/test/buffermgr/sql/buffer_resize.sql
create mode 100644 src/test/buffermgr/t/001_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/003_parallel_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/004_client_join_buffer_resize.pl
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index fe9af45febe..e311e3d266c 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -118,6 +118,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
TupleDesc tupledesc;
TupleDesc expected_tupledesc;
HeapTuple tuple;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
if (SRF_IS_FIRSTCALL())
{
@@ -174,10 +175,10 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
/* Allocate NBuffers worth of BufferCachePagesRec records. */
fctx->record = (BufferCachePagesRec *)
MemoryContextAllocHuge(CurrentMemoryContext,
- sizeof(BufferCachePagesRec) * NBuffers);
+ sizeof(BufferCachePagesRec) * currentNBuffers);
/* Set max calls and remember the user function context. */
- funcctx->max_calls = NBuffers;
+ funcctx->max_calls = currentNBuffers;
funcctx->user_fctx = fctx;
/* Return to original context when allocating transient memory */
@@ -191,13 +192,24 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
* snapshot across all buffers, but we do grab the buffer header
* locks, so the information of each buffer is self-consistent.
*/
- for (i = 0; i < NBuffers; i++)
+ for (i = 0; i < currentNBuffers; i++)
{
BufferDesc *bufHdr;
uint32 buf_state;
CHECK_FOR_INTERRUPTS();
+ /*
+ * TODO: We should just scan the entire buffer descriptor
+ * array instead of relying on curent buffer pool size. But that can
+ * happen if only we setup the descriptor array large enough at the
+ * server startup time.
+ */
+ if (currentNBuffers != pg_atomic_read_u32(&ShmemCtrl->currentNBuffers))
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("number of shared buffers changed during scan of buffer cache")));
+
bufHdr = GetBufferDescriptor(i);
/* Lock each buffer header before inspecting. */
buf_state = LockBufHdr(bufHdr);
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 683f7c36f46..e4cd9b1f555 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -1724,7 +1724,6 @@ include_dir 'conf.d'
that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
(Non-default values of <symbol>BLCKSZ</symbol> change the minimum
value.)
- This parameter can only be set at server start.
</para>
<para>
@@ -1747,6 +1746,49 @@ include_dir 'conf.d'
appropriate, so as to leave adequate space for the operating system.
</para>
+ <para>
+ The shared memory consumed by the buffer pool is allocated and
+ initialized according to the value of the GUC at the time of starting
+ the server. A desired new value of GUC can be loaded while the server is
+ running using <systemitem>SIGHUP</systemitem>. But the buffer pool will
+ not be resized immediately. Use
+ <function>pg_resize_shared_buffers()</function> to dynamically resize
+ the shared buffer pool (see <xref linkend="functions-admin"/> for details).
+ <command>SHOW shared_buffers</command> shows the current number of
+ shared buffers and pending number, if any. Please note that when the GUC
+ is changed, the other GUCS which use this GUCs value to set their
+ defaults will not be changed. They may still require a server restart to
+ consider new value.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-max-shared-buffers" xreflabel="max_shared_buffers">
+ <term><varname>max_shared_buffers</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>max_shared_buffers</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Sets the upper limit for the <varname>shared_buffers</varname> value.
+ The default value is <literal>0</literal>,
+ which means no explicit limit is set and <varname>max_shared_buffers</varname>
+ will be automatically set to the value of <varname>shared_buffers</varname>
+ at server startup.
+ If this value is specified without units, it is taken as blocks,
+ that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
+ This parameter can only be set at server start.
+ </para>
+
+ <para>
+ This parameter determines the amount of memory address space to reserve
+ in each backend for expanding the buffer pool in future. While the
+ memory for buffer pool is allocated on demand as it is resized, the
+ memory required to hold the buffer manager metadata is allocated
+ statically at the server start accounting for the largest buffer pool
+ size allowed by this parameter.
+ </para>
</listitem>
</varlistentry>
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 1b465bc8ba7..0dc89b07c76 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -99,6 +99,63 @@
<returnvalue>off</returnvalue>
</para></entry>
</row>
+
+ <row>
+ <entry role="func_table_entry"><para role="func_signature">
+ <indexterm>
+ <primary>pg_resize_shared_buffers</primary>
+ </indexterm>
+ <function>pg_resize_shared_buffers</function> ()
+ <returnvalue>boolean</returnvalue>
+ </para>
+ <para>
+ Dynamically resizes the shared buffer pool to match the current
+ value of the <varname>shared_buffers</varname> parameter. This
+ function implements a coordinated resize process that ensures all
+ backend processes acknowledge the change before completing the
+ operation. The resize happens in multiple phases to maintain
+ data consistency and system stability. Returns <literal>true</literal>
+ if the resize was successful, or raises an error if the operation
+ fails. This function can only be called by superusers.
+ </para>
+ <para>
+ To resize shared buffers, first update the <varname>shared_buffers</varname>
+ setting and reload the configuration, then verify the new value is loaded
+ before calling this function. For example:
+<programlisting>
+postgres=# ALTER SYSTEM SET shared_buffers = '256MB';
+ALTER SYSTEM
+postgres=# SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers;
+ shared_buffers
+-------------------------
+ 128MB (pending: 256MB)
+(1 row)
+
+postgres=# SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+</programlisting>
+ The <command>SHOW shared_buffers</command> step is important to verify
+ that the configuration reload was successful and the new value is
+ available to the current session before attempting the resize. The
+ output shows both the current and pending values when a change is waiting
+ to be applied.
+ </para></entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/access/transam/slru.c b/src/backend/access/transam/slru.c
index 77676d6d035..73df5909886 100644
--- a/src/backend/access/transam/slru.c
+++ b/src/backend/access/transam/slru.c
@@ -232,7 +232,7 @@ SimpleLruAutotuneBuffers(int divisor, int max)
{
return Min(max - (max % SLRU_BANK_SIZE),
Max(SLRU_BANK_SIZE,
- NBuffers / divisor - (NBuffers / divisor) % SLRU_BANK_SIZE));
+ NBuffersPending / divisor - (NBuffersPending / divisor) % SLRU_BANK_SIZE));
}
/*
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 22d0a2e8c3a..f4363e0035d 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4676,7 +4676,7 @@ XLOGChooseNumBuffers(void)
{
int xbuffers;
- xbuffers = NBuffers / 32;
+ xbuffers = NBuffersPending / 32;
if (xbuffers > (wal_segment_size / XLOG_BLCKSZ))
xbuffers = (wal_segment_size / XLOG_BLCKSZ);
if (xbuffers < 8)
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index fc8638c1b61..226944e4588 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -335,6 +335,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ InitializeMaxNBuffers();
+
CreateSharedMemoryAndSemaphores();
/*
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index cc4b2c80e1a..68de301441b 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,13 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
@@ -98,6 +104,8 @@ typedef enum
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
+volatile bool delay_shmem_resize = false;
+
/*
* Anonymous mapping layout we use looks like this:
*
@@ -124,6 +132,9 @@ void *UsedShmemSegAddr = NULL;
* being counted against memory limits). The mapping serves as an address space
* reservation, into which shared memory segment can be extended and is
* represented by the second /memfd:main with no permissions.
+ *
+ * The reserved space for buffer manager related segments is calculated based on
+ * MaxNBuffers.
*/
/*
@@ -134,6 +145,42 @@ void *UsedShmemSegAddr = NULL;
*/
static bool huge_pages_on = false;
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -156,8 +203,6 @@ MappingName(int shmem_segment)
return "iocv";
case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
return "checkpoint";
- case STRATEGY_SHMEM_SEGMENT:
- return "strategy";
default:
return "unknown";
}
@@ -921,6 +966,114 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the new shared_buffers value (saved
+ * in ShmemCtrl area). The actual segment resizing is done via ftruncate, which
+ * will fail if there is not sufficient space to expand the anon file.
+ *
+ * TODO: Rename this to BufferShmemResize() or something. Only buffer manager's
+ * memory should be resized in this function.
+ *
+ * TODO: This function changes the amount of shared memory used. So it should
+ * also update the show only GUCs shared_memory_size and
+ * shared_memory_size_in_huge_pages in all backends. SetConfigOption() may be
+ * used for that. But it's not clear whether is_reload parameter is safe to use
+ * while resizing is going on; also at what stage it should be done.
+ */
+bool
+AnonymousShmemResize(void)
+{
+ int mmap_flags = PG_MMAP_FLAGS;
+ Size hugepagesize;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /* TODO: This is a hack. NBuffersPending should never be written by anything
+ * other than GUC system. Find a way to pass new NBuffers value to
+ * BufferManagerShmemSize(). */
+ NBuffersPending = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffers, NBuffersPending);
+
+#ifndef MAP_HUGETLB
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+ }
+#endif
+
+ /* Note that BufferManagerShmemSize() indirectly depends on NBuffersPending. */
+ BufferManagerShmemSize(mapping_sizes);
+
+ for(int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ MemoryMappingSizes *mapping = &mapping_sizes[i];
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmem_hdr = segment->ShmemSegHdr;
+
+ /* Main shared memory segment is always static. Ignore it. */
+ if (i == MAIN_SHMEM_SEGMENT)
+ continue;
+
+ round_off_mapping_sizes(mapping);
+ round_off_mapping_sizes_for_hugepages(mapping, hugepagesize);
+
+ /*
+ * Size of the reserved address space should not change, since it depends
+ * upon MaxNBuffers, which can be changed only on restart.
+ */
+ Assert(segment->shmem_reserved == mapping->shmem_reserved);
+#ifdef MAP_HUGETLB
+ if (huge_pages_on && (mapping_sizes->shmem_req_size % hugepagesize != 0))
+ mapping_sizes->shmem_req_size += hugepagesize - (mapping_sizes->shmem_req_size % hugepagesize);
+#endif
+ elog(DEBUG1, "segment[%s]: requested size %zu, current size %zu, reserved %zu",
+ MappingName(i), mapping->shmem_req_size, segment->shmem_size,
+ segment->shmem_reserved);
+
+ if (segment->shmem == NULL)
+ continue;
+
+ if (segment->shmem_size == mapping->shmem_req_size)
+ continue;
+
+ /*
+ * We should have reserved enough address space for resizing. PANIC if
+ * that's not the case.
+ */
+ if (segment->shmem_reserved < mapping->shmem_req_size)
+ ereport(PANIC,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved")));
+
+ elog(DEBUG1, "segment[%s]: resize from %zu to %zu at address %p",
+ MappingName(i), segment->shmem_size,
+ mapping->shmem_req_size, segment->shmem);
+
+ /*
+ * Resize the backing file to resize the allocated memory, and allocate
+ * more memory on supported platforms if required.
+ */
+ if(ftruncate(segment->segment_fd, mapping->shmem_req_size) == -1)
+ ereport(ERROR,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncate anonymous file for \"%s\": %m",
+ MappingName(i))));
+ if (mapping->shmem_req_size > segment->shmem_size)
+ shmem_fallocate(segment->segment_fd, MappingName(i), mapping->shmem_req_size, ERROR);
+
+ segment->shmem_size = mapping->shmem_req_size;
+ shmem_hdr->totalsize = segment->shmem_size;
+ segment->ShmemEnd = segment->shmem + segment->shmem_size;
+ }
+
+ return true;
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1224,3 +1377,22 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ pg_atomic_init_u32(&ShmemCtrl->targetNBuffers, 0);
+ pg_atomic_init_u32(&ShmemCtrl->currentNBuffers, 0);
+ pg_atomic_init_flag(&ShmemCtrl->resize_in_progress);
+
+ ShmemCtrl->coordinator = 0;
+ }
+}
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index e84e8663e96..ef3f84a55f5 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -654,9 +654,12 @@ CheckpointerMain(const void *startup_data, size_t startup_data_len)
static void
ProcessCheckpointerInterrupts(void)
{
- if (ProcSignalBarrierPending)
- ProcessProcSignalBarrier();
-
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
if (ConfigReloadPending)
{
ConfigReloadPending = false;
@@ -676,6 +679,9 @@ ProcessCheckpointerInterrupts(void)
UpdateSharedMemoryConfig();
}
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+
/* Perform logging of memory contexts of this process */
if (LogMemoryContextPending)
ProcessLogMemoryContextInterrupt();
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 7c064cf9fbb..2095713d7c0 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -110,11 +110,15 @@
#include "replication/slotsync.h"
#include "replication/walsender.h"
#include "storage/aio_subsys.h"
+#include "storage/bufmgr.h"
#include "storage/fd.h"
#include "storage/io_worker.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "tcop/backend_startup.h"
#include "tcop/tcopprot.h"
#include "utils/datetime.h"
@@ -125,7 +129,6 @@
#ifdef EXEC_BACKEND
#include "common/file_utils.h"
-#include "storage/pg_shmem.h"
#endif
@@ -958,6 +961,11 @@ PostmasterMain(int argc, char *argv[])
*/
InitializeFastPathLocks();
+ /*
+ * Calculate MaxNBuffers for buffer pool resizing.
+ */
+ InitializeMaxNBuffers();
+
/*
* Give preloaded libraries a chance to request additional shared memory.
*/
diff --git a/src/backend/storage/buffer/Makefile b/src/backend/storage/buffer/Makefile
index fd7c40dcb08..3bc9aee85de 100644
--- a/src/backend/storage/buffer/Makefile
+++ b/src/backend/storage/buffer/Makefile
@@ -17,6 +17,7 @@ OBJS = \
buf_table.o \
bufmgr.o \
freelist.o \
- localbuf.o
+ localbuf.o \
+ buf_resize.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 4fa547f48de..4a354107185 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,7 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
-#include "storage/pg_shmem.h"
+#include "utils/guc.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -62,11 +62,12 @@ CkptSortItem *CkptBufferIds;
/*
* Initialize shared buffer pool
*
- * This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * This is called once during shared-memory initialization.
+ * TODO: Restore this function to it's initial form. This function should see no
+ * change in buffer resize patches, except may be use of NBuffersPending.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
BufferManagerShmemInit(void)
@@ -75,24 +76,25 @@ BufferManagerShmemInit(void)
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
ShmemInitStructInSegment("Buffer Descriptors",
- NBuffers * sizeof(BufferDescPadded),
+ NBuffersPending * sizeof(BufferDescPadded),
&foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
ShmemInitStructInSegment("Buffer Blocks",
- NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ NBuffersPending * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
&foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
ShmemInitStructInSegment("Buffer IO Condition Variables",
- NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ NBuffersPending * sizeof(ConditionVariableMinimallyPadded),
&foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
@@ -104,48 +106,54 @@ BufferManagerShmemInit(void)
*/
CkptBufferIds = (CkptSortItem *)
ShmemInitStructInSegment("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ NBuffersPending * sizeof(CkptSortItem), &foundBufCkpt,
CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
- }
- else
- {
- int i;
-
/*
- * Initialize all the buffer headers.
+ * note: this path is only taken in EXEC_BACKEND case when initializing
+ * shared memory.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
+ }
- ClearBufferTag(&buf->tag);
+ /*
+ * Initialize all the buffer headers.
+ */
+ for (i = 0; i < NBuffersPending; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
- pg_atomic_init_u32(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
- buf->buf_id = i;
+ buf->buf_id = i;
- pgaio_wref_clear(&buf->io_wref);
+ pgaio_wref_clear(&buf->io_wref);
- LWLockInitialize(BufferDescriptorGetContentLock(buf),
- LWTRANCHE_BUFFER_CONTENT);
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
}
- /* Init other shared buffer-management stuff */
+ /*
+ * Init other shared buffer-management stuff.
+ */
StrategyInitialize(!foundDescs);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
&backend_flush_after);
+
+ /* Declare the size of current buffer pool. */
+ NBuffers = NBuffersPending;
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, NBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, NBuffers);
}
/*
@@ -156,6 +164,8 @@ BufferManagerShmemInit(void)
* shared memory segment. The main segment must not allocate anything
* related to buffers, every other segment will receive part of the
* data.
+ *
+ * Also sets the shmem_reserved field for each segment based on MaxNBuffers.
*/
Size
BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
@@ -163,31 +173,222 @@ BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
size_t size;
/* size of buffer descriptors, plus alignment padding */
- size = add_size(0, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ size = add_size(0, mul_size(NBuffersPending, sizeof(BufferDescPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
mapping_sizes[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_req_size = size;
+ size = add_size(0, mul_size(MaxNBuffers, sizeof(BufferDescPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
mapping_sizes[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_reserved = size;
/* size of data pages, plus alignment padding */
size = add_size(0, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ size = add_size(size, mul_size(NBuffersPending, BLCKSZ));
mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
+ size = add_size(0, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(MaxNBuffers, BLCKSZ));
mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_reserved = size;
- /* size of stuff controlled by freelist.c */
- mapping_sizes[STRATEGY_SHMEM_SEGMENT].shmem_req_size = StrategyShmemSize();
- mapping_sizes[STRATEGY_SHMEM_SEGMENT].shmem_reserved = StrategyShmemSize();
-
/* size of I/O condition variables, plus alignment padding */
- size = add_size(0, mul_size(NBuffers,
+ size = add_size(0, mul_size(NBuffersPending,
sizeof(ConditionVariableMinimallyPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
mapping_sizes[BUFFER_IOCV_SHMEM_SEGMENT].shmem_req_size = size;
+ size = add_size(0, mul_size(MaxNBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
mapping_sizes[BUFFER_IOCV_SHMEM_SEGMENT].shmem_reserved = size;
/* size of checkpoint sort array in bufmgr.c */
- mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
- mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_reserved = mul_size(NBuffers, sizeof(CkptSortItem));
+ mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffersPending, sizeof(CkptSortItem));
+ mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_reserved = mul_size(MaxNBuffers, sizeof(CkptSortItem));
+
+ /* Allocations in the main memory segment, at the end. */
+
+ /* size of stuff controlled by freelist.c */
+ size = add_size(0, StrategyShmemSize());
return size;
}
+
+/*
+ * Reinitialize shared buffer manager structures when resizing the buffer pool.
+ *
+ * This function is called in the backend which coordinates buffer resizing
+ * operation.
+ *
+ * TODO: Avoid code duplication with BufferManagerShmemInit() and also assess
+ * which functionality in the latter is required in this function.
+ */
+void
+BufferManagerShmemResize(int currentNBuffers, int targetNBuffers)
+{
+ bool found;
+ int i;
+ void *tmpPtr;
+
+ tmpPtr = (BufferDescPadded *)
+ ShmemUpdateStructInSegment("Buffer Descriptors",
+ targetNBuffers * sizeof(BufferDescPadded),
+ &found, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+ if (BufferDescriptors != tmpPtr || !found)
+ elog(FATAL, "resizing buffer descriptors failed: expected pointer %p, got %p, found=%d",
+ BufferDescriptors, tmpPtr, found);
+
+ tmpPtr = (ConditionVariableMinimallyPadded *)
+ ShmemUpdateStructInSegment("Buffer IO Condition Variables",
+ targetNBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &found, BUFFER_IOCV_SHMEM_SEGMENT);
+ if (BufferIOCVArray != tmpPtr || !found)
+ elog(FATAL, "resizing buffer IO condition variables failed: expected pointer %p, got %p, found=%d",
+ BufferIOCVArray, tmpPtr, found);
+
+ tmpPtr = (CkptSortItem *)
+ ShmemUpdateStructInSegment("Checkpoint BufferIds",
+ targetNBuffers * sizeof(CkptSortItem), &found,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+ if (CkptBufferIds != tmpPtr || !found)
+ elog(FATAL, "resizing checkpoint buffer IDs failed: expected pointer %p, got %p, found=%d",
+ CkptBufferIds, tmpPtr, found);
+
+ tmpPtr = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemUpdateStructInSegment("Buffer Blocks",
+ targetNBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &found, BUFFERS_SHMEM_SEGMENT));
+ if (BufferBlocks != tmpPtr || !found)
+ elog(FATAL, "resizing buffer blocks failed: expected pointer %p, got %p, found=%d",
+ BufferBlocks, tmpPtr, found);
+
+ /*
+ * Initialize the headers for new buffers. If we are shrinking the
+ * buffers, currentNBuffers >= targetNBuffers, thus this loop doesn't execute.
+ */
+ for (i = currentNBuffers; i < targetNBuffers; i++)
+ {
+ BufferDesc *buf = GetBufferDescriptor(i);
+
+ ClearBufferTag(&buf->tag);
+
+ pg_atomic_init_u32(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+
+ buf->buf_id = i;
+
+ LWLockInitialize(BufferDescriptorGetContentLock(buf),
+ LWTRANCHE_BUFFER_CONTENT);
+
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ }
+
+ /*
+ * We do not touch StrategyControl here. Instead it is done by background
+ * writer when handling PROCSIGNAL_BARRIER_SHBUF_EXPAND or
+ * PROCSIGNAL_BARRIER_SHBUF_SHRINK barrier.
+ */
+}
+
+/*
+ * BufferManagerShmemValidate
+ * Validate that buffer manager shared memory structures have correct
+ * pointers and sizes after a resize operation.
+ *
+ * This function is called by backends during ProcessBarrierShmemResizeStruct
+ * to ensure their view of the buffer structures is consistent after memory
+ * remapping.
+ */
+void
+BufferManagerShmemValidate(int targetNBuffers)
+{
+ bool found;
+ void *tmpPtr;
+
+ /* Validate Buffer Descriptors */
+ tmpPtr = (BufferDescPadded *)
+ ShmemInitStructInSegment("Buffer Descriptors",
+ targetNBuffers * sizeof(BufferDescPadded),
+ &found, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+ if (!found || BufferDescriptors != tmpPtr)
+ elog(FATAL, "validating buffer descriptors failed: expected pointer %p, got %p, found=%d",
+ BufferDescriptors, tmpPtr, found);
+
+ /* Validate Buffer IO Condition Variables */
+ tmpPtr = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ targetNBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &found, BUFFER_IOCV_SHMEM_SEGMENT);
+ if (!found || BufferIOCVArray != tmpPtr)
+ elog(FATAL, "validating buffer IO condition variables failed: expected pointer %p, got %p, found=%d",
+ BufferIOCVArray, tmpPtr, found);
+
+ /* Validate Checkpoint BufferIds */
+ tmpPtr = (CkptSortItem *)
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ targetNBuffers * sizeof(CkptSortItem), &found,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+ if (!found || CkptBufferIds != tmpPtr)
+ elog(FATAL, "validating checkpoint buffer IDs failed: expected pointer %p, got %p, found=%d",
+ CkptBufferIds, tmpPtr, found);
+
+ /* Validate Buffer Blocks */
+ tmpPtr = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSegment("Buffer Blocks",
+ targetNBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &found, BUFFERS_SHMEM_SEGMENT));
+ if (!found || BufferBlocks != tmpPtr)
+ elog(FATAL, "validating buffer blocks failed: expected pointer %p, got %p, found=%d",
+ BufferBlocks, tmpPtr, found);
+}
+
+/*
+ * check_shared_buffers
+ * GUC check_hook for shared_buffers
+ *
+ * When reloading the configuration, shared_buffers should not be set to a value
+ * higher than max_shared_buffers fixed at the boot time.
+ */
+bool
+check_shared_buffers(int *newval, void **extra, GucSource source)
+{
+ if (finalMaxNBuffers && *newval > MaxNBuffers)
+ {
+ GUC_check_errdetail("\"shared_buffers\" must be less than \"max_shared_buffers\".");
+ return false;
+ }
+ return true;
+}
+
+/*
+ * show_shared_buffers
+ * GUC show_hook for shared_buffers
+ *
+ * Shows both current and pending buffer counts with proper unit formatting.
+ */
+const char *
+show_shared_buffers(void)
+{
+ static char buffer[128];
+ int64 current_value, pending_value;
+ const char *current_unit, *pending_unit;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+
+ if (currentNBuffers == NBuffersPending)
+ {
+ /* No buffer pool resizing pending. */
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s", current_value, current_unit);
+ }
+ else
+ {
+ /*
+ * New value for NBuffers is loaded but not applied yet, show both
+ * current and pending.
+ */
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ convert_int_from_base_unit(NBuffersPending, GUC_UNIT_BLOCKS, &pending_value, &pending_unit);
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s (pending: " INT64_FORMAT "%s)",
+ current_value, current_unit, pending_value, pending_unit);
+ }
+
+ return buffer;
+}
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
new file mode 100644
index 00000000000..e815600c3ba
--- /dev/null
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -0,0 +1,399 @@
+/*-------------------------------------------------------------------------
+ *
+ * buf_resize.c
+ * shared buffer pool resizing functionality
+ *
+ * This module contains the implementation of shared buffer pool resizing,
+ * including the main resize coordination function and barrier processing
+ * functions that synchronize all backends during resize operations.
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/storage/buffer/buf_resize.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "miscadmin.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
+#include "utils/injection_point.h"
+
+
+/*
+ * Prepare ShmemCtrl for resizing the shared buffer pool.
+ */
+static void
+MarkBufferResizingStart(int targetNBuffers, int currentNBuffers)
+{
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ Assert(pg_atomic_read_u32(&ShmemCtrl->currentNBuffers) == currentNBuffers);
+
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, targetNBuffers);
+ ShmemCtrl->coordinator = MyProcPid;
+}
+
+/*
+ * Reset ShmemCtrl after resizing the shared buffer pool is done.
+ */
+static void
+MarkBufferResizingEnd(int NBuffers)
+{
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ Assert(pg_atomic_read_u32(&ShmemCtrl->currentNBuffers) == NBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, 0);
+ ShmemCtrl->coordinator = -1;
+}
+
+/*
+ * Communicate given buffer pool resize barrier to all other backends and the Postmaster.
+ *
+ * ProcSignalBarrier is not sent to the Postmaster but we need the Postmaster to
+ * update its knowledge about the buffer pool so that it can be inherited by the
+ * child processes.
+ */
+static void
+SharedBufferResizeBarrier(ProcSignalBarrierType barrier, const char *barrier_name)
+{
+ WaitForProcSignalBarrier(EmitProcSignalBarrier(barrier));
+ elog(LOG, "all backends acknowledged %s barrier", barrier_name);
+
+#ifdef USE_INJECTION_POINTS
+ /* Injection point specific to this barrier type */
+ switch (barrier)
+ {
+ case PROCSIGNAL_BARRIER_SHBUF_SHRINK:
+ INJECTION_POINT("pgrsb-shrink-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM:
+ INJECTION_POINT("pgrsb-resize-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_EXPAND:
+ INJECTION_POINT("pgrsb-expand-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED:
+ /* TODO: Add an injection point here. */
+ break;
+ case PROCSIGNAL_BARRIER_SMGRRELEASE:
+ /*
+ * Not relevant in this function but it's here so that the compiler
+ * can detect any missing shared buffer resizing barrier enum here.
+ */
+ break;
+ }
+#endif /* USE_INJECTION_POINTS */
+}
+
+/*
+ * C implementation of SQL interface to update the shared buffers according to
+ * the current values of shared_buffers GUCs.
+ *
+ * The current boundaries of the buffer pool are given by two ranges.
+ *
+ * - [1, StrategyControl::activeNBuffers] is the range of buffers from which new
+ * allocations can happen at any time.
+ *
+ * - [1, ShmemCtrl::currentNBuffers] is the range of valid buffers at any given
+ * time.
+ *
+ * Let's assume that before resizing, the number of buffers in the buffer pool is
+ * NBuffersOld. After resizing it is NBuffersNew. Before resizing
+ * StrategyControl::activeNBuffers == ShmemCtrl::currentNBuffers == NBuffersOld.
+ * After the resizing finishes StrategyControl::activeNBuffers ==
+ * ShmemCtrl::currentNBuffers == NBuffersNew. Thus when no resizing happens these
+ * two ranges are same.
+ *
+ * Following steps are performed by the coordinator during resizing.
+ *
+ * 1. Marks resizing in progress to avoid multiple concurrent invocations of this
+ * function.
+ *
+ * 2. When shrinking the shared buffer pool, the coordinator sends SHBUF_SHRINK
+ * ProcSignalBarrier. In response to this barrier background writer is expected
+ * to set StrategyControl::activeNBuffers = NBuffersNew to restrict the new
+ * buffer allocations only to the new buffer pool size and also reset its
+ * internal state. Once every backend has acknowledged the barrier, the
+ * coordinator can be sure that new allocations will not happen in the buffer
+ * pool area being shrunk. Then it evicts the buffers in that area. Note that
+ * ShmemCtrl::currentNBuffers is still NBuffersOld, since backend may still
+ * access buffers allocated before the resizing started. Buffer eviction may fail
+ * if a buffer being evicted is pinned and the resizing operatino is aborted.
+ * Once the eviction is finished, the extra memory can be freed in the next step.
+ *
+ * 2. This step is executed in both cases, when expanding the buffer pool or
+ * shrinking the buffer pool. The anonymous file backing each of the shared
+ * memory segment containg the buffer pool shared data structures is resized to
+ * the amount of memory required for the new buffer pool size. When expanding the
+ * expanded portion of memory is initialized appropriately.
+ * ShmemCtrl::currentNBuffers is set to NBuffersNew to indicate new range of
+ * valid shared buffers. Every backend is sent SHBUF_RESIZE_MAP_AND_MEM barrier.
+ * All the backends validate that their pointers to the shared buffers structure
+ * are valid and have the right size. Once every backend has acknowledged the
+ * barrier, this step finishes.
+ *
+ * 3. When expanding the buffer pool, the coordinator sends SHBUF_EXPAND barrier
+ * to signal end of expansion. When expadning the background writer, in response
+ * to StrategyControl::activeNBuffers = NBufferNew so that new allocations can
+ * use expanded range of buffer pool.
+ *
+ * TODO: Handle the case when the backend executing this function dies or the
+ * query is cancelled or it hits an error while resizing.
+ */
+Datum
+pg_resize_shared_buffers(PG_FUNCTION_ARGS)
+{
+ bool result = true;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+ int targetNBuffers = NBuffersPending;
+
+ if (currentNBuffers == targetNBuffers)
+ {
+ elog(LOG, "shared buffers are already at %d, no need to resize", currentNBuffers);
+ PG_RETURN_BOOL(true);
+ }
+
+ if (!pg_atomic_test_set_flag(&ShmemCtrl->resize_in_progress))
+ {
+ elog(LOG, "shared buffer resizing already in progress");
+ PG_RETURN_BOOL(false);
+ }
+
+ /*
+ * TODO: What if the NBuffersPending value seen here is not the desired one
+ * because somebody did a pg_reload_conf() between the last pg_reload_conf()
+ * and execution of this function?
+ */
+ MarkBufferResizingStart(targetNBuffers, currentNBuffers);
+ elog(LOG, "resizing shared buffers from %d to %d", currentNBuffers, targetNBuffers);
+
+ INJECTION_POINT("pg-resize-shared-buffers-flag-set", NULL);
+
+ /* Phase 1: SHBUF_SHRINK - Only for shrinking buffer pool */
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * Phase 1: Shrinking - send SHBUF_SHRINK barrier
+ * Every backend sets activeNBuffers = NewNBuffers to restrict
+ * buffer pool allocations to the new size
+ */
+ elog(LOG, "Phase 1: Shrinking buffer pool, restricting allocations to %d buffers", targetNBuffers);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_SHRINK, CppAsString(PROCSIGNAL_BARRIER_SHBUF_SHRINK));
+
+ /* Evict buffers in the area being shrunk */
+ elog(LOG, "evicting buffers %u..%u", targetNBuffers + 1, currentNBuffers);
+ if (!EvictExtraBuffers(targetNBuffers, currentNBuffers))
+ {
+ elog(WARNING, "failed to evict extra buffers during shrinking");
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED, CppAsString(PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED));
+ MarkBufferResizingEnd(currentNBuffers);
+ pg_atomic_clear_flag(&ShmemCtrl->resize_in_progress);
+ PG_RETURN_BOOL(false);
+ }
+
+ /* Update the current NBuffers. */
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, targetNBuffers);
+ }
+
+ /* Phase 2: SHBUF_RESIZE_MAP_AND_MEM - Both expanding and shrinking */
+ elog(LOG, "Phase 2: Remapping shared memory segments and updating structures");
+ if (!AnonymousShmemResize())
+ {
+ /*
+ * This should never fail since address map should already be reserved.
+ * So the failure should be treated as PANIC.
+ */
+ elog(PANIC, "failed to resize anonymous shared memory");
+ }
+
+ /* Update structure pointers and sizes */
+ BufferManagerShmemResize(currentNBuffers, targetNBuffers);
+
+ INJECTION_POINT("pgrsb-after-shmem-resize", NULL);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM, CppAsString(PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM));
+
+ /* Phase 3: SHBUF_EXPAND - Only for expanding buffer pool */
+ if (targetNBuffers > currentNBuffers)
+ {
+ /*
+ * Phase 3: Expanding - send SHBUF_EXPAND barrier
+ * Backends set activeNBuffers = NewNBuffers and start allocating
+ * buffers from the expanded range
+ */
+ elog(LOG, "Phase 3: Expanding buffer pool, enabling allocations up to %d buffers", targetNBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, targetNBuffers);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_EXPAND, CppAsString(PROCSIGNAL_BARRIER_SHBUF_EXPAND));
+ }
+
+ /*
+ * Reset buffer resize control area.
+ */
+ MarkBufferResizingEnd(targetNBuffers);
+
+ pg_atomic_clear_flag(&ShmemCtrl->resize_in_progress);
+
+ elog(LOG, "successfully resized shared buffers to %d", targetNBuffers);
+
+ PG_RETURN_BOOL(result);
+}
+
+bool
+ProcessBarrierShmemShrink(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * Delay adjusting the new active size of buffer pool till this process
+ * becomes ready to resize buffers.
+ */
+ if (delay_shmem_resize)
+ {
+ elog(LOG, "Phase 1: Delaying SHBUF_SHRINK barrier - restricting allocations to %d buffers, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+
+ return false;
+ }
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation statistics
+ * and the strategy control together so that background writer doesn't go
+ * out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody would
+ * reset the strategy control area. So we can't rely on background
+ * worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, targetNBuffers);
+ /* Reset strategy control to new size */
+ StrategyReset(targetNBuffers);
+ }
+
+ elog(LOG, "Phase 1: Processing SHBUF_SHRINK barrier - NBuffers = %d, coordinator is %d",
+ NBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemResizeMapAndMem(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * If buffer pool is being shrunk, we are already working with a smaller
+ * buffer pool, so shrinking address space and shared structures should not
+ * be a problem. When expanding, expanding the address space and shared
+ * structures beyond the current boundaries is not going to be a problem
+ * since we are not accessing that memory yet. So there is no reason to
+ * delay processing this barrier.
+ */
+
+ /*
+ * Coordinator has already adjusted its address map and also updated sizes
+ * of the shared buffer structures, no further validation needed.
+ */
+ if (ShmemCtrl->coordinator == MyProcPid)
+ return true;
+
+ /*
+ * Backends validate that their pointers to shared buffer structures are
+ * still valid and have the correct size after memory remapping.
+ *
+ * TODO: Do want to do this only in assert enabled builds?
+ */
+ BufferManagerShmemValidate(targetNBuffers);
+
+ elog(LOG, "Backend %d successfully validated structure pointers after resize", MyProcPid);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemExpand(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * Delay adjusting the new active size of buffer pool till this process
+ * becomes ready to resize buffers.
+ */
+ if (delay_shmem_resize)
+ {
+ elog(LOG, "Phase 3: delaying SHBUF_EXPAND barrier - enabling allocations up to %d buffers, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+ return false;
+ }
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation statistics
+ * and the strategy control together so that background writer doesn't go
+ * out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody would
+ * reset the strategy control area. So we can't rely on background
+ * worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, targetNBuffers);
+ StrategyReset(targetNBuffers);
+ }
+
+ elog(LOG, "Phase 3: Processing SHBUF_EXPAND barrier - targetNBuffers = %d, ShmemCtrl->coordinator = %d", targetNBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemResizeFailed(void)
+{
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation statistics
+ * and the strategy control together so that background writer doesn't go
+ * out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody would
+ * reset the strategy control area. So we can't rely on background
+ * worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, currentNBuffers);
+ /* Reset strategy control to new size */
+ StrategyReset(currentNBuffers);
+ }
+
+ elog(LOG, "received proc signal indicating failure to resize shared buffers from %d to %d, restoring to %d, coordinator is %d",
+ NBuffers, targetNBuffers, currentNBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
\ No newline at end of file
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 67e87f9935d..18c9c6f336c 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -65,11 +65,18 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
+ /*
+ * The shared buffer look up table is set up only once with maximum possible
+ * entries considering maximum size of the buffer pool. It is not resized
+ * after that even if the buffer pool is resized. Hence it is allocated in
+ * the main shared memory segment and not in a resizeable shared memory
+ * segment.
+ */
SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE,
- STRATEGY_SHMEM_SEGMENT);
+ MAIN_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 327ddb7adc8..6c8f8552a4c 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/read_stream.h"
#include "storage/smgr.h"
@@ -3607,6 +3608,32 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_DONE(NBuffers, num_written, num_to_scan);
}
+/*
+ * Information saved between BgBufferSync() calls so we can determine the
+ * strategy point's advance rate and avoid scanning already-cleaned buffers. The
+ * variables are global instead of static local so that BgBufferSyncReset() can
+ * adjust it when resizing shared buffers.
+ */
+static bool saved_info_valid = false;
+static int prev_strategy_buf_id;
+static uint32 prev_strategy_passes;
+static int next_to_clean;
+static uint32 next_passes;
+
+/* Moving averages of allocation rate and clean-buffer density */
+static float smoothed_alloc = 0;
+static float smoothed_density = 10.0;
+
+void
+BgBufferSyncReset(int currentNBuffers, int targetNBuffers)
+{
+ saved_info_valid = false;
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
+ currentNBuffers, targetNBuffers);
+#endif
+}
+
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
@@ -3626,20 +3653,6 @@ BgBufferSync(WritebackContext *wb_context)
uint32 strategy_passes;
uint32 recent_alloc;
- /*
- * Information saved between calls so we can determine the strategy
- * point's advance rate and avoid scanning already-cleaned buffers.
- */
- static bool saved_info_valid = false;
- static int prev_strategy_buf_id;
- static uint32 prev_strategy_passes;
- static int next_to_clean;
- static uint32 next_passes;
-
- /* Moving averages of allocation rate and clean-buffer density */
- static float smoothed_alloc = 0;
- static float smoothed_density = 10.0;
-
/* Potentially these could be tunables, but for now, not */
float smoothing_samples = 16;
float scan_whole_pool_milliseconds = 120000.0;
@@ -3662,6 +3675,25 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ /*
+ * If buffer pool is being shrunk the buffer being written out may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing out any. Hence wait till
+ * buffer resizing finishes i.e. go into hibernation mode.
+ *
+ * TODO: We may not need this synchronization if background worker itself
+ * becomes the coordinator.
+ */
+ if (!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress))
+ return true;
+
+ /*
+ * Resizing shared buffers while this function is performing an LRU scan on
+ * them may lead to wrong results. Indicate that the resizing should wait for
+ * the LRU scan to complete.
+ */
+ delay_shmem_resize = true;
+
/*
* Find out where the clock-sweep currently is, and how many buffer
* allocations have happened since our last call.
@@ -3679,6 +3711,7 @@ BgBufferSync(WritebackContext *wb_context)
if (bgwriter_lru_maxpages <= 0)
{
saved_info_valid = false;
+ delay_shmem_resize = false;
return true;
}
@@ -3838,8 +3871,17 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
- while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
+ /*
+ * Execute the LRU scan.
+ *
+ * If buffer pool is being shrunk, the buffer being written may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing any. Hence stop what we are doing. This
+ * also unblocks other processes that are waiting for buffer resizing to
+ * finish.
+ */
+ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
+ !pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress))
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -3898,6 +3940,9 @@ BgBufferSync(WritebackContext *wb_context)
#endif
}
+ /* Let the resizing commence. */
+ delay_shmem_resize = false;
+
/* Return true if OK to hibernate */
return (bufs_to_lap == 0 && recent_alloc == 0);
}
@@ -4208,7 +4253,23 @@ DebugPrintBufferRefcount(Buffer buffer)
void
CheckPointBuffers(int flags)
{
+ /* Mark that buffer sync is in progress - delay any shared memory resizing. */
+ /*
+ * TODO: We need to assess whether we should allow checkpoint and buffer
+ * resizing to run in parallel. When expanding buffers it may be fine to let
+ * the checkpointer run in RESIZE_MAP_AND_MEM phase but delay phase EXPAND
+ * phase till the checkpoint finishes, at the same time not allow checkpoint
+ * to run during expansion phase. When shrinking the buffers, we should
+ * delay SHRINK phase till checkpoint finishes and not allow to start
+ * checkpoint till SHRINK phase is done, but allow it to run in
+ * RESIZE_MAP_AND_MEM phase. This needs careful analysis and testing.
+ */
+ delay_shmem_resize = true;
+
BufferSync(flags);
+
+ /* Mark that buffer sync is no longer in progress - allow shared memory resizing */
+ delay_shmem_resize = false;
}
/*
@@ -7466,3 +7527,70 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
+{
+ bool result = true;
+
+ Assert(targetNBuffers < currentNBuffers);
+
+ /*
+ * If the buffer being evicated is locked, this function will need to wait.
+ * This function should not be called from a Postmaster since it can not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = targetNBuffers + 1; buf <= currentNBuffers; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u32(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is
+ * going one hence unlocked precheck should be safe and saves
+ * some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ /*
+ * XXX: Looks like CurrentResourceOwner can be NULL here, find
+ * another one in that case?
+ * */
+ if (CurrentResourceOwner)
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 13ee840ab9f..256521d889a 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -33,10 +33,16 @@ typedef struct
/* Spinlock: protects the values below */
slock_t buffer_strategy_lock;
+ /*
+ * Number of active buffers that can be allocated. During buffer resizing,
+ * this may be different from NBuffers which tracks the global buffer count.
+ */
+ pg_atomic_uint32 activeNBuffers;
+
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo NBuffers.
+ * get an actual buffer, it needs to be used modulo activeNBuffers.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -101,21 +107,27 @@ static inline uint32
ClockSweepTick(void)
{
uint32 victim;
+ int activeBuffers;
/*
- * Atomically move hand ahead one buffer - if there's several processes
- * doing this, this can lead to buffers being returned slightly out of
- * apparent order.
+ * Atomically move hand ahead one buffer - if there's several processes doing
+ * this, this can lead to buffers being returned slightly out of apparent
+ * order. We need to read both the current position of hand and the current
+ * buffer allocation limit together consistently. They may be reset by
+ * concurrent resize.
*/
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
victim =
pg_atomic_fetch_add_u32(&StrategyControl->nextVictimBuffer, 1);
+ activeBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
- if (victim >= NBuffers)
+ if (victim >= activeBuffers)
{
uint32 originalVictim = victim;
/* always wrap what we look up in BufferDescriptors */
- victim = victim % NBuffers;
+ victim = victim % activeBuffers;
/*
* If we're the one that just caused a wraparound, force
@@ -143,7 +155,7 @@ ClockSweepTick(void)
*/
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
- wrapped = expected % NBuffers;
+ wrapped = expected % activeBuffers;
success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
&expected, wrapped);
@@ -228,7 +240,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
pg_atomic_fetch_add_u32(&StrategyControl->numBufferAllocs, 1);
/* Use the "clock sweep" algorithm to find a free buffer */
- trycounter = NBuffers;
+ trycounter = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+
for (;;)
{
uint32 old_buf_state;
@@ -281,7 +294,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
local_buf_state))
{
- trycounter = NBuffers;
+ trycounter = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
break;
}
}
@@ -323,10 +336,12 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
{
uint32 nextVictimBuffer;
int result;
+ uint32 activeNBuffers;
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
- result = nextVictimBuffer % NBuffers;
+ activeNBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+ result = nextVictimBuffer % activeNBuffers;
if (complete_passes)
{
@@ -336,7 +351,7 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
* Additionally add the number of wraparounds that happened before
* completePasses could be incremented. C.f. ClockSweepTick().
*/
- *complete_passes += nextVictimBuffer / NBuffers;
+ *complete_passes += nextVictimBuffer / activeNBuffers;
}
if (num_buf_alloc)
@@ -383,7 +398,7 @@ StrategyShmemSize(void)
Size size = 0;
/* size of lookup hash table ... see comment in StrategyInitialize */
- size = add_size(size, BufTableShmemSize(NBuffers + NUM_BUFFER_PARTITIONS));
+ size = add_size(size, BufTableShmemSize(MaxNBuffers + NUM_BUFFER_PARTITIONS));
/* size of the shared replacement strategy control block */
size = add_size(size, MAXALIGN(sizeof(BufferStrategyControl)));
@@ -391,6 +406,31 @@ StrategyShmemSize(void)
return size;
}
+void
+StrategyReset(int activeNBuffers)
+{
+ Assert(StrategyControl);
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ /* Update the active buffer count for the strategy */
+ pg_atomic_write_u32(&StrategyControl->activeNBuffers, activeNBuffers);
+
+ /* Reset the clock-sweep pointer to start from beginning */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The statistics is viewed in the context of the number of shared buffers.
+ * Reset it as the size of active number of shared buffers changes.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_write_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* TODO: Do we need to seset background writer notifications? */
+ StrategyControl->bgwprocno = -1;
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+}
+
/*
* StrategyInitialize -- initialize the buffer cache replacement
* strategy.
@@ -408,12 +448,21 @@ StrategyInitialize(bool init)
*
* Since we can't tolerate running out of lookup table entries, we must be
* sure to specify an adequate table size here. The maximum steady-state
- * usage is of course NBuffers entries, but BufferAlloc() tries to insert
- * a new entry before deleting the old. In principle this could be
- * happening in each partition concurrently, so we could need as many as
- * NBuffers + NUM_BUFFER_PARTITIONS entries.
+ * usage is of course is as many number of entries as the number of buffers
+ * in the buffer pool. Right now there is no way to free shared memory. Even
+ * if we shrink the buffer lookup table when shrinking the buffer pool the
+ * unused hash table entries can not be freed. When we expand the buffer
+ * pool, more entries can be allocated but we can not resize the hash table
+ * directory without rehashing all the entries. Just allocating more entries
+ * will lead to more contention. Hence we setup the buffer lookup table
+ * considering the maximum possible size of the buffer pool which is
+ * MaxNBuffers.
+ *
+ * Additionally BufferAlloc() tries to insert a new entry before deleting the
+ * old. In principle this could be happening in each partition concurrently,
+ * so we need extra NUM_BUFFER_PARTITIONS entries.
*/
- InitBufTable(NBuffers + NUM_BUFFER_PARTITIONS);
+ InitBufTable(MaxNBuffers + NUM_BUFFER_PARTITIONS);
/*
* Get or create the shared strategy control block
@@ -421,7 +470,7 @@ StrategyInitialize(bool init)
StrategyControl = (BufferStrategyControl *)
ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found, STRATEGY_SHMEM_SEGMENT);
+ &found, MAIN_SHMEM_SEGMENT);
if (!found)
{
@@ -432,6 +481,8 @@ StrategyInitialize(bool init)
SpinLockInit(&StrategyControl->buffer_strategy_lock);
+ /* Initialize the active buffer count */
+ pg_atomic_init_u32(&StrategyControl->activeNBuffers, NBuffersPending);
/* Initialize the clock-sweep pointer */
pg_atomic_init_u32(&StrategyControl->nextVictimBuffer, 0);
@@ -669,12 +720,23 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot has been
+ * invalidated when buffer pool was shrunk, tell the caller to allocate a new
+ * buffer with the normal allocation strategy. He will then fill this slot
+ * by calling AddBufferToRing with the new buffer.
+ *
+ * TODO: Ideally we would want to check for bufnum > NBuffers only once
+ * after every time the buffer pool is shrunk so as to catch any runtime
+ * bugs that introduce invalid buffers in the ring. But that is complicated.
+ * The BufferAccessStrategy objects are not accessible outside the
+ * ScanState. Hence we can not purge the buffers while evicting the buffers.
+ * After the resizing is finished, it's not possible to notice when we touch
+ * the first of those objects and the last of objects. See if this can
+ * fixed.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer ||
+ bufnum > pg_atomic_read_u32(&StrategyControl->activeNBuffers))
return NULL;
buf = GetBufferDescriptor(bufnum - 1);
diff --git a/src/backend/storage/buffer/meson.build b/src/backend/storage/buffer/meson.build
index 448976d2400..2fc58db5a91 100644
--- a/src/backend/storage/buffer/meson.build
+++ b/src/backend/storage/buffer/meson.build
@@ -6,4 +6,5 @@ backend_sources += files(
'bufmgr.c',
'freelist.c',
'localbuf.c',
+ 'buf_resize.c',
)
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 41190f96639..23e9b53ea07 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -154,6 +154,14 @@ CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
size = add_size(size, AioShmemSize());
size = add_size(size, WaitLSNShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -168,8 +176,7 @@ CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
/* might as well round it off to a multiple of a typical page size */
for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
{
- mapping_sizes[segment].shmem_req_size = add_size(mapping_sizes[segment].shmem_req_size, 8192 - (mapping_sizes[segment].shmem_req_size % 8192));
- mapping_sizes[segment].shmem_reserved = add_size(mapping_sizes[segment].shmem_reserved, 8192 - (mapping_sizes[segment].shmem_reserved % 8192));
+ round_off_mapping_sizes(&mapping_sizes[segment]);
/* Compute the total size of all segments */
size = size + mapping_sizes[segment].shmem_req_size;
}
@@ -313,6 +320,8 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
+ /* TODO: This should be part of BufferManagerShmemInit() */
+ ShmemControlInit();
BufferManagerShmemInit();
/*
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 087821311cc..c7c36f2be67 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -24,9 +24,11 @@
#include "port/pg_bitutils.h"
#include "replication/logicalworker.h"
#include "replication/walsender.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -109,6 +111,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -170,6 +176,43 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -576,6 +619,18 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_SMGRRELEASE:
processed = ProcessBarrierSmgrRelease();
break;
+ case PROCSIGNAL_BARRIER_SHBUF_SHRINK:
+ processed = ProcessBarrierShmemShrink();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM:
+ processed = ProcessBarrierShmemResizeMapAndMem();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_EXPAND:
+ processed = ProcessBarrierShmemExpand();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED:
+ processed = ProcessBarrierShmemResizeFailed();
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index f303a9328df..eafcb665ba9 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -69,11 +69,19 @@
#include "funcapi.h"
#include "miscadmin.h"
#include "port/pg_numa.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
#include "storage/shmem.h"
#include "storage/spin.h"
#include "utils/builtins.h"
+#include "utils/injection_point.h"
+#include "utils/wait_event.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
@@ -493,8 +501,7 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. The size better be the same as the size we are trying to
*/
if (result->size != size)
{
@@ -504,6 +511,7 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
" \"%s\": expected %zu, actual %zu",
name, size, result->size)));
}
+
structPtr = result->location;
}
else
@@ -538,6 +546,59 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
return structPtr;
}
+/*
+ * ShmemUpdateStructInSegment -- Update the size of a structure in shared memory.
+ *
+ * This function updates the size of an existing shared memory structure. It
+ * finds the structure in the shmem index and updates its size information while
+ * preserving the existing memory location.
+ *
+ * Returns: pointer to the existing structure location.
+ */
+void *
+ShmemUpdateStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
+{
+ ShmemIndexEnt *result;
+ void *structPtr;
+ Size delta;
+
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+
+ Assert(ShmemIndex);
+
+ /* Look up the structure in the shmem index */
+ result = (ShmemIndexEnt *)
+ hash_search(ShmemIndex, name, HASH_FIND, foundPtr);
+
+ Assert(*foundPtr);
+ Assert(result);
+ Assert(result->shmem_segment == shmem_segment);
+
+ delta = size - result->size;
+ /* Store the existing structure pointer */
+ structPtr = result->location;
+
+ /* Update the size information.
+ TODO: Ideally we should implement repalloc kind of functionality for shared memory which will return allocated size. */
+ result->size = size;
+ result->allocated_size = size;
+
+ /* Reflect size change in the shared segment */
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
+ Segments[shmem_segment].ShmemSegHdr->freeoffset += delta;
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
+ LWLockRelease(ShmemIndexLock);
+
+ /* Verify the structure is still in the correct segment */
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
+ Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
+
+ return structPtr;
+}
+
+
+
/*
* Add two Size values, checking for overflow
*/
@@ -871,4 +932,3 @@ pg_get_shmem_segments(PG_FUNCTION_ARGS)
return (Datum) 0;
}
-
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 2bd89102686..00c8afb9fe9 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -63,6 +63,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4138,6 +4139,9 @@ PostgresSingleUserMain(int argc, char *argv[],
/* Initialize size of fast-path lock cache. */
InitializeFastPathLocks();
+ /* Initialize MaxNBuffers for buffer pool resizing. */
+ InitializeMaxNBuffers();
+
/*
* Give preloaded libraries a chance to request additional shared memory.
*/
@@ -4328,6 +4332,13 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /*
+ * TODO: The new backend should fetch the shared buffers status. If the
+ * resizing is going on, it should bring itself upto speed with it. If not,
+ * simply fetch the latest pointers are sizes. Is this the right place to do
+ * that?
+ */
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index c1ac71ff7f2..ee5887496ba 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -162,6 +162,7 @@ WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
WAL_RECEIVER_WAIT_START "Waiting for startup process to send initial data for streaming replication."
WAL_SUMMARY_READY "Waiting for a new WAL summary to be generated."
XACT_GROUP_UPDATE "Waiting for the group leader to update transaction status at transaction end."
+PM_BUFFER_RESIZE_WAIT "Waiting for the postmaster to complete shared buffer pool resize operations."
ABI_compatibility:
@@ -358,6 +359,7 @@ InjectionPoint "Waiting to read or update information related to injection point
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
WaitLSN "Waiting to read or update shared Wait-for-LSN state."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index d31cb45a058..419c7fad890 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -139,7 +139,10 @@ int max_parallel_maintenance_workers = 2;
* MaxBackends is computed by PostmasterMain after modules have had a chance to
* register background workers.
*/
-int NBuffers = 16384;
+int NBuffers = 0;
+int NBuffersPending = 16384;
+bool finalMaxNBuffers = false;
+int MaxNBuffers = 0;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 98f9598cd78..46a8a8a3faa 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -595,6 +595,55 @@ InitializeFastPathLocks(void)
pg_nextpower2_32(FastPathLockGroupsPerBackend));
}
+/*
+ * Initialize MaxNBuffers variable with validation.
+ *
+ * This must be called after GUCs have been loaded but before shared memory size
+ * is determined.
+ *
+ * Since MaxNBuffers limits the size of the buffer pool, it must be at least as
+ * much as NBuffersPending. If MaxNBuffers is 0 (default), set it to
+ * NBuffersPending. Otherwise, validate that MaxNBuffers is not less than
+ * NBuffersPending.
+ */
+void
+InitializeMaxNBuffers(void)
+{
+ if (MaxNBuffers == 0) /* default/boot value */
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", NBuffersPending);
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+
+ /*
+ * We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
+ * However, if the DBA explicitly set max_shared_buffers = 0 in
+ * the config file, then PGC_S_DYNAMIC_DEFAULT will fail to override
+ * that and we must force the matter with PGC_S_OVERRIDE.
+ */
+ if (MaxNBuffers == 0) /* failed to apply it? */
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+ else
+ {
+ if (MaxNBuffers < NBuffersPending)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("max_shared_buffers (%d) cannot be less than current shared_buffers (%d)",
+ MaxNBuffers, NBuffersPending),
+ errhint("Increase max_shared_buffers or decrease shared_buffers.")));
+ }
+ }
+
+ Assert(MaxNBuffers > 0);
+ Assert(!finalMaxNBuffers);
+ finalMaxNBuffers = true;
+}
+
/*
* Early initialization of a backend (either standalone or under postmaster).
* This happens even before InitPostgres.
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 7e2b17cc04e..c26a02e4a42 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -2599,7 +2599,7 @@ convert_to_base_unit(double value, const char *unit,
* the value without loss. For example, if the base unit is GUC_UNIT_KB, 1024
* is converted to 1 MB, but 1025 is represented as 1025 kB.
*/
-static void
+void
convert_int_from_base_unit(int64 base_value, int base_unit,
int64 *value, const char **unit)
{
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index 1128167c025..539b29f0065 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2013,6 +2013,15 @@
max => 'MAX_BACKENDS /* XXX? */',
},
+{ name => "max_shared_buffers", type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+ short_desc => 'Sets the upper limit for the shared_buffers value.',
+ flags => 'GUC_UNIT_BLOCKS',
+ variable => 'MaxNBuffers',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX / 2',
+},
+
{ name => 'max_slot_wal_keep_size', type => 'int', context => 'PGC_SIGHUP', group => 'REPLICATION_SENDING',
short_desc => 'Sets the maximum WAL size that can be reserved by replication slots.',
long_desc => 'Replication slots will be marked as failed, and segments released for deletion or recycling, if this much space is occupied by WAL on disk. -1 means no maximum.',
@@ -2581,13 +2590,15 @@
# We sometimes multiply the number of shared buffers by two without
# checking for overflow, so we mustn't allow more than INT_MAX / 2.
-{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
- variable => 'NBuffers',
+ variable => 'NBuffersPending',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
+ check_hook => 'check_shared_buffers',
+ show_hook => 'show_shared_buffers',
},
{ name => 'shared_memory_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 411043ca750..119d9dd5880 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12612,4 +12612,10 @@
proargnames => '{pid,io_id,io_generation,state,operation,off,length,target,handle_data_len,raw_result,result,target_desc,f_sync,f_localmem,f_buffered}',
prosrc => 'pg_get_aios' },
+{ oid => '9999', descr => 'resize shared buffers according to the value of GUC `shared_buffers`',
+ proname => 'pg_resize_shared_buffers',
+ provolatile => 'v',
+ prorettype => 'bool',
+ proargtypes => '',
+ prosrc => 'pg_resize_shared_buffers'},
]
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 9a7d733ddef..b4dc2c4ba57 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,7 +173,11 @@ extern PGDLLIMPORT bool ExitOnAnyError;
extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
+/* TODO: This is no more a GUC variable; should be moved somewhere else. */
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersPending;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
@@ -502,6 +506,7 @@ extern PGDLLIMPORT ProcessingMode Mode;
extern void pg_split_opts(char **argv, int *argcp, const char *optstr);
extern void InitializeMaxBackends(void);
extern void InitializeFastPathLocks(void);
+extern void InitializeMaxNBuffers(void);
extern void InitPostgres(const char *in_dbname, Oid dboid,
const char *username, Oid useroid,
bits32 flags,
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 519692702a0..4c53194e13e 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -513,6 +513,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyReset(int activeNBuffers);
/* buf_table.c */
extern Size BufTableShmemSize(int size);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 3769f4db7dc..774cf8f38ed 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -21,6 +21,7 @@
#include "storage/bufpage.h"
#include "storage/pg_shmem.h"
#include "storage/relfilelocator.h"
+#include "utils/guc.h"
#include "utils/relcache.h"
#include "utils/snapmgr.h"
@@ -159,6 +160,7 @@ typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersPending;
/* in bufmgr.c */
extern PGDLLIMPORT bool zero_damaged_pages;
@@ -205,6 +207,11 @@ extern PGDLLIMPORT int32 *LocalRefCount;
#define BUFFER_LOCK_SHARE 1
#define BUFFER_LOCK_EXCLUSIVE 2
+/*
+ * prototypes for functions in buf_init.c
+ */
+extern const char *show_shared_buffers(void);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
/*
* prototypes for functions in bufmgr.c
@@ -308,6 +315,7 @@ extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
extern bool BgBufferSync(WritebackContext *wb_context);
+extern void BgBufferSyncReset(int currentNBuffers, int targetNBuffers);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
@@ -324,10 +332,13 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
int32 *buffers_evicted,
int32 *buffers_flushed,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int targetNBuffers, int currentNBuffers);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
extern Size BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes);
+extern void BufferManagerShmemResize(int currentNBuffers, int targetNBuffers);
+extern void BufferManagerShmemValidate(int targetNBuffers);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
@@ -376,7 +387,7 @@ extern void FreeAccessStrategy(BufferAccessStrategy strategy);
static inline bool
BufferIsValid(Buffer bufnum)
{
- Assert(bufnum <= NBuffers);
+ Assert(bufnum <= (Buffer) pg_atomic_read_u32(&ShmemCtrl->currentNBuffers));
Assert(bufnum >= -NLocBuffer);
return bufnum != InvalidBuffer;
@@ -430,4 +441,11 @@ BufferGetPage(Buffer buffer)
#endif /* FRONTEND */
+/* buf_resize.c */
+extern Datum pg_resize_shared_buffers(PG_FUNCTION_ARGS);
+extern bool ProcessBarrierShmemShrink(void);
+extern bool ProcessBarrierShmemResizeMapAndMem(void);
+extern bool ProcessBarrierShmemExpand(void);
+extern bool ProcessBarrierShmemResizeFailed(void);
+
#endif /* BUFMGR_H */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index d73f1b407db..6dbbb9ad064 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -66,6 +66,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
@@ -85,5 +86,7 @@ extern void CreateSharedMemoryAndSemaphores(void);
extern void AttachSharedMemoryStructs(void);
#endif
extern void InitializeShmemGUCs(void);
+extern void CoordinateShmemResize(void);
+extern bool AnonymousShmemResize(void);
#endif /* IPC_H */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index 5b0ce383408..9c4b928441c 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -86,6 +86,7 @@ PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
PG_LWLOCK(54, WaitLSN)
+PG_LWLOCK(55, ShmemResize)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index beee0a53d2d..36900068820 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,8 +24,13 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "port/atomics.h"
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
+#include "storage/procsignal.h"
#include "storage/spin.h"
+#include "storage/shmem.h"
+#include "utils/guc.h"
typedef struct MemoryMappingSizes
{
@@ -65,15 +70,39 @@ typedef struct ShmemSegment
} ShmemSegment;
/* Number of available segments for anonymous memory mappings */
-#define NUM_MEMORY_MAPPINGS 6
+#define NUM_MEMORY_MAPPINGS 5
extern PGDLLIMPORT ShmemSegment Segments[NUM_MEMORY_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ *
+ * TODO: I think we need a lock to protect this structure. If we do so, do we
+ * need to use atomic integers?
+ */
+typedef struct
+{
+ pg_atomic_flag resize_in_progress; /* true if resizing is in progress. false otherwise. */
+ pg_atomic_uint32 currentNBuffers; /* Original NBuffers value before resize started */
+ pg_atomic_uint32 targetNBuffers;
+ pid_t coordinator;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -113,6 +142,17 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
+/*
+ * round off mapping size to a multiple of a typical page size.
+ */
+static inline void
+round_off_mapping_sizes(MemoryMappingSizes *mapping_sizes)
+{
+ mapping_sizes->shmem_req_size = add_size(mapping_sizes->shmem_req_size, 8192 - (mapping_sizes->shmem_req_size % 8192));
+ mapping_sizes->shmem_reserved = add_size(mapping_sizes->shmem_reserved, 8192 - (mapping_sizes->shmem_reserved % 8192));
+}
+
+
extern PGShmemHeader *PGSharedMemoryCreate(MemoryMappingSizes *mapping_sizes, int segment_id,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
@@ -122,6 +162,13 @@ extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
void PrepareHugePages(void);
+bool ProcessBarrierShmemResize(Barrier *barrier);
+const char *show_shared_buffers(void);
+bool check_shared_buffers(int *newval, void **extra, GucSource source);
+void AdjustShmemSize(void);
+extern void WaitOnShmemBarrier(void);
+extern void ShmemControlInit(void);
+
/*
* To be able to dynamically resize largest parts of the data stored in shared
* memory, we split it into multiple shared memory mappings segments. Each
@@ -144,7 +191,4 @@ void PrepareHugePages(void);
/* Checkpoint BufferIds */
#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
-/* Buffer strategy status */
-#define STRATEGY_SHMEM_SEGMENT 5
-
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 428aa3fd68a..5ced2a83537 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -42,9 +42,10 @@ typedef enum
PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */
PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */
PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */
+ PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */
} PMSignalReason;
-#define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1)
+#define NUM_PMSIGNALS (PMSIGNAL_SHMEM_RESIZE+1)
/*
* Reasons why the postmaster would send SIGQUIT to its children.
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index afeeb1ca019..4de11faf12d 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,10 @@ typedef enum
typedef enum
{
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
+ PROCSIGNAL_BARRIER_SHBUF_SHRINK, /* shrink buffer pool - restrict allocations to new size */
+ PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM, /* remap shared memory segments and update structure pointers */
+ PROCSIGNAL_BARRIER_SHBUF_EXPAND, /* expand buffer pool - enable allocations in new range */
+ PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED, /* signal backends that the shared buffer resizing failed. */
} ProcSignalBarrierType;
/*
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index c56712555f0..d59e5ba6dcd 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -49,11 +49,14 @@ extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
extern void *ShmemInitStructInSegment(const char *name, Size size,
bool *foundPtr, int shmem_segment);
+extern void *ShmemUpdateStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
extern PGDLLIMPORT Size pg_get_shmem_pagesize(void);
+
/* ipci.c */
extern void RequestAddinShmemSpace(Size size);
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index f21ec37da89..08a84373fb7 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -459,6 +459,8 @@ extern config_handle *get_config_handle(const char *name);
extern void AlterSystemSetConfigFile(AlterSystemStmt *altersysstmt);
extern char *GetConfigOptionByName(const char *name, const char **varname,
bool missing_ok);
+extern void convert_int_from_base_unit(int64 base_value, int base_unit,
+ int64 *value, const char **unit);
extern void TransformGUCArray(ArrayType *array, List **names,
List **values);
diff --git a/src/test/Makefile b/src/test/Makefile
index 511a72e6238..95f8858a818 100644
--- a/src/test/Makefile
+++ b/src/test/Makefile
@@ -12,7 +12,7 @@ subdir = src/test
top_builddir = ../..
include $(top_builddir)/src/Makefile.global
-SUBDIRS = perl postmaster regress isolation modules authentication recovery subscription
+SUBDIRS = perl postmaster regress isolation modules authentication recovery subscription buffermgr
ifeq ($(with_icu),yes)
SUBDIRS += icu
diff --git a/src/test/README b/src/test/README
index afdc7676519..77f11607ff7 100644
--- a/src/test/README
+++ b/src/test/README
@@ -15,6 +15,9 @@ examples/
Demonstration programs for libpq that double as regression tests via
"make check"
+buffermgr/
+ Tests for resizing buffer pool without restarting the server
+
isolation/
Tests for concurrent behavior at the SQL level
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
new file mode 100644
index 00000000000..eb275027fa6
--- /dev/null
+++ b/src/test/buffermgr/Makefile
@@ -0,0 +1,30 @@
+#-------------------------------------------------------------------------
+#
+# Makefile for src/test/buffermgr
+#
+# Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+# Portions Copyright (c) 1994, Regents of the University of California
+#
+# src/test/buffermgr/Makefile
+#
+#-------------------------------------------------------------------------
+
+EXTRA_INSTALL = contrib/pg_buffercache
+
+REGRESS = buffer_resize
+
+# Custom configuration for buffer manager tests
+TEMP_CONFIG = $(srcdir)/buffermgr_test.conf
+
+subdir = src/test/buffermgr
+top_builddir = ../../..
+include $(top_builddir)/src/Makefile.global
+
+check:
+ $(prove_check)
+
+installcheck:
+ $(prove_installcheck)
+
+clean distclean:
+ rm -rf tmp_check
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
new file mode 100644
index 00000000000..c375ad80989
--- /dev/null
+++ b/src/test/buffermgr/README
@@ -0,0 +1,26 @@
+src/test/buffermgr/README
+
+Regression tests for buffer manager
+===================================
+
+This directory contains a test suite for resizing buffer manager without restarting the server.
+
+
+Running the tests
+=================
+
+NOTE: You must have given the --enable-tap-tests argument to configure.
+
+Run
+ make check
+or
+ make installcheck
+You can use "make installcheck" if you previously did "make install".
+In that case, the code in the installation tree is tested. With
+"make check", a temporary installation tree is built from the current
+sources and then tested.
+
+Either way, this test initializes, starts, and stops a test Postgres
+cluster.
+
+See src/test/perl/README for more info about running these tests.
diff --git a/src/test/buffermgr/buffermgr_test.conf b/src/test/buffermgr/buffermgr_test.conf
new file mode 100644
index 00000000000..b7c0065c80b
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.conf
@@ -0,0 +1,11 @@
+# Configuration for buffer manager regression tests
+
+# Even if max_shared_buffers is set multiple times only the last one is used to
+# as the limit on shared_buffers.
+max_shared_buffers = 128kB
+# Set initial shared_buffers as expected by test
+shared_buffers = 128MB
+# Set a larger value for max_shared_buffers to allow testing resize operations
+max_shared_buffers = 300MB
+# Turn huge pages off, since that affects the size of memory segments
+huge_pages = off
\ No newline at end of file
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
new file mode 100644
index 00000000000..d5cb9d78437
--- /dev/null
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -0,0 +1,329 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+-- TODO: The test sets shared_buffers values in MBs. Instead it could use values
+-- in kBs so that the test runs on very small machines.
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, mapping_size, mapping_reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | descriptors | 1048576 | 1048576
+ Buffer IO Condition Variables | iocv | 262144 | 262144
+ Checkpoint BufferIds | checkpoint | 327680 | 327680
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 134225920 | 134225920 | 314580992
+ checkpoint | 335872 | 335872 | 770048
+ descriptors | 1056768 | 1056768 | 2465792
+ iocv | 270336 | 270336 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | descriptors | 1048576 | 1048576
+ Buffer IO Condition Variables | iocv | 262144 | 262144
+ Checkpoint BufferIds | checkpoint | 327680 | 327680
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 134225920 | 134225920 | 314580992
+ checkpoint | 335872 | 335872 | 770048
+ descriptors | 1056768 | 1056768 | 2465792
+ iocv | 270336 | 270336 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 128MB (pending: 64MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 64MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+----------+----------------
+ Buffer Blocks | buffers | 67112960 | 67112960
+ Buffer Descriptors | descriptors | 524288 | 524288
+ Buffer IO Condition Variables | iocv | 131072 | 131072
+ Checkpoint BufferIds | checkpoint | 163840 | 163840
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+----------+--------------+-----------------------
+ buffers | 67117056 | 67117056 | 314580992
+ checkpoint | 172032 | 172032 | 770048
+ descriptors | 532480 | 532480 | 2465792
+ iocv | 139264 | 139264 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 8192
+(1 row)
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 64MB (pending: 256MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 268439552 | 268439552
+ Buffer Descriptors | descriptors | 2097152 | 2097152
+ Buffer IO Condition Variables | iocv | 524288 | 524288
+ Checkpoint BufferIds | checkpoint | 655360 | 655360
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 268443648 | 268443648 | 314580992
+ checkpoint | 663552 | 663552 | 770048
+ descriptors | 2105344 | 2105344 | 2465792
+ iocv | 532480 | 532480 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 32768
+(1 row)
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 256MB (pending: 100MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 100MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 104861696 | 104861696
+ Buffer Descriptors | descriptors | 819200 | 819200
+ Buffer IO Condition Variables | iocv | 204800 | 204800
+ Checkpoint BufferIds | checkpoint | 256000 | 256000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+-----------+--------------+-----------------------
+ buffers | 104865792 | 104865792 | 314580992
+ checkpoint | 262144 | 262144 | 770048
+ descriptors | 827392 | 827392 | 2465792
+ iocv | 212992 | 212992 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 12800
+(1 row)
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 100MB (pending: 128kB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+--------+----------------
+ Buffer Blocks | buffers | 135168 | 135168
+ Buffer Descriptors | descriptors | 1024 | 1024
+ Buffer IO Condition Variables | iocv | 256 | 256
+ Checkpoint BufferIds | checkpoint | 320 | 320
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | mapping_size | mapping_reserved_size
+-------------+--------+--------------+-----------------------
+ buffers | 139264 | 139264 | 314580992
+ checkpoint | 8192 | 8192 | 770048
+ descriptors | 8192 | 8192 | 2465792
+ iocv | 8192 | 8192 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16
+(1 row)
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+ERROR: invalid value for parameter "shared_buffers": 51200
+DETAIL: "shared_buffers" must be less than "max_shared_buffers".
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
new file mode 100644
index 00000000000..c24bff721e6
--- /dev/null
+++ b/src/test/buffermgr/meson.build
@@ -0,0 +1,23 @@
+# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+
+tests += {
+ 'name': 'buffermgr',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'buffer_resize',
+ ],
+ 'regress_args': ['--temp-config', files('buffermgr_test.conf')],
+ },
+ 'tap': {
+ 'env': {
+ 'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
+ },
+ 'tests': [
+ 't/001_resize_buffer.pl',
+ 't/003_parallel_resize_buffer.pl',
+ 't/004_client_join_buffer_resize.pl',
+ ],
+ },
+}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
new file mode 100644
index 00000000000..dfaaeabfcbb
--- /dev/null
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -0,0 +1,95 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+-- TODO: The test sets shared_buffers values in MBs. Instead it could use values
+-- in kBs so that the test runs on very small machines.
+
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, mapping_size, mapping_reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+SHOW max_shared_buffers;
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
new file mode 100644
index 00000000000..a0d7f094171
--- /dev/null
+++ b/src/test/buffermgr/t/001_resize_buffer.pl
@@ -0,0 +1,135 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Minimal test testing shared_buffer resizing under load
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Function to resize buffer pool and verify the change.
+sub apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ # Use the new pg_resize_shared_buffers() interface which handles everything synchronously
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ # If resize function fails, try a few times before giving up
+ my $max_retries = 5;
+ my $retry_delay = 1; # seconds
+ my $success = 0;
+ for my $attempt (1..$max_retries) {
+ my $result = $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()");
+ if ($result eq 't') {
+ $success = 1;
+ last;
+ }
+
+ # If not the last attempt, wait before retrying
+ if ($attempt < $max_retries) {
+ note "Resizing buffer pool to $new_size, attempt $attempt failed, retrying after $retry_delay seconds...";
+ sleep($retry_delay);
+ }
+ }
+
+ is($success, 1, 'resizing to ' . $new_size . ' succeeded after retries');
+ is($node->safe_psql('postgres', "SHOW shared_buffers"), $new_size,
+ 'SHOW after resizing to '. $new_size . ' succeeded');
+}
+
+# Initialize a cluster and start pgbench in the background for concurrent load.
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Permit resizing up to 1GB for this test and let the server start with 128MB.
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 1GB
+shared_buffers = 128MB
+log_statement = none
+});
+
+$node->start;
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+my $pgb_scale = 10;
+my $pgb_duration = 120;
+my $pgb_num_clients = 10;
+$node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$pgb_scale --quiet",
+ 0,
+ [qr{^$}],
+ [ # stderr patterns to verify initialization stages
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$pgb_scale)"
+);
+my ($pgbench_stdin, $pgbench_stdout, $pgbench_stderr) = ('', '', '');
+# Use --exit-on-abort so that the test stops on the first server crash or error,
+# thus making it easy to debug the failure. Use -C to increase the chances of a
+# new backend being created while resizing the buffer pool.
+my $pgbench_process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-T', $pgb_duration,
+ '-c', $pgb_num_clients,
+ '-C',
+ '--exit-on-abort',
+ 'postgres'
+ ],
+ '<' => \$pgbench_stdin,
+ '>' => \$pgbench_stdout,
+ '2>' => \$pgbench_stderr
+);
+
+ok($pgbench_process, "pgbench started successfully");
+
+# Allow pgbench to establish connections and start generating load.
+#
+# TODO: When creating new backends is known to work well with buffer pool
+# resizing, this wait should be removed.
+sleep(1);
+
+# Resize buffer pool to various sizes while pgbench is running in the
+# background.
+#
+# TODO: These are pseudo-randomly picked sizes, but we can do better.
+my $tests_completed = 0;
+my @buffer_sizes = ('900MB', '500MB', '250MB', '400MB', '120MB', '600MB');
+for my $target_size (@buffer_sizes)
+{
+ # Verify workload generator is still running
+ if (!$pgbench_process->pumpable) {
+ ok(0, "pgbench is still running");
+ last;
+ }
+
+ apply_and_verify_buffer_change($node, $target_size);
+ $tests_completed++;
+
+ # Wait for the resized buffer pool to stabilize. If the resized buffer pool
+ # is utilized fully, it might hit any wrongly initialized areas of shared
+ # memory.
+ sleep(2);
+}
+is($tests_completed, scalar(@buffer_sizes), "All buffer sizes were tested");
+
+# Make sure that pgbench can end normally.
+$pgbench_process->signal('TERM');
+IPC::Run::finish $pgbench_process;
+ok(grep { $pgbench_process->result == $_ } (0, 15), "pgbench exited gracefully");
+
+# Log any error output from pgbench for debugging
+diag("pgbench stderr:\n$pgbench_stderr");
+diag("pgbench stdout:\n$pgbench_stdout");
+
+# Ensure database is still functional after all the buffer changes
+$node->connect_ok("dbname=postgres",
+ "Database remains accessible after $tests_completed buffer resize operations");
+
+done_testing();
\ No newline at end of file
diff --git a/src/test/buffermgr/t/003_parallel_resize_buffer.pl b/src/test/buffermgr/t/003_parallel_resize_buffer.pl
new file mode 100644
index 00000000000..9cbb5452fd2
--- /dev/null
+++ b/src/test/buffermgr/t/003_parallel_resize_buffer.pl
@@ -0,0 +1,71 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test that only one pg_resize_shared_buffers() call succeeds when multiple
+# sessions attempt to resize buffers concurrently
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Initialize a cluster
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', 'shared_buffers = 128kB');
+$node->append_conf('postgresql.conf', 'max_shared_buffers = 256kB');
+$node->start;
+
+# Load injection points extension for test coordination
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Test 1: Two concurrent pg_resize_shared_buffers() calls
+# Set up injection point to pause the first resize call
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('pg-resize-shared-buffers-flag-set', 'wait')");
+
+# Change shared_buffers for the resize operation
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '144kB'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+# Start first resize session (will pause at injection point)
+my $session1 = $node->background_psql('postgres');
+$session1->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+);
+
+# Wait until session actually reaches the injection point
+$node->wait_for_event('client backend', 'pg-resize-shared-buffers-flag-set');
+
+# Start second resize session (should fail immediately since resize is in progress)
+my $result2 = $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()");
+
+# The second call should return false (already in progress)
+is($result2, 'f', 'Second concurrent resize call returns false');
+
+# Wake up the first session
+$node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('pg-resize-shared-buffers-flag-set')");
+
+# The pg_resize_shared_buffers() in session1 should now complete successfully
+# We can't easily capture the return value from query_until, but we can
+# verify the session completes without error and the resize actually happened
+$session1->quit;
+
+# Detach injection point
+$node->safe_psql('postgres',
+ "SELECT injection_points_detach('pg-resize-shared-buffers-flag-set')");
+
+done_testing();
\ No newline at end of file
diff --git a/src/test/buffermgr/t/004_client_join_buffer_resize.pl b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
new file mode 100644
index 00000000000..06f0de6b409
--- /dev/null
+++ b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
@@ -0,0 +1,241 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test shared_buffer resizing coordination with client connections joining using injection points
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+use Time::HiRes qw(sleep);
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Function to calculate the size of test table required to fill up maximum
+# buffer pool when populating it.
+sub calculate_test_sizes
+{
+ my ($node, $block_size) = @_;
+
+ # Get the maximum buffer pool size from configuration
+ my $max_shared_buffers = $node->safe_psql('postgres', "SHOW max_shared_buffers");
+ my ($max_val, $max_unit) = ($max_shared_buffers =~ /(\d+)(\w+)/);
+ my $max_size_bytes;
+ if (lc($max_unit) eq 'kb') {
+ $max_size_bytes = $max_val * 1024;
+ } elsif (lc($max_unit) eq 'mb') {
+ $max_size_bytes = $max_val * 1024 * 1024;
+ } elsif (lc($max_unit) eq 'gb') {
+ $max_size_bytes = $max_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $max_size_bytes = $max_val * 1024;
+ }
+
+ # Fill more pages than minimally required to increase the chances of pages
+ # from the test table filling the buffer cache.
+ $max_size_bytes = $max_size_bytes;
+ my $pages_needed = int($max_size_bytes / $block_size) + 10; # Add some extra to ensure buffers are filled
+ my $rows_to_insert = $pages_needed * 100; # Assuming roughly 100 rows per page for our table structure
+
+ return ($max_size_bytes, $pages_needed, $rows_to_insert);
+}
+
+# Function to calculate expected buffer count from size string
+sub calculate_buffer_count
+{
+ my ($size_string, $block_size) = @_;
+
+ # Parse size and convert to bytes
+ my ($size_val, $unit) = ($size_string =~ /(\d+)(\w+)/);
+ my $size_bytes;
+ if (lc($unit) eq 'kb') {
+ $size_bytes = $size_val * 1024;
+ } elsif (lc($unit) eq 'mb') {
+ $size_bytes = $size_val * 1024 * 1024;
+ } elsif (lc($unit) eq 'gb') {
+ $size_bytes = $size_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $size_bytes = $size_val * 1024;
+ }
+
+ return int($size_bytes / $block_size);
+}
+
+# Initialize cluster with very small buffer sizes for testing
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Configure for buffer resizing with very small buffer pool sizes for faster tests.
+# TODO: for some reason parallel workers try to load default number of shared_buffers which doesn't work with lower max_shared_buffers. We need to fix that - somewhere it's picking default value of shared buffers. For now disable parallelism
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 512kB
+shared_buffers = 320kB
+max_parallel_workers_per_gather = 0
+});
+
+$node->start;
+
+# Enable injection points
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Get the block size (this is fixed for the binary)
+my $block_size = $node->safe_psql('postgres', "SHOW block_size");
+
+# Try to create pg_buffercache extension for buffer analysis
+eval {
+ $node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+};
+if ($@) {
+ $node->stop;
+ plan skip_all => 'pg_buffercache extension not available - cannot verify buffer usage';
+}
+
+# Create a small test table, and fetch its properties for later reference if required.
+$node->safe_psql('postgres', qq{
+ CREATE TABLE client_test (c1 int, data char(50));
+});
+
+my $table_oid = $node->safe_psql('postgres', "SELECT oid FROM pg_class WHERE relname = 'client_test'");
+my $table_relfilenode = $node->safe_psql('postgres', "SELECT relfilenode FROM pg_class WHERE relname = 'client_test'");
+note("Test table client_test: OID = $table_oid, relfilenode = $table_relfilenode");
+my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $block_size);
+
+# Create dedicated sessions for injection point handling and test queries,
+# so that we don't create new backends for test operations after starting
+# resize operation. Only one backend, which tests new backend synchronization
+# with resizing operation, should start after resizing has commenced.
+my $injection_session = $node->background_psql('postgres');
+my $query_session = $node->background_psql('postgres');
+my $resize_session = $node->background_psql('postgres');
+
+# Function to run a single injection point test
+sub run_injection_point_test
+{
+ my ($test_name, $injection_point, $target_size, $operation_type) = @_;
+
+ note("Test with $test_name ($operation_type)");
+
+ # Calculate test parameters before starting resize
+ my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $target_size, $block_size);
+
+ # Update buffer pool size and wait for it to reflect pending state
+ $resize_session->query_safe("ALTER SYSTEM SET shared_buffers = '$target_size'");
+ $resize_session->query_safe("SELECT pg_reload_conf()");
+ my $pending_size_str = "pending: $target_size";
+ $resize_session->poll_query_until("SELECT substring(current_setting('shared_buffers'), '$pending_size_str')", $pending_size_str);
+
+ # Set up injection point in injection session
+ $injection_session->query_safe("SELECT injection_points_attach('$injection_point', 'wait')");
+
+ # Trigger resize
+ $resize_session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+ );
+
+ # Wait until resize actually reaches the injection point using the query session
+ $query_session->wait_for_event('client backend', $injection_point);
+
+ # Start a client while resize is paused
+ my $client = $node->background_psql('postgres');
+ note("Background client backend PID: " . $client->query_safe("SELECT pg_backend_pid()"));
+
+ # Wake up the injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_wakeup('$injection_point')");
+
+ # Test buffer functionality immediately after waking up injection point
+ # Insert data to test buffer pool functionality during/after resize
+ $client->query_safe("INSERT INTO client_test SELECT i, 'test_data_' || i FROM generate_series(1, $rows_to_insert) i");
+ # Verify the data was inserted correctly and can be read back
+ is($client->query_safe("SELECT COUNT(*) FROM client_test"), $rows_to_insert, "inserted $rows_to_insert during $test_name ($operation_type) successful");
+
+ # Verify table size is reasonable (should be substantial for testing)
+ ok($query_session->query_safe("SELECT pg_total_relation_size('client_test')") >= $max_size_bytes,"table size is large enough to overflow buffer pool in test $test_name ($operation_type)");
+
+ # Wait for the resize operation to complete. There is no direct way to do so
+ # in background_psql. Hence fire a psql command and wait for it to finish
+ $resize_session->query(q(\echo 'done'));
+
+ # Detach injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_detach('$injection_point')");
+
+ # Verify resize completed successfully
+ is($query_session->query_safe("SELECT current_setting('shared_buffers')"), $target_size,
+ "resize completed successfully to $target_size");
+
+ # Check buffer pool size using pg_buffercache after resize completion
+ is($query_session->query_safe("SELECT COUNT(*) FROM pg_buffercache"), calculate_buffer_count($target_size, $block_size), "all buffers in the buffer pool used in $test_name ($operation_type)");
+
+ # Wait for client to complete
+ ok($client->quit, "client succeeded during $test_name ($operation_type)");
+
+ # Clean up for next test
+ $query_session->query_safe("DELETE FROM client_test");
+}
+
+# Test injection points during buffer resize with client connections
+my @common_injection_tests = (
+ {
+ name => 'flag setting phase',
+ injection_point => 'pg-resize-shared-buffers-flag-set',
+ },
+ {
+ name => 'memory remap phase',
+ injection_point => 'pgrsb-after-shmem-resize',
+ },
+ {
+ name => 'resize map barrier complete',
+ injection_point => 'pgrsb-resize-barrier-sent',
+ },
+);
+
+# Test common injection points for both shrinking and expanding
+foreach my $test (@common_injection_tests)
+{
+ # Test shrinking scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '272kB', 'shrinking');
+
+ # Test expanding scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '400kB', 'expanding');
+}
+
+my @shrink_only_tests = (
+ {
+ name => 'shrink barrier complete',
+ injection_point => 'pgrsb-shrink-barrier-sent',
+ size => '200kB',
+ }
+);
+foreach my $test (@shrink_only_tests)
+{
+ run_injection_point_test($test->{name}, $test->{injection_point}, $test->{size}, 'shrinking only');
+}
+
+my @expand_only_tests = (
+ {
+ name => 'expand barrier complete',
+ injection_point => 'pgrsb-expand-barrier-sent',
+ size => '416kB',
+ }
+);
+foreach my $test (@expand_only_tests)
+{
+ run_injection_point_test($test->{name}, $test->{injection_point}, $test->{size}, 'expanding only');
+}
+
+$injection_session->quit;
+$query_session->quit;
+$resize_session->quit;
+
+done_testing();
\ No newline at end of file
diff --git a/src/test/meson.build b/src/test/meson.build
index ccc31d6a86a..2a5ba1dec39 100644
--- a/src/test/meson.build
+++ b/src/test/meson.build
@@ -4,6 +4,7 @@ subdir('regress')
subdir('isolation')
subdir('authentication')
+subdir('buffermgr')
subdir('postmaster')
subdir('recovery')
subdir('subscription')
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index 60bbd5dd445..16625e94d92 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -61,6 +61,7 @@ use Config;
use IPC::Run;
use PostgreSQL::Test::Utils qw(pump_until);
use Test::More;
+use Time::HiRes qw(usleep);
=pod
@@ -371,4 +372,79 @@ sub set_query_timer_restart
return $self->{query_timer_restart};
}
+=pod
+
+=item $session->poll_query_until($query [, $expected ])
+
+Run B<$query> repeatedly in this background session, until it returns the
+B<$expected> result ('t', or SQL boolean true, by default).
+Continues polling if the query returns an error result.
+Times out after a reasonable number of attempts.
+Returns 1 if successful, 0 if timed out.
+
+=cut
+
+sub poll_query_until
+{
+ my ($self, $query, $expected) = @_;
+
+ $expected = 't' unless defined($expected); # default value
+
+ my $max_attempts = 10 * $PostgreSQL::Test::Utils::timeout_default;
+ my $attempts = 0;
+ my ($stdout, $stderr_flag);
+
+ while ($attempts < $max_attempts)
+ {
+ ($stdout, $stderr_flag) = $self->query($query);
+
+ chomp($stdout);
+
+ # If query succeeded and returned expected result
+ if (!$stderr_flag && $stdout eq $expected)
+ {
+ return 1;
+ }
+
+ # Wait 0.1 second before retrying.
+ usleep(100_000);
+
+ $attempts++;
+ }
+
+ # Give up. Print the output from the last attempt, hopefully that's useful
+ # for debugging.
+ my $stderr_output = $stderr_flag ? $self->{stderr} : '';
+ diag qq(poll_query_until timed out executing this query:
+$query
+expecting this output:
+$expected
+last actual query output:
+$stdout
+with stderr:
+$stderr_output);
+ return 0;
+}
+
+=item $session->wait_for_event(backend_type, wait_event_name)
+
+Poll pg_stat_activity until backend_type reaches wait_event_name using this
+background session.
+
+=cut
+
+sub wait_for_event
+{
+ my ($self, $backend_type, $wait_event_name) = @_;
+
+ $self->poll_query_until(qq[
+ SELECT count(*) > 0 FROM pg_stat_activity
+ WHERE backend_type = '$backend_type' AND wait_event = '$wait_event_name'
+ ])
+ or die
+ qq(timed out when waiting for $backend_type to reach wait event '$wait_event_name');
+
+ return;
+}
+
1;
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 432509277c9..f7ce00990cc 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2774,6 +2774,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.34.1
[text/x-patch] 0002-Memory-and-address-space-management-for-buf-20251114.patch (70.1K, ../../CAExHW5sVxEwQsuzkgjjJQP9-XVe0H2njEVw1HxeYFdT7u7J+eQ@mail.gmail.com/5-0002-Memory-and-address-space-management-for-buf-20251114.patch)
download | inline diff:
From a7f25c62ef900b2b115c575c2d8aa158ec825c69 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH 2/4] Memory and address space management for buffer resizing
This has three changes
1. Allow to use multiple shared memory mappings
============================================
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
A new shared memory API is introduced, extended with a segment as a new
parameter. As a path of least resistance, the original API is kept in
place, utilizing the main shared memory segment.
Modifies pg_shmem_allocations to report shared memory segment as well.
Adds pg_shmem_segments to report shared memory segment information.
2. Address space reservation for shared memory
============================================
Currently the shared memory layout is designed to pack everything tight
together, leaving no space between mappings for resizing. Here is how it
looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
...
Make the layout more dynamic via splitting every shared memory segment
into two parts:
* An anonymous file, which actually contains shared memory content.
Such an anonymous file is created via memfd_create, it lives in
memory, behaves like a regular file and semantically equivalent to an
anonymous memory allocated via mmap with MAP_ANONYMOUS.
* A reservation mapping, which size is much larger than required shared
segment size. This mapping is created with flag MAP_NORESERVE (to not
count the reserved space against memory limits). The anonymous file is
mapped into this reservation mapping.
If we have to change the address maps while resizing the shared buffer
pool, it is needed to be done in Postmaster too, so that the new
backends will inherit the resized address space from the Postmaster.
However, Postmaster is not invovled in ProcSignalBarrier mechanism and
we don't want it to spend time in things other than its core
functionality. To achive that, maximum required address space maps are
setup upfront with read and write access when starting the server. When
resizing the buffer pool only the backing file object is resized from
the coordinator. This also makes the ProcSignalBarrier handling code
light for backends other than the coordinator.
The resulting layout looks like this:
00400000-00490000 /path/bin/postgres
...
3f526000-3f590000 rw-p [heap]
7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
To resize a shared memory segment in this layout it's possible to use
ftruncate on the memory mapped file.
This approach also do not impact the actual memory usage as reported by
the kernel.
TODO: Verify that Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched within
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
20770816 (~19.8 MB)
There are also few unrelated advantages of using memory mapped files:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps.
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process. - Some hackers have expressed concerns over it.
The downside is that memfd_create is Linux specific.
3. Refactor CalculateShmemSize()
=============================
This function calls many functions which return the amount of shared
memory required for different shared memory data structures. Up until
now, the returned total of these sizes was used to create a single
shared memory segment. With this change, CalculateShmemSize() needs to
estimate memory requirements for each of the segments. It now takes an
array of MemoryMappingSizes, containing as many elements as the number
of segments, as an argument. The sizes returned by all the function it
calls, except BufferManagerShmemSize(), are added and saved in the first
element (index 0) of the array. BufferManagerShmemSize() is modified to
save the amount of memory required for buffer manager related segments
in the corresponding array element. Additionally it also saves the
amount of reserved space. For now, the amount of reserved address space
is same as the amount of required memory but that is expected to change
with the next commit which implements buffer pool resize.
CalculateShmemSize() now returns the total of sizes corresponding to all
the sizes.
Author: Dmitrii Dolgov and Ashutosh Bapat
Reviewed-by: Tomas Vondra
---
doc/src/sgml/system-views.sgml | 9 +
src/backend/catalog/system_views.sql | 7 +
src/backend/port/sysv_shmem.c | 425 +++++++++++++++++++------
src/backend/port/win32_sema.c | 2 +-
src/backend/port/win32_shmem.c | 14 +-
src/backend/storage/buffer/buf_init.c | 56 ++--
src/backend/storage/buffer/buf_table.c | 6 +-
src/backend/storage/buffer/freelist.c | 5 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 99 ++++--
src/backend/storage/ipc/shmem.c | 243 ++++++++++----
src/backend/storage/lmgr/lwlock.c | 15 +-
src/include/catalog/pg_proc.dat | 12 +-
src/include/portability/mem.h | 2 +-
src/include/storage/bufmgr.h | 3 +-
src/include/storage/ipc.h | 4 +-
src/include/storage/pg_shmem.h | 60 +++-
src/include/storage/shmem.h | 12 +
src/test/regress/expected/rules.out | 10 +-
19 files changed, 755 insertions(+), 233 deletions(-)
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 8f3e2741051..bc70a3ee6c9 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -4233,6 +4233,15 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
</para></entry>
</row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>segment</structfield> <type>text</type>
+ </para>
+ <para>
+ The name of the shared memory segment concerning the allocation.
+ </para></entry>
+ </row>
+
<row>
<entry role="catalog_table_entry"><para role="column_definition">
<structfield>off</structfield> <type>int8</type>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 059e8778ca7..59145066647 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -668,6 +668,13 @@ GRANT SELECT ON pg_shmem_allocations TO pg_read_all_stats;
REVOKE EXECUTE ON FUNCTION pg_get_shmem_allocations() FROM PUBLIC;
GRANT EXECUTE ON FUNCTION pg_get_shmem_allocations() TO pg_read_all_stats;
+CREATE VIEW pg_shmem_segments AS
+ SELECT * FROM pg_get_shmem_segments();
+
+REVOKE ALL ON pg_shmem_segments FROM PUBLIC;
+GRANT SELECT ON pg_shmem_segments TO pg_read_all_stats;
+REVOKE EXECUTE ON FUNCTION pg_get_shmem_segments() FROM PUBLIC;
+GRANT EXECUTE ON FUNCTION pg_get_shmem_segments() TO pg_read_all_stats;
CREATE VIEW pg_shmem_allocations_numa AS
SELECT * FROM pg_get_shmem_allocations_numa();
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..cc4b2c80e1a 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -90,12 +90,49 @@ typedef enum
SHMSTATE_UNATTACHED, /* pertinent to DataDir, no attached PIDs */
} IpcMemoryState;
-
+/*
+ * TODO: These should be moved into ShmemSegment, now that there can be multiple
+ * shared memory segments. But there's windows specific code which will need
+ * adjustment, so leaving it here.
+ */
unsigned long UsedShmemSegID = 0;
void *UsedShmemSegAddr = NULL;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+/*
+ * Anonymous mapping layout we use looks like this:
+ *
+ * 00400000-00c2a000 r-xp /bin/postgres
+ * ...
+ * 3f526000-3f590000 rw-p [heap]
+ * 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted)
+ * 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted)
+ * 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
+ * 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
+ * ...
+ *
+ * We need to place shared memory mappings in such a way, that there will be
+ * gaps between them in the address space. Those gaps have to be large enough
+ * to resize the mapping up to certain size, without counting towards the total
+ * memory consumption.
+ *
+ * To achieve this, for each shared memory segment we first create an anonymous
+ * file of specified size using memfd_create, which will accomodate actual
+ * shared memory mapping content. It is represented by the first /memfd:main
+ * with rw permissions. Then we create a mapping for this file using mmap, with
+ * size much larger than required and flags PROT_NONE (allows to make sure the
+ * reserved space will not be used) and MAP_NORESERVE (prevents the space from
+ * being counted against memory limits). The mapping serves as an address space
+ * reservation, into which shared memory segment can be extended and is
+ * represented by the second /memfd:main with no permissions.
+ */
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -104,6 +141,27 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId,
void *attachAt,
PGShmemHeader **addr);
+const char*
+MappingName(int shmem_segment)
+{
+ switch (shmem_segment)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
+ default:
+ return "unknown";
+ }
+}
/*
* InternalIpcMemoryCreate(memKey, size)
@@ -470,19 +528,20 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* hugepage sizes, we might want to think about more invasive strategies,
* such as increasing shared_buffers to absorb the extra space.
*
- * Returns the (real, assumed or config provided) page size into
- * *hugepagesize, and the hugepage-related mmap flags to use into
- * *mmap_flags if requested by the caller. If huge pages are not supported,
- * *hugepagesize and *mmap_flags are set to 0.
+ * Returns the (real, assumed or config provided) page size into *hugepagesize,
+ * the hugepage-related mmap and memfd flags to use into *mmap_flags and
+ * *memfd_flags if requested by the caller. If huge pages are not supported,
+ * *hugepagesize, *mmap_flags and *memfd_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -541,6 +600,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -551,7 +611,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
{
int shift = pg_ceil_log2_64(hugepagesize_local);
- mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
+ }
+#endif
+
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT;
}
#endif
@@ -560,6 +629,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -567,6 +638,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -588,83 +661,242 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
return true;
}
+/*
+ * Wrapper around posix_fallocate() to allocate memory for a given shared memory
+ * segment.
+ *
+ * Performs retry on EINTR, and raises error upon failure.
+ */
+static void
+shmem_fallocate(int fd, const char *mapping_name, Size size, int elevel)
+{
+#if defined(HAVE_POSIX_FALLOCATE) && defined(__linux__)
+ int ret;
+
+
+ /*
+ * If there is not enough memory, trying to access a hole in address space
+ * will cause SIGBUS. If supported, avoid that by allocating memory upfront.
+ *
+ * We still use a traditional EINTR retry loop to handle SIGCONT.
+ * posix_fallocate() doesn't restart automatically, and we don't want this to
+ * fail if you attach a debugger.
+ */
+ do
+ {
+ ret = posix_fallocate(fd, 0, size);
+ } while (ret == EINTR);
+
+ if (ret != 0)
+ {
+ ereport(elevel,
+ (errmsg("segment[%s]: could not allocate space for anonymous file: %s",
+ mapping_name, strerror(ret)),
+ (ret == ENOMEM) ?
+ errhint("This error usually means that PostgreSQL's request "
+ "for a shared memory segment exceeded available memory, "
+ "swap space, or huge pages. To reduce the request size "
+ "(currently %zu bytes), reduce PostgreSQL's shared "
+ "memory usage, perhaps by reducing \"shared_buffers\" or "
+ "\"max_connections\".",
+ size) : 0));
+ }
+#endif /* HAVE_POSIX_FALLOCATE && __linux__ */
+}
+
+/*
+ * Round up the required amount of memory and the amount of required reserved
+ * address space to the nearest huge page size.
+ */
+static inline void
+round_off_mapping_sizes_for_hugepages(MemoryMappingSizes *mapping, int hugepagesize)
+{
+ if (hugepagesize == 0)
+ return;
+
+ if (mapping->shmem_req_size % hugepagesize != 0)
+ mapping->shmem_req_size += hugepagesize -
+ (mapping->shmem_req_size % hugepagesize);
+
+ if (mapping->shmem_reserved % hugepagesize != 0)
+ mapping->shmem_reserved = mapping->shmem_reserved + hugepagesize -
+ (mapping->shmem_reserved % hugepagesize);
+}
+
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested. If needed,
+ * it also rounds up the mapping reserved size to be a multiple of huge page
+ * size.
+ *
+ * Note that we do not fallback from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
+ *
+ * TODO: Update the prologue to be consistent with the code.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(MemoryMappingSizes *mapping, int segment_id)
{
- Size allocsize = *size;
void *ptr = MAP_FAILED;
- int mmap_errno = 0;
+ int save_errno = 0;
+ int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0;
+ ShmemSegment *segment = &Segments[segment_id];
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
- int mmap_flags;
- GetHugePageSize(&hugepagesize, &mmap_flags);
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
- if (allocsize % hugepagesize != 0)
- allocsize += hugepagesize - (allocsize % hugepagesize);
+ /* Round up the request size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags);
+ round_off_mapping_sizes_for_hugepages(mapping, hugepagesize);
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ /* Verify that the new size is withing the reserved boundaries */
+ Assert(mapping->shmem_reserved >= mapping->shmem_req_size);
+
+ mmap_flags = PG_MMAP_FLAGS | mmap_flags;
}
#endif
/*
- * Report whether huge pages are in use. This needs to be tracked before
- * the second mmap() call if attempting to use huge pages failed
- * previously.
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
*/
- SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
- PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ segment->segment_fd = memfd_create(MappingName(segment_id), memfd_flags);
+ if (segment->segment_fd == -1)
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not create anonymous shared memory file: %m",
+ MappingName(segment_id))));
- if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
- {
- /*
- * Use the original size, not the rounded-up value, when falling back
- * to non-huge pages.
- */
- allocsize = *size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS, -1, 0);
- mmap_errno = errno;
- }
+ elog(DEBUG1, "segment[%s]: mmap(%zu)", MappingName(segment_id), mapping->shmem_req_size);
+ /*
+ * Reserve maximum required address space for future expansion of this
+ * memory segment. MAP_NORESERVE ensures that no memory is allocated. The
+ * whole address space will be setup for read/write access, so that memory
+ * allocated to this address space can be read or written to even if it is
+ * resized.
+ */
+ ptr = mmap(NULL, mapping->shmem_reserved, PROT_READ | PROT_WRITE,
+ mmap_flags | MAP_NORESERVE, segment->segment_fd, 0);
if (ptr == MAP_FAILED)
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ MappingName(segment_id))));
+
+ /*
+ * Resize the backing file to the required size. On platforms where it is
+ * supported, we also allocate the required memory upfront. On other
+ * platform the memory upto the size of file will be allocated on demand.
+ */
+ if(ftruncate(segment->segment_fd, mapping->shmem_req_size) == -1)
{
- errno = mmap_errno;
+ save_errno = errno;
+
+ close(segment->segment_fd);
+
+ errno = save_errno;
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
- (mmap_errno == ENOMEM) ?
+ (errmsg("segment[%s]: could not truncate anonymous file to size %zu: %m",
+ MappingName(segment_id), mapping->shmem_req_size),
+ (save_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
"swap space, or huge pages. To reduce the request size "
"(currently %zu bytes), reduce PostgreSQL's shared "
"memory usage, perhaps by reducing \"shared_buffers\" or "
"\"max_connections\".",
- allocsize) : 0));
+ mapping->shmem_req_size) : 0));
}
+ shmem_fallocate(segment->segment_fd, MappingName(segment_id), mapping->shmem_req_size, FATAL);
- *size = allocsize;
- return ptr;
+ segment->shmem = ptr;
+ segment->shmem_size = mapping->shmem_req_size;
+ segment->shmem_reserved = mapping->shmem_reserved;
+}
+
+/*
+ * PrepareHugePages
+ *
+ * Figure out if there are enough huge pages to allocate all shared memory
+ * segments, and report that information via huge_pages_status and
+ * huge_pages_on. It needs to be called before creating shared memory segments.
+ *
+ * It is necessary to maintain the same semantic (simple on/off) for
+ * huge_pages_status, even if there are multiple shared memory segments: all
+ * segments either use huge pages or not, there is no mix of segments with
+ * different page size. The latter might be actually beneficial, in particular
+ * because only some segments may require large amount of memory, but for now
+ * we go with a simple solution.
+ */
+void
+PrepareHugePages()
+{
+ void *ptr = MAP_FAILED;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+
+ CalculateShmemSize(mapping_sizes);
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize, total_size = 0;
+ int mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &mmap_flags, NULL);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory and
+ * decide whether we will allocate huge pages or not.
+ */
+ for(int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ Size segment_size = mapping_sizes[segment].shmem_req_size;
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0);
+ }
+#endif
+
+ /*
+ * Report whether huge pages are in use. This needs to be tracked before
+ * creating shared memory segments.
+ */
+ SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
}
/*
@@ -674,20 +906,25 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for(int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ ShmemSegment *segment = &Segments[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (segment->shmem != NULL)
+ {
+ if (munmap(segment->shmem, segment->shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ segment->shmem, segment->shmem_size);
+ segment->shmem = NULL;
+ }
}
}
/*
* PGSharedMemoryCreate
*
- * Create a shared memory segment of the given size and initialize its
+ * Create a shared memory segment for the given mapping and initialize its
* standard header. Also, register an on_shmem_exit callback to release
* the storage.
*
@@ -697,7 +934,7 @@ AnonymousShmemDetach(int status, Datum arg)
* postmaster or backend.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(MemoryMappingSizes *mapping, int segment_id,
PGShmemHeader **shim)
{
IpcMemoryKey NextShmemSegID;
@@ -705,6 +942,7 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ ShmemSegment *segment = &Segments[segment_id];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -717,14 +955,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -732,12 +962,12 @@ PGSharedMemoryCreate(Size size,
errmsg("huge pages not supported with the current \"shared_memory_type\" setting")));
/* Room for a header? */
- Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ Assert(mapping->shmem_req_size > MAXALIGN(sizeof(PGShmemHeader)));
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(mapping, segment_id);
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -747,7 +977,7 @@ PGSharedMemoryCreate(Size size,
}
else
{
- sysvsize = size;
+ sysvsize = mapping->shmem_req_size;
/* huge pages are only available with mmap */
SetConfigOption("huge_pages_status", "off",
@@ -760,7 +990,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + segment_id;
for (;;)
{
@@ -852,13 +1082,13 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_req_size;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ segment->seg_addr = memAddress;
+ segment->seg_id = (unsigned long) NextShmemSegID;
/*
* If AnonymousShmem is NULL here, then we're not using anonymous shared
@@ -866,10 +1096,10 @@ PGSharedMemoryCreate(Size size,
* block. Otherwise, the System V shared memory block is only a shim, and
* we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (segment->shmem == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(segment->shmem, hdr, sizeof(PGShmemHeader));
+ return (PGShmemHeader *) segment->shmem;
}
#ifdef EXEC_BACKEND
@@ -969,23 +1199,28 @@ PGSharedMemoryNoReAttach(void)
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for(int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ ShmemSegment *segment = &Segments[i];
+
+ if (segment->seg_addr != NULL)
+ {
+ if ((shmdt(segment->seg_addr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", segment->seg_addr);
+ segment->seg_addr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (segment->shmem != NULL)
+ {
+ if (munmap(segment->shmem, segment->shmem_size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ segment->shmem, segment->shmem_size);
+ segment->shmem = NULL;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index 5854ad1f54d..e7365ff8060 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 4dee856d6bd..5c0c32babaf 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -204,7 +204,7 @@ EnableLockPagesPrivilege(int elevel)
* standard header.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(MemoryMappingSizes *mapping_sizes, int segment_id,
PGShmemHeader **shim)
{
void *memAddress;
@@ -216,9 +216,10 @@ PGSharedMemoryCreate(Size size,
DWORD size_high;
DWORD size_low;
SIZE_T largePageSize = 0;
- Size orig_size = size;
+ Size size = mapping_sizes->shmem_req_size;
DWORD flProtect = PAGE_READWRITE;
DWORD desiredAccess;
+ ShmemSegment *segment = &Segments[segment_id]
ShmemProtectiveRegion = VirtualAlloc(NULL, PROTECTIVE_REGION_SIZE,
MEM_RESERVE, PAGE_NOACCESS);
@@ -304,7 +305,7 @@ retry:
* Use the original size, not the rounded-up value, when
* falling back to non-huge pages.
*/
- size = orig_size;
+ size = mapping_sizes->shmem_req_size;
flProtect = PAGE_READWRITE;
goto retry;
}
@@ -393,6 +394,11 @@ retry:
hdr->dsm_control = 0;
/* Save info for possible future use */
+ segment->shmem_size = size;
+ segment->seg_addr = memAddress;
+ segment->shmem = (Pointer) hdr;
+ segment->seg_id = (unsigned long) hmap2;
+
UsedShmemSegAddr = memAddress;
UsedShmemSegSize = size;
UsedShmemSegID = hmap2;
@@ -627,7 +633,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6fd3a6bbac5..4fa547f48de 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -62,7 +63,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -74,22 +78,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
+ ShmemInitStructInSegment("Buffer Descriptors",
NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
+ ShmemInitStructInSegment("Buffer Blocks",
NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -99,8 +103,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -147,33 +152,42 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
{
- Size size = 0;
+ size_t size;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
+ /* size of buffer descriptors, plus alignment padding */
+ size = add_size(0, mul_size(NBuffers, sizeof(BufferDescPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
+ mapping_sizes[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_reserved = size;
/* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(0, PG_IO_ALIGN_SIZE);
size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_reserved = size;
/* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ mapping_sizes[STRATEGY_SHMEM_SEGMENT].shmem_req_size = StrategyShmemSize();
+ mapping_sizes[STRATEGY_SHMEM_SEGMENT].shmem_reserved = StrategyShmemSize();
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
+ /* size of I/O condition variables, plus alignment padding */
+ size = add_size(0, mul_size(NBuffers,
sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
size = add_size(size, PG_CACHE_LINE_SIZE);
+ mapping_sizes[BUFFER_IOCV_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[BUFFER_IOCV_SHMEM_SEGMENT].shmem_reserved = size;
/* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
+ mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_reserved = mul_size(NBuffers, sizeof(CkptSortItem));
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index f0c39ec2822..67e87f9935d 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -25,6 +25,7 @@
#include "funcapi.h"
#include "storage/buf_internals.h"
#include "storage/lwlock.h"
+#include "storage/pg_shmem.h"
#include "utils/rel.h"
#include "utils/builtins.h"
@@ -64,10 +65,11 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
- SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
+ SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table",
size, size,
&info,
- HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE);
+ HASH_ELEM | HASH_BLOBS | HASH_PARTITION | HASH_FIXED_SIZE,
+ STRATEGY_SHMEM_SEGMENT);
}
/*
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 28d952b3534..13ee840ab9f 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -418,9 +419,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
+ ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found);
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index 2704e80b3a7..1965b2d3eb4 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -61,6 +61,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -68,7 +70,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index b23d0c19360..41190f96639 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -81,10 +81,17 @@ RequestAddinShmemSpace(Size size)
/*
* CalculateShmemSize
- * Calculates the amount of shared memory needed.
+ * Calculates the amount of shared memory needed.
+ *
+ * The amount of shared memory required per segment is saved in mapping_sizes,
+ * which is expected to be an array of size NUM_MEMORY_MAPPINGS. The total
+ * amount of memory needed across all the segments is returned. For the memory
+ * mappings which reserve address space for future expansion, the required
+ * amount of reserved space is saved in mapping_sizes of those segments.
+ * This memory is not included in the returned value.
*/
Size
-CalculateShmemSize(void)
+CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
{
Size size;
@@ -102,7 +109,13 @@ CalculateShmemSize(void)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+
+ /*
+ * Buffer manager adds estimates for memory requirements for every shared
+ * memory segment that it uses in the corresponding AnonymousMappings.
+ * Consider size required from only the main shared memory segment here.
+ */
+ size = add_size(size, BufferManagerShmemSize(mapping_sizes));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
@@ -144,8 +157,22 @@ CalculateShmemSize(void)
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
+ /*
+ * All the shared memory allocations considered so far happen in the main
+ * shared memory segment.
+ */
+ mapping_sizes[MAIN_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[MAIN_SHMEM_SEGMENT].shmem_reserved = size;
+
+ size = 0;
/* might as well round it off to a multiple of a typical page size */
- size = add_size(size, 8192 - (size % 8192));
+ for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ mapping_sizes[segment].shmem_req_size = add_size(mapping_sizes[segment].shmem_req_size, 8192 - (mapping_sizes[segment].shmem_req_size % 8192));
+ mapping_sizes[segment].shmem_reserved = add_size(mapping_sizes[segment].shmem_reserved, 8192 - (mapping_sizes[segment].shmem_reserved % 8192));
+ /* Compute the total size of all segments */
+ size = size + mapping_sizes[segment].shmem_req_size;
+ }
return size;
}
@@ -191,32 +218,44 @@ CreateSharedMemoryAndSemaphores(void)
{
PGShmemHeader *shim;
PGShmemHeader *seghdr;
- Size size;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize();
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+ CalculateShmemSize(mapping_sizes);
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
-
- /*
- * Make sure that huge pages are never reported as "unknown" while the
- * server is running.
- */
- Assert(strcmp("unknown",
- GetConfigOption("huge_pages_status", false, false)) != 0);
-
- InitShmemAccess(seghdr);
+ /* Decide if we use huge pages or regular size pages */
+ PrepareHugePages();
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for(int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ MemoryMappingSizes *mapping = &mapping_sizes[segment];
+
+ /* Compute the size of the shared-memory block */
+ elog(DEBUG3, "invoking IpcMemoryCreate(segment %s, size=%zu, reserved address space=%zu)",
+ MappingName(segment), mapping->shmem_req_size, mapping->shmem_reserved);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(mapping, segment, &shim);
+
+ /*
+ * Make sure that huge pages are never reported as "unknown" while the
+ * server is running.
+ */
+ Assert(strcmp("unknown",
+ GetConfigOption("huge_pages_status", false, false)) != 0);
+
+ InitShmemAccessInSegment(seghdr, segment);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocationInSegment(segment);
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -334,7 +373,9 @@ CreateOrAttachShmemStructs(void)
* InitializeShmemGUCs
*
* This function initializes runtime-computed GUCs related to the amount of
- * shared memory required for the current configuration.
+ * shared memory required for the current configuration. It assumes that the
+ * memory required by the shared memory segments is already calculated and is
+ * available in AnonymousMappings.
*/
void
InitializeShmemGUCs(void)
@@ -343,11 +384,13 @@ InitializeShmemGUCs(void)
Size size_b;
Size size_mb;
Size hp_size;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize();
+ size_b = CalculateShmemSize(mapping_sizes);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
@@ -356,7 +399,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 0f18beb6ad4..f303a9328df 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -76,20 +76,19 @@
#include "utils/builtins.h"
static void *ShmemAllocRaw(Size size, Size *allocated_size);
-static void *ShmemAllocUnlocked(Size size);
+static void *ShmemAllocRawInSegment(Size size, Size *allocated_size,
+ int shmem_segment);
/* shared memory global variables */
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ShmemSegment Segments[NUM_MEMORY_MAPPINGS];
-static void *ShmemBase; /* start address of shared memory */
-
-static void *ShmemEnd; /* end+1 address of shared memory */
-
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
-
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -102,9 +101,17 @@ Datum pg_numa_available(PG_FUNCTION_ARGS);
void
InitShmemAccess(PGShmemHeader *seghdr)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment)
+{
+ PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr;
+ ShmemSegment *seg = &Segments[shmem_segment];
+ seg->ShmemSegHdr = shmhdr;
+ seg->ShmemBase = (void *) shmhdr;
+ seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize;
}
/*
@@ -115,7 +122,13 @@ InitShmemAccess(PGShmemHeader *seghdr)
void
InitShmemAllocation(void)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT);
+}
+
+void
+InitShmemAllocationInSegment(int shmem_segment)
+{
+ PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr;
char *aligned;
Assert(shmhdr != NULL);
@@ -124,9 +137,9 @@ InitShmemAllocation(void)
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment);
- SpinLockInit(ShmemLock);
+ SpinLockInit(Segments[shmem_segment].ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -151,16 +164,22 @@ InitShmemAllocation(void)
*/
void *
ShmemAlloc(Size size)
+{
+ return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemAllocInSegment(Size size, int shmem_segment)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ MappingName(shmem_segment), size)));
return newSpace;
}
@@ -185,6 +204,12 @@ ShmemAllocNoError(Size size)
*/
static void *
ShmemAllocRaw(Size size, Size *allocated_size)
+{
+ return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT);
+}
+
+static void *
+ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -204,22 +229,22 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[shmem_segment].ShmemLock);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[shmem_segment].ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -228,15 +253,16 @@ ShmemAllocRaw(Size size, Size *allocated_size)
}
/*
- * ShmemAllocUnlocked -- allocate max-aligned chunk from shared memory
+ * ShmemAllocUnlockedInSegment
+ * allocate max-aligned chunk from given shared memory segment
*
* Allocate space without locking ShmemLock. This should be used for,
* and only for, allocations that must happen before ShmemLock is ready.
*
* We consider maxalign, rather than cachealign, sufficient here.
*/
-static void *
-ShmemAllocUnlocked(Size size)
+void *
+ShmemAllocUnlockedInSegment(Size size, int shmem_segment)
{
Size newStart;
Size newFree;
@@ -247,19 +273,19 @@ ShmemAllocUnlocked(Size size)
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(Segments[shmem_segment].ShmemSegHdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
- ShmemSegHdr->freeoffset = newFree;
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ MappingName(shmem_segment), size)));
+ Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -274,7 +300,13 @@ ShmemAllocUnlocked(Size size)
bool
ShmemAddrIsValid(const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT);
+}
+
+bool
+ShmemAddrIsValidInSegment(const void *addr, int shmem_segment)
+{
+ return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd);
}
/*
@@ -335,6 +367,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
int64 max_size, /* max size of the table */
HASHCTL *infoP, /* info about key and bucket size */
int hash_flags) /* info about infoP */
+{
+ return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags,
+ MAIN_SHMEM_SEGMENT);
+}
+
+HTAB *
+ShmemInitHashInSegment(const char *name, /* table string name for shmem index */
+ long init_size, /* initial table size */
+ long max_size, /* max size of the table */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags, /* info about infoP */
+ int shmem_segment) /* in which segment to keep the table */
{
bool found;
void *location;
@@ -351,9 +395,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
+ location = ShmemInitStructInSegment(name,
hash_get_shared_size(infoP, hash_flags),
- &found);
+ &found, shmem_segment);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -386,6 +430,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
+ int shmem_segment)
{
ShmemIndexEnt *result;
void *structPtr;
@@ -394,7 +445,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr;
/* Must be trying to create/attach to ShmemIndex itself */
Assert(strcmp(name, "ShmemIndex") == 0);
@@ -417,7 +468,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInSegment(size, shmem_segment);
shmemseghdr->index = structPtr;
*foundPtr = false;
}
@@ -434,8 +485,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, shmem_segment)));
}
if (*foundPtr)
@@ -460,7 +511,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -475,18 +526,18 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
result->size = size;
result->allocated_size = allocated_size;
result->location = structPtr;
+ result->shmem_segment = shmem_segment;
}
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -527,13 +578,14 @@ mul_size(Size s1, Size s2)
Datum
pg_get_shmem_allocations(PG_FUNCTION_ARGS)
{
-#define PG_GET_SHMEM_SIZES_COLS 4
+#define PG_GET_SHMEM_SIZES_COLS 5
ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
HASH_SEQ_STATUS hstat;
ShmemIndexEnt *ent;
- Size named_allocated = 0;
+ Size named_allocated[NUM_MEMORY_MAPPINGS] = {0};
Datum values[PG_GET_SHMEM_SIZES_COLS];
bool nulls[PG_GET_SHMEM_SIZES_COLS];
+ int i;
InitMaterializedSRF(fcinfo, 0);
@@ -546,29 +598,40 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
- values[2] = Int64GetDatum(ent->size);
- values[3] = Int64GetDatum(ent->allocated_size);
- named_allocated += ent->allocated_size;
+ values[1] = CStringGetTextDatum(MappingName(ent->shmem_segment));
+ values[2] = Int64GetDatum((char *) ent->location - (char *) Segments[ent->shmem_segment].ShmemSegHdr);
+ values[3] = Int64GetDatum(ent->size);
+ values[4] = Int64GetDatum(ent->allocated_size);
+ named_allocated[ent->shmem_segment] += ent->allocated_size;
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
values, nulls);
}
/* output shared memory allocated but not counted via the shmem index */
- values[0] = CStringGetTextDatum("<anonymous>");
- nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ values[0] = CStringGetTextDatum("<anonymous>");
+ values[1] = CStringGetTextDatum(MappingName(i));
+ nulls[2] = true;
+ values[3] = Int64GetDatum(Segments[i].ShmemSegHdr->freeoffset - named_allocated[i]);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
/* output as-of-yet unused shared memory */
- nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
- nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ memset(nulls, 0, sizeof(nulls));
+
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGShmemHeader *shmhdr = Segments[i].ShmemSegHdr;
+ nulls[0] = true;
+ values[1] = CStringGetTextDatum(MappingName(i));
+ values[2] = Int64GetDatum(shmhdr->freeoffset);
+ values[3] = Int64GetDatum(shmhdr->totalsize - shmhdr->freeoffset);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
LWLockRelease(ShmemIndexLock);
@@ -593,7 +656,7 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size os_page_size;
void **page_ptrs;
int *pages_status;
- uint64 shm_total_page_count,
+ uint64 shm_total_page_count = 0,
shm_ent_page_count,
max_nodes;
Size *nodes;
@@ -628,7 +691,12 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ for(int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ PGShmemHeader *shmhdr = Segments[segment].ShmemSegHdr;
+ shm_total_page_count += (shmhdr->totalsize / os_page_size) + 1;
+ }
+
page_ptrs = palloc0(sizeof(void *) * shm_total_page_count);
pages_status = palloc(sizeof(int) * shm_total_page_count);
@@ -751,7 +819,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
@@ -761,3 +829,46 @@ pg_numa_available(PG_FUNCTION_ARGS)
{
PG_RETURN_BOOL(pg_numa_init() != -1);
}
+
+/* SQL SRF showing shared memory segments */
+Datum
+pg_get_shmem_segments(PG_FUNCTION_ARGS)
+{
+#define PG_GET_SHMEM_SEGS_COLS 6
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ Datum values[PG_GET_SHMEM_SEGS_COLS];
+ bool nulls[PG_GET_SHMEM_SEGS_COLS];
+ int i;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* output all allocated entries */
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+ int j;
+
+ if (shmhdr == NULL)
+ {
+ for (j = 0; j < PG_GET_SHMEM_SEGS_COLS; j++)
+ nulls[j] = true;
+ }
+ else
+ {
+ memset(nulls, 0, sizeof(nulls));
+ values[0] = Int32GetDatum(i);
+ values[1] = CStringGetTextDatum(MappingName(i));
+ values[2] = Int64GetDatum(shmhdr->totalsize);
+ values[3] = Int64GetDatum(shmhdr->freeoffset);
+ values[4] = Int64GetDatum(segment->shmem_size);
+ values[5] = Int64GetDatum(segment->shmem_reserved);
+ }
+
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
+
+ return (Datum) 0;
+}
+
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index b017880f5e4..c25dd13b63a 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -612,12 +614,15 @@ LWLockNewTrancheId(const char *name)
/*
* We use the ShmemLock spinlock to protect LWLockCounter and
* LWLockTrancheNames.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
*/
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
if (*LWLockCounter - LWTRANCHE_FIRST_USER_DEFINED >= MAX_NAMED_TRANCHES)
{
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
ereport(ERROR,
(errmsg("maximum number of tranches already registered"),
errdetail("No more than %d tranches may be registered.",
@@ -628,7 +633,7 @@ LWLockNewTrancheId(const char *name)
LocalLWLockCounter = *LWLockCounter;
strlcpy(LWLockTrancheNames[result - LWTRANCHE_FIRST_USER_DEFINED], name, NAMEDATALEN);
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
@@ -750,9 +755,9 @@ GetLWTrancheName(uint16 trancheId)
*/
if (trancheId >= LocalLWLockCounter)
{
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
LocalLWLockCounter = *LWLockCounter;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock);
if (trancheId >= LocalLWLockCounter)
elog(ERROR, "tranche %d is not registered", trancheId);
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 5cf9e12fcb9..411043ca750 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8576,8 +8576,8 @@
{ oid => '5052', descr => 'allocations from the main shared memory segment',
proname => 'pg_get_shmem_allocations', prorows => '50', proretset => 't',
provolatile => 'v', prorettype => 'record', proargtypes => '',
- proallargtypes => '{text,int8,int8,int8}', proargmodes => '{o,o,o,o}',
- proargnames => '{name,off,size,allocated_size}',
+ proallargtypes => '{text,text,int8,int8,int8}', proargmodes => '{o,o,o,o,o}',
+ proargnames => '{name,segment,off,size,allocated_size}',
prosrc => 'pg_get_shmem_allocations' },
{ oid => '4099', descr => 'Is NUMA support available?',
@@ -8600,6 +8600,14 @@
proargmodes => '{o,o,o}', proargnames => '{name,type,size}',
prosrc => 'pg_get_dsm_registry_allocations' },
+# shared memory segments
+{ oid => '5101', descr => 'shared memory segments',
+ proname => 'pg_get_shmem_segments', prorows => '6', proretset => 't',
+ provolatile => 'v', prorettype => 'record', proargtypes => '',
+ proallargtypes => '{int4,text,int8,int8,int8,int8}', proargmodes => '{o,o,o,o,o,o}',
+ proargnames => '{id,name,size,freeoffset,mapping_size,mapping_reserved_size}',
+ prosrc => 'pg_get_shmem_segments' },
+
# memory context of local backend
{ oid => '2282',
descr => 'information about all memory contexts of local backend',
diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h
index ef9800732d9..40588ff6968 100644
--- a/src/include/portability/mem.h
+++ b/src/include/portability/mem.h
@@ -38,7 +38,7 @@
#define MAP_NOSYNC 0
#endif
-#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
+#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
/* Some really old systems don't define MAP_FAILED. */
#ifndef MAP_FAILED
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index b5f8f3c5d42..3769f4db7dc 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -19,6 +19,7 @@
#include "storage/block.h"
#include "storage/buf.h"
#include "storage/bufpage.h"
+#include "storage/pg_shmem.h"
#include "storage/relfilelocator.h"
#include "utils/relcache.h"
#include "utils/snapmgr.h"
@@ -326,7 +327,7 @@ extern void EvictRelUnpinnedBuffers(Relation rel,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index 2a8a8f0eabd..d73f1b407db 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -18,6 +18,8 @@
#ifndef IPC_H
#define IPC_H
+#include "storage/pg_shmem.h"
+
typedef void (*pg_on_exit_callback) (int code, Datum arg);
typedef void (*shmem_startup_hook_type) (void);
@@ -77,7 +79,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(void);
+extern Size CalculateShmemSize(MemoryMappingSizes *mapping_sizes);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 5f7d4b83a60..beee0a53d2d 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,6 +25,13 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
+
+typedef struct MemoryMappingSizes
+{
+ Size shmem_req_size; /* Required size of the segment */
+ Size shmem_reserved; /* Required size of the reserved address space. */
+} MemoryMappingSizes;
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
@@ -41,6 +48,27 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ void *ShmemBase; /* start address of shared memory */
+ void *ShmemEnd; /* end+1 address of shared memory */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+ int segment_fd; /* fd for the backing anon file */
+ unsigned long seg_id; /* IPC key */
+ int shmem_segment; /* TODO: Do we really need it? */
+ Size shmem_size; /* Size of the actually used memory */
+ Size shmem_reserved; /* Size of the reserved mapping */
+ Pointer shmem; /* Pointer to the start of the mapped memory */
+ Pointer seg_addr; /* SysV shared memory for the header */
+} ShmemSegment;
+
+/* Number of available segments for anonymous memory mappings */
+#define NUM_MEMORY_MAPPINGS 6
+
+extern PGDLLIMPORT ShmemSegment Segments[NUM_MEMORY_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -85,10 +113,38 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(Size size,
+extern PGShmemHeader *PGSharedMemoryCreate(MemoryMappingSizes *mapping_sizes, int segment_id,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern const char *MappingName(int shmem_segment);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+void PrepareHugePages(void);
+
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, which size depends on
+ * NBuffers.
+ */
+
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 70a5b8b172c..c56712555f0 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -30,14 +30,25 @@ extern PGDLLIMPORT slock_t *ShmemLock;
typedef struct PGShmemHeader PGShmemHeader; /* avoid including
* storage/pg_shmem.h here */
extern void InitShmemAccess(PGShmemHeader *seghdr);
+extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr,
+ int shmem_segment);
extern void InitShmemAllocation(void);
+extern void InitShmemAllocationInSegment(int shmem_segment);
extern void *ShmemAlloc(Size size);
+extern void *ShmemAllocInSegment(Size size, int shmem_segment);
extern void *ShmemAllocNoError(Size size);
+extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment);
extern void InitShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
HASHCTL *infoP, int hash_flags);
+extern HTAB *ShmemInitHashInSegment(const char *name, long init_size,
+ long max_size, HASHCTL *infoP,
+ int hash_flags, int shmem_segment);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int shmem_segment);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
@@ -59,6 +70,7 @@ typedef struct
void *location; /* location in shared mem */
Size size; /* # bytes requested for the structure */
Size allocated_size; /* # bytes actually allocated */
+ int shmem_segment; /* segment in which the structure is allocated */
} ShmemIndexEnt;
#endif /* SHMEM_H */
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 7c52181cbcb..bd877df5f3b 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1765,14 +1765,22 @@ pg_shadow| SELECT pg_authid.rolname AS usename,
LEFT JOIN pg_db_role_setting s ON (((pg_authid.oid = s.setrole) AND (s.setdatabase = (0)::oid))))
WHERE pg_authid.rolcanlogin;
pg_shmem_allocations| SELECT name,
+ segment,
off,
size,
allocated_size
- FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size);
+ FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, segment, off, size, allocated_size);
pg_shmem_allocations_numa| SELECT name,
numa_node,
size
FROM pg_get_shmem_allocations_numa() pg_get_shmem_allocations_numa(name, numa_node, size);
+pg_shmem_segments| SELECT id,
+ name,
+ size,
+ freeoffset,
+ mapping_size,
+ mapping_reserved_size
+ FROM pg_get_shmem_segments() pg_get_shmem_segments(id, name, size, freeoffset, mapping_size, mapping_reserved_size);
pg_stat_activity| SELECT s.datid,
d.datname,
s.pid,
--
2.34.1
[text/x-csrc] mfdtruncate.c (2.1K, ../../CAExHW5sVxEwQsuzkgjjJQP9-XVe0H2njEVw1HxeYFdT7u7J+eQ@mail.gmail.com/6-mfdtruncate.c)
download | inline:
#define _GNU_SOURCE 1 /* See feature_test_macros(7) */
#include <errno.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mman.h>
#include <unistd.h>
#include <stdbool.h>
#include <fcntl.h>
#define MBSIZE(x) ((long) (x) * 1024 * 1024)
static void
resize_map_and_test(void *memaddr, int fd, size_t size, char *localmem)
{
int ret;
if (ftruncate(fd, size) < 0)
{
printf("ftruncate failed with errno %d on fd = %d and size = %ld\n", errno, fd, size);
exit(__LINE__);
}
ret = posix_fallocate(fd, 0, size);
if (ret != 0)
{
printf("fallocate failed with errno %d\n", ret);
exit(__LINE__);
}
memset(memaddr, 1, size);
if (memcmp(memaddr, localmem, size) != 0)
{
printf("mmap memory and local memory are not equal upto size = %ld\n", size);
exit(__LINE__);
}
printf("Check memory maps and enter (size %ld):", size);
getchar();
/* Causes segfault: memset(memaddr, 0, size + 1); */
}
int
main(int argc, char **argv)
{
int flags = MAP_NORESERVE;
void *memaddr;
pid_t pid = getpid();
char *localmem;
int fd = memfd_create("mmap_fd_exp", MFD_HUGETLB);
size_t maxsize = MBSIZE(400);
size_t size = MBSIZE(300);
printf("pid = %d\n", pid);
localmem = malloc(maxsize);
memset(localmem, 1, maxsize);
printf("Check initial memory maps and enter:");
getchar();
if (fd < 0)
{
printf("memfd_create failed with errno %d\n", errno);
exit(__LINE__);
}
memaddr = mmap(NULL, maxsize, PROT_WRITE | PROT_READ /* PROT_NONE */, MAP_SHARED | MAP_NORESERVE | MAP_HUGETLB, fd, 0);
if (memaddr == MAP_FAILED)
{
printf("mmap failed with error %m\n");
exit(__LINE__);
}
printf("Initial mapped file size: %ld\n", lseek(fd, 0, SEEK_END));
printf("Check memory maps and enter:");
getchar();
resize_map_and_test(memaddr, fd, size, localmem);
resize_map_and_test(memaddr, fd, MBSIZE(100), localmem);
/* causes a segmentation fault: memset(memaddr, 1, MBSIZE(200)); */
resize_map_and_test(memaddr, fd, MBSIZE(200), localmem);
resize_map_and_test(memaddr, fd, maxsize, localmem);
resize_map_and_test(memaddr, fd, MBSIZE(200), localmem);
close(fd);
}
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-01-28 13:19 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-06 09:25 ` Re: Changing shared_buffers without restart Bowen Shi <zxwsbg12138@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
0 siblings, 3 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2026-01-28 13:19 UTC (permalink / raw)
To: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; +Cc: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On Fri, Nov 14, 2025 at 5:23 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> Hi,
> PFA new patchset with some TODOs from previous email addressed:
>
> On Mon, Oct 13, 2025 at 9:28 PM Ashutosh Bapat
> <ashutosh.bapat.oss@gmail.com> wrote:
> > 1. New backends join while the synchronization is going on.
>
> Done. Explained the solution below in detail.
>
> > An existing backend exiting.
>
> Not tested specifically, but should work.
>
> > 2. Failure or crash in the backend which is executing pg_resize_buffer_pool()
>
> still a TODO
>
> > 3. Fix crashes in the tests.
>
> core regression passes, pg_buffercache regression tests pass and the
> tests for buffer resizing pass most of the time. So far I have seen
> two issues
> 1. An assertion from AIO worker - which happened only once and I
> couldn't reproduce again. Need to study interaction of AIO worker with
> buffer resizing.
> 2. checkpointer crashes - which is one of the TODOs listed below.
> 3. Also there's an shared memory id related failure, which I don't
> understand but happen more frequently than the first one. Need to look
> into that.
>
> > go through Tomas's detailed comments and address those
> > which still apply.
>
> Still a TODO. But since many of those patches are revised heavily, I
> think many of the comments may have been addressed, some may not apply
> anymore.
Thanks a lot Tomas for these review comments and summary of the
discussion as of your response. Sorry, it took so long to fully
respond to this email, but I wanted the patches to be in a reasonable
shape before responding. It's been a while since I have posted the
last patchset. Hence attaching patchset along with responses to your
comments as well as the earlier comments that you mentioned in your
response. Dmitry has responded to some parts of these emails already
but I am responding to all the comments in the context of the latest
implementation which has been revised heavily since his responses. The
latest design and UI is described in details in [1]. Please note that
that email also responds to some earlier comments from Andres and
others. I have not repeated those responses here.
The earlier patches did not build in EXEC_BACKEND mode. The attached
patches build in EXEC_BACKEND mode and also the regression tests pass
in that build. I haven't tried running whole test suite though.
>
> I agree it'd be useful to be able to resize shared buffers, without
> having to restart the instance (which is obviously very disruptive). So
> if we can make this work reliably, with reasonable trade offs (both on
> the backends, and also the risks/complexity introduced by the feature).
Agreed. That's the goal - to make it work reliably with reasonable
trade-offs. Adding some details to this goal as follows:
1. Performance
--------------
a. The impact of resizing on the performance of the concurrent
database operations should be reasonable and acceptable.
b. The code changes should not cause significant performance
degradation during normal operations
Palak Chaturvedi has already done some benchmarking. I am listing her
findings here
1. The code changes do not have any noticeable performance degradation
during normal operations.
2. Resizing itself is also reasonably fast and with very minimal,
almost unnoticeable, performance impact on the concurrent
transactions.
We will share the benchmarks once the code is cleaner and more complete stable.
2. Reliability
--------------
We are adding tests to make sure that resizing does not introduce any
crashes, hazards or data corruption. Some tests are already part of
the patch set.
>
> I'm far from an expert on mmap() and similar low-level stuff, but the
> current appproach (reserving a big chunk of shared memory and slicing
> it by mmap() into smaller segments) seems reasonable.
The latest patches, we reserve address space by mmap and manage memory
within that address space using file backed memory.
>
> But I'm getting a bit lost in how exactly this interacts with things
> like overcommit, system memory accounting / OOM killer and this sort of
> stuff. I went through the thread and it seems to me the reserve+map
> approach works OK in this regard (and the messages on linux-mm seem to
> confirm this). But this information is scattered over many messages and
> it's hard to say for sure, because some of this might be relevant for
> an earlier approach, or a subtly different variant of it.
Looking at the documentation and the messages on linux-mm, it seems
that the mmap with MAP_NORESERVE approach works well with overcommit,
system memory accounting and OOM killer. I plan to add tests verify
these aspects, so that any changes in the future do not cause any
regressions. I am looking for ways to write a small test program for
the same, which we can add to our test battery. But no success so far.
Any idea?
>
> A similar question is portability. The comments and commit messages
> seem to suggest most of this is linux-specific, and other platforms just
> don't have these capabilities. But there's a bunch of messages (mostly
> by Thomas Munro) that hint FreeBSD might be capable of this too, even if
> to some limited extent. And possibly even Windows/EXEC_BACKEND, although
> that seems much trickier.
>
> FWIW I think it's perfectly fine to only support resizing on selected
> platforms, especially considering Linux is the most widely used system
> for running Postgres. We still need to be able to build/run on other
> systems, of course. And maybe it'd be good to be able to disable this
> even on Linux, if that eliminates some overhead and/or risks for people
> who don't need the feature. Just a thought.
+1. Plan is to make it work on Linux first and disable it on the other
platforms. I think it's a good idea to be able to disable it at initdb
time say. But I am not yet sure if that's going to avoid any overhead
or risks. I haven't yet wrapped my head around the platform dependency
completely.
>
> Anyway, my main point is that this information is important, but very
> scattered over the thread. It's a bit foolish to expect everyone who
> wants to do a review to read the whole thread (which will inevitably
> grow longer over time), and assemble all these pieces again an again,
> following all the changes in the design etc. Few people will get over
> that hurdle, IMHO.
>
> So I think it'd be very helpful to write a README, explaining the
> currnent design/approach, and summarizing all these aspects in a single
> place. Including things like portability, interaction with the OS
> accounting, OOM killer, this kind of stuff. Some of this stuff may be
> already mentioned in code comments, but you it's hard to find those.
>
> Especially worth documenting are the states the processes need to go
> through (using the barriers), and the transacitons between them (i.e.
> what is allowed in each phase, what blocks can be visible, etc.).
>
That's a great idea. There is already a README in
src/backend/storage/buffer/, I have added a section of buffer pool
resizing. The section has pointers to relevant code comments. I will
revise this to be as self sufficient as possible as the patches
mature. The code managing the shared memory is scattered across
backend/portability, storage/ipc etc. It's slightly troublesome to
find one good place to add a README which explains everything.
>
> I'll go over some higher-level items first, and then over some comments
> for individual patches.
>
>
> 1) no user docs
>
> There are no user .sgml docs, and maybe it's time to write some,
> explaining how to use this thing - how to configure it, how to trigger
> the resizing, etc. It took me a while to realize I need to do ALTER
> SYSTEM + pg_reload_conf() to kick this off.
Latest patches have updated config.sgml and func-admin.sgml to
document the functionality. Please review them.
>
> It should also document the user-visible limitations, e.g. what activity
> is blocked during the resizing, etc.
>
We are still figuring this out. But we will document the limitations
in func-admin.sgml. Right now there are no limitations and my
intention is to keep the limitations as limited as possible.
>
> 2) pending GUC changes
>
> I'm somewhat skeptical about the GUC approach. I don't think it was
> designed with this kind of use case in mind, and so I think it's quite
> likely it won't be able to handle it well.
>
> For example, there's almost no validation of the values, so how do you
> ensure the new value makes sense? Because if it doesn't, it can easily
> crash the system (I've seen such crashes repeatedly, I'll get to that).
> Sure, you may do ALTER SYSTEM to set shared_buffers to nonsense and it
> won't start after restart/reboot, but crashing an instance is maybe a
> little bit more annoying.
>
> Let's say we did the ALTER SYSTEM + pg_reload_conf(), and it gets stuck
> waiting on something (can't evict a buffer or something). How do you
> cancel it, when the change is already written to the .auto.conf file?
> Can you simply do ALTER SYSTEM + pg_reload_conf() again?
>
> It also seems a bit strange that the "switch" gets to be be driven by a
> randomly selected backend (unless I'm misunderstanding this bit). It
> seems to be true for the buffer eviction during shrinking, at least.
>
> Perhaps this should be a separate utility command, or maybe even just
> a new ALTER SYSTEM variant? Or even just a function, similar to what
> the "online checksums" patch did, possibly combined with a bgworder
> (but probably not needed, there are no db-specific tasks to do).
>
As documented in func-admin.sgml, and config.sgml, resizing the buffer
pool requires ALTER SYSTEM + pg_reload_conf() followed by
pg_resize_shared_buffers(). There is no pending flag, no arbitrary
backend driving the process. The validation, failures are handled by
pg_resize_shared_buffers(). The backend where
pg_resize_shared_buffers() is called is responsible for driving the
resizing process. I think, the new implementation takes care of all
the concerns mentioned above.
>
> 3) max_available_memory
>
> Speaking of GUCs, I dislike how max_available_memory works. It seems a
> bit backwards to me. I mean, we're specifying shared_buffers (and some
> other parameters), and the system calculates the amount of shared memory
> needed. But the limit determines the total limit?
>
> I think the GUC should specify the maximum shared_buffers we want to
> allow, and then we'd work out the total to pre-allocate? Considering
> we're only allowing to resize shared_buffers, that should be pretty
> trivial. Yes, it might happen that the "total limit" happens to exceed
> the available memory or something, but we already have the problem
> with shared_buffers. Seems fine if we explain this in the docs, and
> perhaps print the calculated memory limit on start.
>
> In any case, we should not allow setting a value that ends up
> overflowing the internal reserved space. It's true we don't have a good
> way to do checks for GUcs, but it's a bit silly to crash because of
> hitting some non-obvious internal limit that we necessarily know about.
>
> Maybe this is a reason why GUC hooks are not a good way to set this.
Latest patches introduce max_shared_buffers which specifies the
maximum shared buffers allowed. max_available_memory, that some
previous patchsets had, has been removed. This GUC is mentioned in the
documentation changes.
One relatively small thing to think about is what do we call
shared_buffers GUC - it's not SIGHUP per say since it requires a
function to be called after the reload but it's not POSTMASTER either
since restart is not required. I think we need a new PGC_ for it but
haven't yet thought of a good name. Do you have any suggestions?
>
>
> 4) SHMEM_RESIZE_RATIO
>
> The SHMEM_RESIZE_RATIO thing seems a bit strange too. There's no way
> these ratios can make sense. For example, BLCKSZ is 8192 but the buffer
> descriptor is 64B. That's 128x difference, but the ratios says 0.6 and
> 0.1, so 6x. Sure, we'll actually allocate only the memory we need, and
> the rest is only "reserved".
>
> However, that just makes the max_available_memory a bit misleading,
> because you can't ever use it. You can use the 60% for shared buffers
> (which is not mentioned anywhere, and good luck not overflowing that,
> as it's never checked), but those smaller regions are guaranteed to be
> mostly unused. Unfortunate.
>
> And it's not just a matter of fixing those ratios, because then someone
> rebuilds with 32kB blocks and you're in the same situation.
>
> Moreover, all of the above is for mappings sized based on NBuffers. But
> if we allocate 10% for MAIN_SHMEM_SEGMENT, won't that be a problem the
> moment someone increases of max_connection, max_locks_per_transaction
> and possibly some other stuff?
>
+1. Fixed. max_shared_buffers only deals with the shared buffers.
>
> 5) no tests
>
> I mentioned no "user docs", but the patch has 0 tests too. Which seems
> a bit strange for a patch of this age.
>
> A really serious part of the patch series seems to be the coordination
> of processes when going through the phases, enforced by the barriers.
> This seems like a perfect match for testing using injection points, and
> I know we did something like this in the online checksums patch, which
> needs to coordinate processes in a similar way.
>
> But even just a simple TAP test that does a bunch of (random?) resizes
> while running a pgbench seem better than no tests. (That's what I did
> manually, and it crashed right away.)
>
> There's a lot more stuff to test here, I think. Idle sessions with
> buffers pinned by open cursors, multiple backends doing ALTER SYSTEM
> + pg_reload_conf concurrently, other kinds of failures.
>
+1. The latest patches already have a few tests but we are adding more
tests to cover different scenarios.
>
> 6) SIGBUS failures
>
> As mentioned, I did some simple tests with shrink/resize with a pgbench
> in the background, and it almost immediately crashed for me :-( With a
> SIGBUS, which I think is fairly rare on x86 (definitely much less common
> than e.g. SIGSEGV).
>
> An example backtrace attached.
>
There is a test which runs pgbench concurrently with resizing and it
fails much less frequently (1 in 20 or even lower). Quite likely the
crash you have seen is also fixed. Please let me know if the test
fails for you and if there is something that is in your test but not
that test.
>
> 7) EXEC_BACKEND, FreeBSD
>
> We clearly need to keep this working on systems without the necessary
> bits (so likely EXEC_BACKEND, FreeBSD etc.). But the builds currently
> fail in both cases, it seems.
>
> I think it's fine to not support resizing on every platform, then we'd
> never get it, but it still needs to build. It would be good to not have
> two very different code versions, one for resizing and one without it,
> though. I wonder if we can just have the "no-resize" use the same struct
> (with the segments/mapping, ...) and all that, but skipping the space
> reservation.
>
I have not thought fully through support on platforms other than
Linux. However, my current plan is to not support GUC
max_shared_buffers and the resizing function
pg_resize_shared_buffers() on platforms which do not support the
necessary bits. The code will still build and run on those platforms,
but resizing shared buffers without restart won't be possible.
>
> 8) monitoring
>
> So, let's say I start a resize of shared buffers. How will I know what
> it's currently doing, how much longer it might take, what it's waiting
> for, etc.? I think it'd be good to have progress monitoring, through
> the regular system view (e.g. pg_stat_shmem_resize_progress?).
>
When pg_resize_shared_buffers() finishes, user knows that the resizing
is finished. It usually takes a few seconds for the resizing to
finish. I am not sure whether we will need progress reporting. But
there might be cases where the resizing has to wait for something or
it may take longer for a given phase e.g. eviction. So we may require
progress reporting. Have added TODO in the code for the same.
>
> 10) what to do about stuck resize?
>
> AFAICS the resize can get stuck for various reasons, e.g. because it
> can't evict pinned buffers, possibly indefinitely. Not great, it's not
> clear to me if there's a way out (canceling the resize) after a timeout,
> or something like that? Not great to start an "online resize" only to
> get stuck with all activity blocked for indefinite amount of time, and
> get to restart anyway.
>
> Seems related to Thomas' message [2], but AFAICS the patch does not do
> anything about this yet, right? What's the plan here?
>
pg_resize_shared_buffers() is affected by timeouts like
statement_timeouts as well as query cancellation. However, current
implementation lacks graceful handling of these events. Another TODO.
>
> 11) preparatory actions?
>
> Even if it doesn't get stuck, some of the actions can take a while, like
> evicting dirty buffers before shrinking, etc. This is similar to what
> happens on restart, when the shutdown checkpoint can take a while, while
> the system is (partly) unavailable.
>
> The common mitigation is to do an explicit checkpoint right before the
> restart, to make the shutdown checkpoint cheap. Could we do something
> similar for the shrinking, e.g. flush buffers from the part to be
> removed before actually starting the resize?
>
Eviction is carried out before changes to the memory or shared buffer
metadata. If eviction fails, the resizing is rolled back. The system
continues to work with old buffer pool size. A test particularly
testing this aspect remains to be added.
>
> 12) does this affect e.g. fork() costs?
>
> I wonder if this affects the cost of fork() in some undesirable way?
> Could it make fork() more measurably more expensive?
>
A new backend needs to attach to extra shared memory segments, AFAIK
and detach those when exiting. But I don't think there's any other
extra work that a backend needs to do when starting or exiting.
Windows might be different. I haven't seen any noticeable difference
in pgbench performance which uses new connection for every
transaction. However, we haven't tried a benchmark with empty
transactions so that fork() and exit() are exercised at a higher
frequency. Another TODO.
>
> 13) resize "state" is all over the place
>
> For me, a big hurdle when reasoning about the resizing correctness is
> that there's quite a lot of distinct pieces tracking what the current
> "state" is. I mean, there's:
>
> - ShmemCtrl->NSharedBuffers
> - NBuffers
> - NBuffersOld
> - NBuffersPending
> - ... (I'm sure I missed something)
>
> There's no cohesive description how this fits together, it seems a bit
> "ad hoc". Could be correct, but I find it hard to reason about.
>
Next set of patches will consolidate the state only in two places
NBuffersPending and StrategyControl, which needs a new name. But even
in the attached patches it's consolidated in NBuffersPending,
ShmemCtrl and StrategyControl. There are instances of NBuffers which
will be replaced by variables from ShmemCtrl or StrategyControl.
>
> 14) interesting messages from the thread
>
> While reading through the thread, I noticed a couple messages that I
> think are still relevant:
>
> - I see Peter E posted some review in 2024/11 [3], but it seems his
> comments were mostly ignored. I agree with most of them.
Find detailed reply to that email at the end.
>
> - Robert mentioned a couple interesting failure scenarios in [4], not
> sure if all of this was handled. He howerver assumes pointers would
> not be stable (and that's something we should not allow, and the
> current approach works OK in this regard, I think). He also outlines
> how it'd happen in phases - this would be useful for the design README
> I think. It also reminds me the "phases" in the checksums patch.
>
stable pointers in the latest patch. The phases are explained in the
prologue of pg_resize_shared_buffers().
> - Robert asked [5] if Linux might abruptly break this, but I find that
> unlikely. We'd point out we rely on this, and they'd likely rethink.
> This would be made safer if this was specified by POSIX - taking that
> away once implemented seems way harder than for custom extensions.
> It's likely they'd not take away the feature without an alternative
> way to achieve the same effect, I think (yes, harder to maintain).
> Tom suggests [7] this is not in POSIX.
>
> - Matthias mentioned [6] similar flags on other operating systems. Could
> some of those be used to implement the same resizing?
Part of portability TODO.
>
> - Andres had an interesting comment about how overcommit interacts with
> MAP_NORESERVE. AFAIK it means we need the flag to not break overcommit
> accounting. There's also some comments about from linux-mm people [9].
>
New implementation uses MAP_NORESERVE. See my earlier response about overcommit.
> - There seem to be some issues with releasing memory backing a mapping
> with hugetlb [10]. With the fd (and truncating the file), this seems
> to release the memory, but it's linux-specific? But most of this stuff
> is specific to linux, it seems. So is this a problem? With this it
> should be working even for hugetlb ...
>
Right. mmap() + ftruncate(), instead of mmap() + mremap() allows us to
avoid going through postmaster, which makes the implementation much
simpler. And also support hugetlb. But it will be linux only.
> - It seems FreeBSD has MFD_HUGETLB [11], so maybe we could use this and
> make the hugetlb stuff work just like on Linux? Unclear. Also, I
> thought the mfd stuff is linux-specific ... or am I confused?
>
Portability TODO.
> - Andres objected to any approach without pointer stability, and I agree
> with that. If we can figure out such solution, of course.
Since we call mmap only once, the address of the mappping does not
change; it is always stable. So pointer stability is guaranteed.
>
> - Thomas asked [13] why we need to stop all the backends, instead of
> just waiting for them to acknowledge the new (smaller) NBuffers value
> and then let them continue. I also don't quite see why this should
> not work, and it'd limit the disruption when we have to wait for
> eviction of buffers pinned by paused cursors, etc.
>
Approach in the latest patches does not stop all backends. Backends
continue to work while resizing is in progress. They are synchronized
at certain points using barriers. The details of synchronization for
each subsystem and worker backend need to be worked out. We are
working on that.
>
>
> Now, some comments about the individual patches (some of this may be a
> bit redundant with the earlier points):
Since these patches have been heavily rewritten because we have
redesigned and reimplemented the resizing process, some of the
comments may not be relevant any more. I will try to answer the
underlying concerns wherever applicable.
>
>
> v5-0001-Process-config-reload-in-AIO-workers.patch
>
> 1) Hmmm, so which other workers may need such explicit handling? Do all
> other processes participate in procsignal stuff, or does anything
> need an explicit handling?
>
As Andres mentioned in [2], this is not needed. It's not part of the
latest patches.
>
> v5-0003-Introduce-pss_barrierReceivedGeneration.patch
>
> 1) Do we actually need this? Isn't it enough to just have two barriers?
> Or a barrier + condition variable, or something like that.
>
> 2) The comment talks about "coordinated way" when processing messages,
> but it's not very clear to me. It should explain what is needed and
> not possible with the current barrier code.
>
> 3) This very much reminds me what the online checksums patch needed to
> do, and we managed to do it using plain barriers. So why does this
> need this new thing? (No opinion on whether it's correct.)
>
This isn't required in the latest implementation. Removed from the
latest patchset.
>
> v5-0004-Allow-to-use-multiple-shared-memory-mappings.patch
>
> 1) "int shmem_segment" - wouldn't it be better to have a separate enum
> for this? I mean, we'll have a predefined list of segments, right?
>
+1. Andres suggested [2] to keep only two segments. So enum may be
superfluous, but may be good if we need more segments in future. TODO
for now.
> 2) typedef struct AnonymousMapping would deserve some comment
This structure no more exists. Instead we use MemoryMappingSizes to
hold the required sizes of shared memory segments.
>
> 3) ANON_MAPPINGS - Probably should be MAX_ANON_MAPPINGS? But we'll know
> how many we have, so why not to allocate exactly the right number?
> Or even just an array of structs, like in similar cases?
>
+1. Renamed as NUM_MEMORY_MAPPINGS and is used to declare and traverse
corresponding arrays.
> 4) static int next_free_segment = 0;
>
> We exactly know what segments we'll create and in which order, no? So
> why do we even bother with this next_free_segment thing? Can't we
> simply declare an array of AnonymousMapping elements, with all the
> elements, and then just walk it and calculate the sizes/pointers?
next_free_segment is removed from the latest patches as it's not needed.
>
> 5) I'm a bit confused about the segment/mapping difference. The patch
> seems to randomly mix those, or maybe I'm just confused. I mean,
> we are creating just shmem segment, and the pieces are mappings,
> right? So why do we index them by "shmem_segment"?
>
> Also, consider
>
> CreateAnonymousSegment(AnonymousMapping *mapping)
>
> so is that creating a segment or mapping? Or what's the difference?
>
> Or are we creating multiple segments, and I missed that? Or are there
> different "segment" concepts, or what?
>
> 6) There should probably be some sort of API wrapping the mappings, so
> that the various places don't need to mess with next_free_segments
> directly, etc. Perhaps PGSharedMemoryCreate() shouldn't do this, and
> should just pass size to CreateAnonymousSegment(), and that finding
> empty slot in Mappings, etc.? Not sure that'll work, but it's a bit
> error-prone if a struct is modified from multiple places like this.
Fixed this confusion in the latest patches. There are multiple
segments, each mapped to a different address space. Since we are using
mmap() + ftruncate(), the memory allocated for each segment is
controlled by a separate fds. If we use a single address space
reservation and carve multiple segments out of it, we can not resize
each segment separately. Just to clarify, we do not reserve a large
part of address space and then carve it into smaller segments.
>
> 7) We should remember which segments got to use huge pages and which
> did not. And we should make it optional for each segment. Although,
> maybe I'm just confused about the "segment" definition - if we only
> have one, that's where huge pages are applied.
>
> If we could have multiple segments for different segments (whatever
> that means), not sure what we'll report for cases when some segments
> get to use huge pages and others don't. Either because we don't want
> to use that for some segments, or because we happen to run out of
> the available huge pages.
This is an interesting idea. In the current implementation either we
use huge pages for all the segments or none of them. I think, per
segment huge page usage will be an add-on feature in the next version.
What do you think? Using separate segments for reserving separate
address spaces will make it easy to use huge pages for some and not
for others.
>
> 8) It seems PGSharedMemoryDetach got some significant changes, but the
> comment was not modified at all. I'd guess that means the comment is
> perhaps stale, or maybe there's something we should mention.
>
Done.
> 9) I doubt the Assert on GetConfigOption needs to be repeated for all
> segments (in CreateSharedMemoryAndSemaphores).
>
Done.
> 10) Why do we have the Mapping and Segments indexed in different ways?
> I mean, Mappings seem to be filled in FIFO (just grab the next free
> slot), while Segments are indexed by segment ID.
>
Fixed this confusion in the latest patches. Both are indexed by segment ID.
> 11) Actually, what's the difference between the contents of Mappings
> and Segments? Isn't that the same thing, indexed in the same way?
> Or could it be unified? Or are they conceptually different thing?
>
See explanation above.
> 12) I believe we'll have a predefined list of segments, with fixed IDs,
> so why not just have a MAX of those IDs as the capacity?
>
yes. Fixed in the latest patches.
> 13) Would it be good to have some checks on shmem_segment values? That
> it's valid with respect to defined segments, etc. An assert, maybe?
> What about some asserts on the Mapping/Segment elements? To check
> that the element is sensible, and that the arrays "match" (if we
> need both).
>
That's a good idea. Added Asserts to that effect.
> 14) Some of the lines got pretty long, e.g. in pg_get_shmem_allocations.
> I suggest we define some macros to make this shorter, or something
> like that.
>
Done.
> 15) I'd maybe rename ShmemSegment to PGShmemSegment, for consistency
> with PGShmemHeader?
Actually we track information about the shared memory segments at
multiple places. shmem.c has APIs similar to MemoryContext for shared
memory. The information there is consolidated into ShmemSegment
structure. pg_shmem.h has platform independent APIs to manage shared
memory segments and then each implementation has implementation
specific information. Some of that information is passed from parent
process (postmaster) to child process (postgresql backends). I have
created PGInhShmemSegment for the information that is passed from
postmaster to backends and then AnonymousShmemSegment structure for
tracking information about anonymous shared memory in sysv_shmem.c. I
don't like PGInhShmemSegment name, but I haven't figured out a better
name. I may rearrange these structures a bit more in the next patches.
>
> 16) Is MAIN_SHMEM_SEGMENT something we want to expose in a public header
> file? Seems very much like an internal thing, people should access
> it only through APIs ...
>
If we convert it to an enum, we can't avoid exposing it in a public
header file. In future, we may allow users to allocate memory in
specific segments. So I think it's fine to expose it.
>
> v5-0005-Address-space-reservation-for-shared-memory.patch
>
> 1) Shouldn't reserved_offset and huge_pages_on really be in the segment
> info? Or maybe even in mapping info? (again, maybe I'm confused
> about what these structs store)
I can't find reserved_offset in the latest patches. huge_pages_on is
now a global flag, either all segments use huge pages or none of them
do. As mentioned earlier, per segment huge page usage may be an add-on
separate feature.
>
> 2) CreateSharedMemoryAndSemaphores comment is rather light on what it
> does, considering it now reserves space and then carves is into
> segments.
>
The detailed comment is in PGSharedMemoryCreate(), which seems to be a
better place for explaining memory reservation etc. I just change the
comment here to mention plural shared memory segments and
differentiate between shared memory segments and the shared memory
structures. Please let me know if this looks good.
> 3) So ReserveAnonymousMemory is what makes decisions about huge pages,
> for the whole reserved space / all segments in it. That's a bit
> unfortunate with respect to the desirability of some segments
> benefiting from huge pages and others not. Maybe we should have two
> "reserved" areas, one with huge pages, one without?
>
See my responses above about mapping vs segment and per segment huge page usage.
> I guess we don't want too many segments, because that might make
> fork() more expensive, etc. Just guessing, though. Also, how would
> this work with threading?
See my earlier response about fork() performance and also limiting
number of segments to just 2. If we do see fork() is getting
expensive, we can limit the number of segments to 1.
>
> 4) Any particular reason to define max_available_memory as
> GUC_UNIT_BLOCKS and not GUC_UNIT_MB? Of course, if we change this
> to have "max shared buffers limit" then it'd make sense to use
> blocks, but "total limit" is not in blocks.
>
Right. See response about max_shared_buffers above.
> 5) The general approach seems sound to me, but I'm not expert on this.
> I wonder how portable this behavior is. I mean, will it work on other
> Unix systems / Windows? Is it POSIX or Linux extension?
See my earlier response about portability.
>
> 6) It might be a good idea to have Assert procedures to chech mappings
> and segments (that it doesn't overflow reserved space, etc.). It
> took me ages to realize I can change shared_buffers to >60% of the
> limit, it'll happily oblige and then just crash with OOM when
> calling mprotect().
>
This shouldn't happen with the latest patches. There are tests for the
same. Please let me know if you still face the issue again in your
tests.
>
> v5-0006-Introduce-multiple-shmem-segments-for-shared-buff.patch
>
> 1) I suspect the SHMEM_RESIZE_RATIO is the wrong direction, because it
> entirely ignores relationships between the parts. See the earlier
> comment about this.
>
> 2) In fact, what happens if the user tries to resize to a value that is
> too large for one of the segments? How would the system know before
> starting the resize (and failing)?
>
+1. No SHMEM_RESIZE_RATIO in the latest patches. Reservation is
entirely based on max_shared_buffers.
> 3) It seems wrong to modify the BufferManagerShmemSize like this. It's
> probably better to have a "...SegmentSize" function for individual
> segments, and let BufferManagerShmemSize() to still return a sum of
> all segments.
>
I have reimplemented CalculateShmemSize() in the latest patches. Please review.
> 4) I think MaxAvailableMemory is the wrong abstraction, because that's
> not what people specify. See earlier comment.
>
+1. No MaxAvailableMemory in the latest patches.
> 5) Let's say we change the shared memory size (ALTER SYSTEM), trigger
> the config reload (pg_reload_conf). But then we find that we can't
> actually shrink the buffers, for some unpredictable reason (e.g.
> there's pinned buffers). How do we "undo" the change? We can't
> really undo the ALTER SYSTEM, that's already written in the .conf
> and we don't know the old value, IIRC. Is it reasonable to start
> killing backends from the assign_hook or something? Seems weird.
>
The buffer pool resizing is done using the function
pg_resize_shared_buffers(), which does not need rolling back the
config value. As mentioned earlier, appropriate error handling is
still a TBD.
>
> v5-0007-Allow-to-resize-shared-memory-without-restart.patch
>
> 1) Why would AdjustShmemSize be needed? Isn't that a sign of a bug
> somewhere in the resizing?
>
It has been replaced by BufferManagerShmemResize() in the latest
patches. BufferManagerShmemResize() only resizes the buffer manager
related shared memory segments and data structures.
> 2) Isn't the pg_memory_barrier() in CoordinateShmemResize a bit weird?
> Why is it needed, exactly? If it's to flush stuff for processes
> consuming EmitProcSignalBarrier, it's that too late? What if a
> process consumes the barrier between the emit and memory barrier?
>
> 3) WaitOnShmemBarrier seem a bit under-documented.
>
Both of these functions are not required in the new implementation.
Removed from the latest patches.
> 4) Is this actually adding buffers to the freelist? I see buf_init only
> links the new buffers by seeting freeNext, but where are the new
> buffers added to the existing freelist?
>
Freelist does not exist anymore.
> 5) The issue with a new backend seeing an old NBuffers value reminds me
> of the "support enabling checksums online" thread, where we ran into
> similar race conditions. See message [1], the part about race #2
> (the other race might be relevant too, not sure). It's been a while,
> but I think our conclusion ini that thread was that the "best" fix
> would be to change the order of steps in InitPostgres(), i.e. setup
> the ProcSignal stuff first, and only then "copy" the NBuffers value.
> And handle the possibility that we receive a "duplicate" barriers.
>
I plan to remove NBuffers entirely and instead use shared memory
variables so that we don't have to worry about backends inheriting
stale values from Postmaster or even involving Postmaster in the
resizing process. This is being worked upon and will be available in a
future patchset.
> 6) In fact, the online checksums thread seems like a possible source of
> inspiration for some of the issues, because it needs to do similar
> stuff (e.g. make sure all backends follow steps in a synchronized
> way, etc.). And it didn't need new types of Barrier to do that.
>
Right. New implementation does not require any changes to the barrier
implementation.
> 7) Also, this seems like a perfect match for testing using injection
> points. In fact, there's not a single test in the whole patch series.
> Or a single line of .sgml docs, for that matter. It took me a while
> to realize I'm supposed to change the size by ALTER SYSTEM + reload
> the config.
>
There are some tests in the latest patches. More tests are being added.
>
> v5-0008-Support-shrinking-shared-buffers.patch
>
> 1) Why is ShmemCtrl->evictor_pid reset in AnonymousShmemResize? Isn't
> there a place starting it and waiting for it to complete? Why
> shouldn't it do EvictExtraBuffers itself?
>
evictor_pid is not used in the latest patches. Eviction is done in
pg_resize_shared_buffers() itself by the backend which executes that
function.
> 2) Isn't the change to BufferManagerShmemInit wrong? How do we know the
> last buffer is still at the end of the freelist? Seems unlikely.
>
No freelist anymore.
> 3) Seems a bit strange to do it from a random backend. Shouldn't it
> be the responsibility of a process like checkpointer/bgwriter, or
> maybe a dedicated dynamic bgworker? Can we even rely on a backend
> to be available?
>
The backend which executes pg_resize_shared_buffers() coordinates the
resizing itself. I have not seen a need for it to use another
background worker yet.
> 4) Unsolved issues with buffers pinned for a long time. Could be an
> issue if the buffer is pinned indefinitely (e.g. cursor in idle
> connection), and the resizing blocks some activity (new connections
> or stuff like that).
>
Resizing stops immediately after it encounters a pinned buffer. More
sophisticated handling of such situations is a TBD and mostly v2.
> 5) Funny that "AI suggests" something, but doesn't the block fail to
> reset nextVictimBuffer of the clocksweep? It may point to a buffer
> we're removing, and it'll be invalid, no?
>
TODO:
> 6) It's not clear to me in what situations this triggers (in the call
> to BufferManagerShmemInit)
>
> if (FirstBufferToInit < NBuffers) ...
>
That code does not exist in the latest patches.
>
> v5-0009-Reinitialize-StrategyControl-after-resizing-buffe.patch
>
> 1) IMHO this should be included in the earlier resize/shrink patches,
> I don't see a reason to keep it separate (assuming this is the
> correct way, and the "init" is not).
>
Right. Merged into the resizing patch in the latest patches.
> 2) Doesn't StrategyPurgeFreeList already do some of this for the case
> of shrinking memory?
No freelist anymore.
>
> 3) Not great adding a bunch of static variables to bufmgr.c. Why do we
> need to make "everything" static global? Isn't it enough to make
> only the "valid" flag global? The rest can stay local, no?
>
> If everything needs to be global for some reason, could we at least
> make it a struct, to group the fields, not just separate random
> variables? And maybe at the top, not half-way throught the file?
>
> 4) Isn't the name BgBufferSyncAdjust misleading? It's not adjusting
> anything, it's just invalidating the info about past runs.
Being worked upon.
>
> 5) I don't quite understand why BufferSync needs to do the dance with
> delay_shmem_resize. I mean, we certainly should not run BufferSync
> from the code that resizes buffers, right? Certainly not after the
> eviction, from the part that actually rebuilds shmem structs etc.
> So perhaps something could trigger resize while we're running the
> BufferSync()? Isn't that a bit strange? If this flag is needed, it
> seems more like a band-aid for some issue in the architecture.
>
> 6) Also, why should it be fine to get into situation that some of the
> buffers might not be valid, during shrinking? I mean, why should
> this check (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers).
> It seems better to ensure we never get into "sync" in a way that
> might lead some of the buffers invalid. Seems way too lowlevel to
> care about whether resize is happening.
>
> 7) I don't understand the new condition for "Execute the LRU scan".
> Won't this stop LRU scan even in cases when we want it to happen?
> Don't we want to scan the buffers in the remaining part (after
> shrinking), for example? Also, we already checked this shmem flag at
> the beginning of the function - sure, it could change (if some other
> process modifies it), but does that make sense? Wouldn't it cause
> problems if it can change at an arbitrary point while running the
> BufferSync? IMHO just another sign it may not make sense to allow
> this, i.e. buffer sync should not run during the "actual" resize.
>
Resizing runs for a few seconds. With default BufferSync frequency of
200ms, there will be a few cycles of BufferSync during resizing. If we
don't let BufferSync run during resizing, we might have more dirty
buffers to tackle post-resizing. I think it will be wise to let it run
during resizing in the portion of the buffer pool which is not being
removed. Working on that.
>
> v5-0010-Additional-validation-for-buffer-in-the-ring.patch
>
> 1) So the problem is we might create a ring before shrinking shared
> buffers, and then GetBufferFromRing will see bogus buffers? OK, but
> we should be more careful with these checks, otherwise we'll miss
> real issues when we incorrectly get an invalid buffer. Can't the
> backends do this only when they for sure know we did shrink the
> shared buffers? Or maybe even handle that during the barrier?
>
AFAIK, the buffer rings are inside the Scan nodes, which are not
accessible globally or to the barrier handling functions. So we can't
purge the rings only once during or after resizing. If we already have
a mechanism to access rings globally and hence in the barrier code, or
we can develop such a mechanism, we can purge the rings only once.
Please let me know if you have any ideas on this.
> 2) IMHO a sign there's the "transitions" between different NBuffers
> values may not be clear enough, and we're allowing stuff to happen
> in the "blurry" area. I think that's likely to cause bugs (it did
> cause issues for the online checksums patch, I think).
>
NBuffers is going to be replaced by shared memory variables. The
current code relies on buffer pool size being static for the entire
duration of the server. We are now changing that assumption. So, yes
there will be some "blurry" areas and bugs to tackle. We are
developing tests to uncover bugs and fix them.
Here's reply to email by Peter E [3].
On Tue, Nov 19, 2024 at 6:27 PM Peter Eisentraut <peter@eisentraut.org> wrote:
>
> On 18.10.24 21:21, Dmitry Dolgov wrote:
> > v1-0001-Allow-to-use-multiple-shared-memory-mappings.patch
> >
> > Preparation, introduces the possibility to work with many shmem mappings. To
> > make it less invasive, I've duplicated the shmem API to extend it with the
> > shmem_slot argument, while redirecting the original API to it. There are
> > probably better ways of doing that, I'm open for suggestions.
>
> After studying this a bit, I tend to think you should just change the
> existing APIs in place. So for example,
>
> void *ShmemAlloc(Size size);
>
> becomes
>
> void *ShmemAlloc(int shmem_slot, Size size);
>
> There aren't that many callers, and all these duplicated interfaces
> almost add more new code than they save.
>
> It might be worth making exceptions for interfaces that are likely to be
> used by extensions. For example, I see pg_stat_statements using
> ShmemInitStruct() and ShmemInitHash(). But that seems to be it. Are
> there any other examples out there? Maybe there are many more that I
> don't see right now. But at least for the initialization functions, it
> doesn't seem worth it to preserve the existing interfaces exactly.
>
> In any case, I think the slot number should be the first argument. This
> matches how MemoryContextAlloc() or also talloc() work.
>
Fixed. Please take a look at shmem.c changes.
> (Now here is an idea: Could these just be memory contexts? Instead of
> making six shared memory slots, could you make six memory contexts with
> a special shared memory type. And ShmemAlloc becomes the allocation
> function, etc.?)
I don't see a need to do that right now. But will revisit this idea.
>
> I noticed the existing code made inconsistent use of PGShmemHeader * vs.
> void *, which also bled into your patch. I made the attached little
> patch to clean that up a bit.
>
The latest patches are rebased on top of your commit.
> I suggest splitting the struct ShmemSegment into one struct for the
> three memory addresses and a separate array just for the slock_t's. The
> former struct can then stay private in storage/ipc/shmem.c, only the
> locks need to be exported.
> Also, maybe some of this should be declared in storage/shmem.h rather
> than in storage/pg_shmem.h. We have the existing ShmemLock in there, so
> it would be a bit confusing to have the per-segment locks elsewhere.
>
I have described the new structures above. Please review.
>
> Maybe rename ANON_MAPPINGS to something like NUM_ANON_MAPPINGS.
Done.
>
> > v1-0003-Introduce-multiple-shmem-slots-for-shared-buffers.patch
> >
> > Splits shared_buffers into multiple slots, moving out structures that depend on
> > NBuffers into separate mappings. There are two large gaps here:
> >
> > * Shmem size calculation for those mappings is not correct yet, it includes too
> > many other things (no particular issues here, just haven't had time).
> > * It makes hardcoded assumptions about what is the upper limit for resizing,
> > which is currently low purely for experiments. Ideally there should be a new
> > configuration option to specify the total available memory, which would be a
> > base for subsequent calculations.
>
> Yes, I imagine a shared_buffers_hard_limit setting. We could maybe
> default that to the total available memory, but it would also be good to
> be able to specify it directly, for testing.
>
Latest patches introduce max_shared_buffers instead.
>
> > v1-0005-Use-anonymous-files-to-back-shared-memory-segment.patch
> >
> > Allows an anonyous file to back a shared mapping. This makes certain things
> > easier, e.g. mappings visual representation, and gives an fd for possible
> > future customizations.
>
> I think this could be a useful patch just by itself, without the rest of
> the series, because of
>
> > * By default, Linux will not add file-backed shared mappings into a
> > core dump, making it more convenient to work with them in PostgreSQL:
> > no more huge dumps to process.
>
> This could be significant operational benefit.
>
> When you say "by default", is this adjustable? Does someone actually
> want the whole shared memory in their core file? (If it's adjustable,
> is it also adjustable for anonymous mappings?)
'man core' mentions that /proc/[PID]/coredump_filter file controls the
memory segments written to the core file. The value in the file is a
bitmask of memory mapping types. The bits correspond to the flags
passed to mmap(). This file can be written to by the program through
/proc/self/coredump_filter or externally by something like echo >
/proc/PID/coredump_filter. A boot time setting becomes default for all
the programs. A child process inherits the setting from parent through
fork() and execve() does not change it. So it looks like DBAs should
be able to control what gets dumped in postgresql core dumps by using
any of those options. If we use file backed shared memory as the
patches do, DBAs may have to change their default setting, which may
not have considered filed backed shared memory. We will need to
highlight this in the release notes. I have added a TODO for the same
in the code for now. We will add this note in the commit message or
update relevant document as the patches get committable.
The man page says that bits 0, 1, 4 are set by default (anonymous
private and shared mappings and ELF headers). But on my machine I see
that the default is different. It's possible that different
installations use different configurations.
>
> I'm wondering about this change:
>
> -#define PG_MMAP_FLAGS
> (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
> +#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
>
> It looks like this would affect all mmap() calls, not only the one
> you're changing. But that's the only one that uses this macro! I don't
> understand why we need this; I don't see anything in the commit log
> about this ever being used for any portability. I think we should just
> get rid of it and have mmap() use the right flags directly.
This is committed as c100340729b66dc46d4f9d68a794957bf2c468d8.
>
> I see that FreeBSD has a memfd_create() function. Might be worth a try.
> Obviously, this whole thing needs a configure test for memfd_create()
> anyway.
We are using memfd_create in the latest patches. I will look into
creating a configure test for memfd_create(). I have added a TODO
comment for the same.
>
> I see that memfd_create() has a MFD_HUGETLB flag. It's not very clear
> how that interacts with the MAP_HUGETLB flag for mmap(). Do you need to
> specify both of them if you want huge pages?
We need both. I used toy program attached to [0] to verify that. Maybe
we could use similar program for configure test.
There are three important areas of puzzle that I will work on next:
1. Synchronization: The coordinator sends Proc signal barriers to all
the concurrent backends during resizing. The code which deals with the
buffer pool needs to react to these barriers and may need to adjust
its course of action. Checkpointer, background writer, new buffer
allocation, code scanning the buffer pool sequentially are a few
examples that need to react to the barrier code.
2. Graceful handling of errors and failures during resizing.
3. Portability: A config test to check whether a platform has the APIs
need to support this feature and enable the feature in the
corresponding build. On other platforms build succeeds, regression
runs but this feature is disabled.
All the TODOs that mentioned above and those not covered by the above
three points are added to the code. One of them being to reduce the
number of segments to just two (main and shared buffer blocks). I plan
to address them as the feature matures.
[0] https://www.postgresql.org/message-id/CAExHW5sVxEwQsuzkgjjJQP9-XVe0H2njEVw1HxeYFdT7u7J%2BeQ%40mail.g...
[1] https://www.postgresql.org/message-id/CAExHW5sOu8%2B9h6t7jsA5jVcQ--N-LCtjkPnCw%2BrpoN0ovT6PHg%40mail...
[2] https://www.postgresql.org/message-id/qltuzcdxapofdtb5mrd4em3bzu2qiwhp3cdwdsosmn7rhrtn4u%40yaogvphfw...
[3] https://www.postgresql.org/message-id/12add41a-7625-4639-a394-a5563e349322%40eisentraut.org
[4] https://www.postgresql.org/message-id/CAExHW5vTWABxuM5fbQcFkGuTLwaxuZDEE2vtx2WuMUWk6JnF4g%40mail.gma...
[5] https://www.postgresql.org/message-id/CAEze2WiMkmXUWg10y%2B_oGhJzXirZbYHB5bw0%3DVWte%2BYHwSBa%3DA%40...
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] v20260128-0001-Add-a-view-to-read-contents-of-shared-buff.patch (14.8K, ../../CAExHW5s8s=UhjqNa_Tz1PFCRLzt3=5nvd5vD1wFdKWMQCmFySQ@mail.gmail.com/2-v20260128-0001-Add-a-view-to-read-contents-of-shared-buff.patch)
download | inline diff:
From 42ed43e5ad572cf72b44a38ea3a511ca606d8687 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 25 Aug 2025 19:23:50 +0530
Subject: [PATCH v20260128 1/5] Add a view to read contents of shared buffer
lookup table
The view exposes the contents of the shared buffer lookup table for
debugging, testing and investigation.
This helped me in debugging issues where the buffer descriptor array and
buffer lookup table were out of sync; either the buffer lookup table had
a mapping page->buffer which wasn't present in the buffer descriptor
array or a page in the buffer descriptor array didn't have corresponding
entry in the buffer lookup table. pg_buffercache doesn't help with those
kind of issues. Also doing that under the debugger in very painful.
I intend to keep this patch while the rest of the code matures. If it is
found useful as a debugging tool, we may consider make it committable
and commit it.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../expected/pg_buffercache.out | 39 ++++++++
.../pg_buffercache--1.5--1.6.sql | 24 +++++
contrib/pg_buffercache/pg_buffercache_pages.c | 18 ++++
contrib/pg_buffercache/sql/pg_buffercache.sql | 20 +++++
doc/src/sgml/system-views.sgml | 89 +++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 58 ++++++++++++
src/include/storage/buf_internals.h | 2 +
7 files changed, 250 insertions(+)
diff --git a/contrib/pg_buffercache/expected/pg_buffercache.out b/contrib/pg_buffercache/expected/pg_buffercache.out
index 886dea770f6..f0df2d2d3bc 100644
--- a/contrib/pg_buffercache/expected/pg_buffercache.out
+++ b/contrib/pg_buffercache/expected/pg_buffercache.out
@@ -33,6 +33,26 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
t
(1 row)
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -46,6 +66,10 @@ SELECT * FROM pg_buffercache_summary();
ERROR: permission denied for function pg_buffercache_summary
SELECT * FROM pg_buffercache_usage_counts();
ERROR: permission denied for function pg_buffercache_usage_counts
+SELECT * FROM pg_buffercache_lookup_table_entries();
+ERROR: permission denied for function pg_buffercache_lookup_table_entries
+SELECT * FROM pg_buffercache_lookup_table;
+ERROR: permission denied for view pg_buffercache_lookup_table
RESET role;
-- Check that pg_monitor is allowed to query view / function
SET ROLE pg_monitor;
@@ -73,6 +97,21 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts();
t
(1 row)
+RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
RESET role;
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
index 458f054a691..9bf58567878 100644
--- a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
+++ b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
@@ -44,3 +44,27 @@ CREATE FUNCTION pg_buffercache_evict_all(
OUT buffers_skipped int4)
AS 'MODULE_PATHNAME', 'pg_buffercache_evict_all'
LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Add the buffer lookup table function
+CREATE FUNCTION pg_buffercache_lookup_table_entries(
+ OUT tablespace oid,
+ OUT database oid,
+ OUT relfilenode oid,
+ OUT forknum int2,
+ OUT blocknum int8,
+ OUT bufferid int4)
+RETURNS SETOF record
+AS 'MODULE_PATHNAME', 'pg_buffercache_lookup_table_entries'
+LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Create a view for convenient access.
+CREATE VIEW pg_buffercache_lookup_table AS
+ SELECT * FROM pg_buffercache_lookup_table_entries();
+
+-- Don't want these to be available to public.
+REVOKE ALL ON FUNCTION pg_buffercache_lookup_table_entries() FROM PUBLIC;
+REVOKE ALL ON pg_buffercache_lookup_table FROM PUBLIC;
+
+-- Grant access to monitoring role.
+GRANT EXECUTE ON FUNCTION pg_buffercache_lookup_table_entries() TO pg_read_all_stats;
+GRANT SELECT ON pg_buffercache_lookup_table TO pg_read_all_stats;
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 89b86855243..f60f797a9b4 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -16,6 +16,7 @@
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "utils/rel.h"
+#include "utils/tuplestore.h"
#define NUM_BUFFERCACHE_PAGES_MIN_ELEM 8
@@ -107,6 +108,7 @@ PG_FUNCTION_INFO_V1(pg_buffercache_evict_all);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_relation);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_all);
+PG_FUNCTION_INFO_V1(pg_buffercache_lookup_table_entries);
/* Only need to touch memory once per backend process lifetime */
@@ -958,3 +960,19 @@ pg_buffercache_mark_dirty_all(PG_FUNCTION_ARGS)
PG_RETURN_DATUM(result);
}
+
+/*
+ * Return lookup table content as a set of records.
+ */
+Datum
+pg_buffercache_lookup_table_entries(PG_FUNCTION_ARGS)
+{
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* Fill the tuplestore */
+ BufTableGetContents(rsinfo->setResult, rsinfo->setDesc);
+
+ return (Datum) 0;
+}
diff --git a/contrib/pg_buffercache/sql/pg_buffercache.sql b/contrib/pg_buffercache/sql/pg_buffercache.sql
index 127d604905c..22e255c9721 100644
--- a/contrib/pg_buffercache/sql/pg_buffercache.sql
+++ b/contrib/pg_buffercache/sql/pg_buffercache.sql
@@ -18,6 +18,18 @@ from pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -26,6 +38,8 @@ SELECT * FROM pg_buffercache_os_pages;
SELECT * FROM pg_buffercache_pages() AS p (wrong int);
SELECT * FROM pg_buffercache_summary();
SELECT * FROM pg_buffercache_usage_counts();
+SELECT * FROM pg_buffercache_lookup_table_entries();
+SELECT * FROM pg_buffercache_lookup_table;
RESET role;
-- Check that pg_monitor is allowed to query view / function
@@ -36,6 +50,12 @@ SELECT buffers_used + buffers_unused > 0 FROM pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts();
RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+RESET role;
+
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 8b4abef8c68..c5683068470 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -71,6 +71,11 @@
<entry>backend memory contexts</entry>
</row>
+ <row>
+ <entry><link linkend="view-pg-buffer-lookup-table"><structname>pg_buffer_lookup_table</structname></link></entry>
+ <entry>shared buffer lookup table</entry>
+ </row>
+
<row>
<entry><link linkend="view-pg-config"><structname>pg_config</structname></link></entry>
<entry>compile-time configuration parameters</entry>
@@ -929,6 +934,90 @@ AND c1.path[c2.level] = c2.path[c2.level];
</para>
</sect1>
+ <sect1 id="view-pg-buffer-lookup-table">
+ <title><structname>pg_buffer_lookup_table</structname></title>
+ <indexterm>
+ <primary>pg_buffer_lookup_table</primary>
+ </indexterm>
+ <para>
+ The <structname>pg_buffer_lookup_table</structname> view exposes the current
+ contents of the shared buffer lookup table. Each row represents an entry in
+ the lookup table mapping a relation page to the ID of buffer in which it is
+ cached. The shared buffer lookup table is locked for a short duration while
+ reading so as to ensure consistency. This may affect performance if this view
+ is queried very frequently.
+ </para>
+ <table id="pg-buffer-lookup-table-view" xreflabel="pg_buffer_lookup_table">
+ <title><structname>pg_buffer_lookup_table</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>tablespace</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the tablespace containing the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>database</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database containing the relation (zero for shared relations)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relfilenode</structfield> <type>oid</type>
+ </para>
+ <para>
+ relfilenode identifying the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>forknum</structfield> <type>int2</type>
+ </para>
+ <para>
+ Fork number within the relation (see <xref linkend="storage-file-layout"/>)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>blocknum</structfield> <type>int8</type>
+ </para>
+ <para>
+ Block number within the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>bufferid</structfield> <type>int4</type>
+ </para>
+ <para>
+ ID of the buffer caching the page
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ <para>
+ Access to this view is restricted to members of the
+ <literal>pg_read_all_stats</literal> role by default.
+ </para>
+ </sect1>
+
<sect1 id="view-pg-config">
<title><structname>pg_config</structname></title>
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 23d85fd32e2..5089c7322f3 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -21,7 +21,12 @@
*/
#include "postgres.h"
+#include "fmgr.h"
+#include "funcapi.h"
#include "storage/buf_internals.h"
+#include "storage/lwlock.h"
+#include "utils/rel.h"
+#include "utils/builtins.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -159,3 +164,56 @@ BufTableDelete(BufferTag *tagPtr, uint32 hashcode)
if (!result) /* shouldn't happen */
elog(ERROR, "shared buffer hash table corrupted");
}
+
+/*
+ * BufTableGetContents
+ * Fill the given tuplestore with contents of the shared buffer lookup table
+ *
+ * This function is used by pg_buffercache extension to expose buffer lookup
+ * table contents via SQL. The caller is responsible for setting up the
+ * tuplestore and result set info.
+ */
+void
+BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc)
+{
+/* Expected number of attributes of the buffer lookup table entry. */
+#define BUFTABLE_CONTENTS_COLS 6
+
+ HASH_SEQ_STATUS hstat;
+ BufferLookupEnt *ent;
+ Datum values[BUFTABLE_CONTENTS_COLS];
+ bool nulls[BUFTABLE_CONTENTS_COLS];
+ int i;
+
+ memset(nulls, 0, sizeof(nulls));
+
+ Assert(tupdesc->natts == BUFTABLE_CONTENTS_COLS);
+
+ /*
+ * Lock all buffer mapping partitions to ensure a consistent view of the
+ * hash table during the scan. Must grab LWLocks in partition-number order
+ * to avoid LWLock deadlock.
+ */
+ for (i = 0; i < NUM_BUFFER_PARTITIONS; i++)
+ LWLockAcquire(BufMappingPartitionLockByIndex(i), LW_SHARED);
+
+ hash_seq_init(&hstat, SharedBufHash);
+ while ((ent = (BufferLookupEnt *) hash_seq_search(&hstat)) != NULL)
+ {
+ values[0] = ObjectIdGetDatum(ent->key.spcOid);
+ values[1] = ObjectIdGetDatum(ent->key.dbOid);
+ values[2] = ObjectIdGetDatum(ent->key.relNumber);
+ values[3] = ObjectIdGetDatum(ent->key.forkNum);
+ values[4] = Int64GetDatum(ent->key.blockNum);
+ values[5] = Int32GetDatum(ent->id);
+
+ tuplestore_putvalues(tupstore, tupdesc, values, nulls);
+ }
+
+ /*
+ * Release all buffer mapping partition locks in the reverse order so as
+ * to avoid LWLock deadlock.
+ */
+ for (i = NUM_BUFFER_PARTITIONS - 1; i >= 0; i--)
+ LWLockRelease(BufMappingPartitionLockByIndex(i));
+}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 27f12502d19..f4e9e703b8b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -29,6 +29,7 @@
#include "storage/spin.h"
#include "utils/relcache.h"
#include "utils/resowner.h"
+#include "utils/tuplestore.h"
/*
* Buffer state is a single 64-bit variable where following data is combined.
@@ -581,6 +582,7 @@ extern uint32 BufTableHashCode(BufferTag *tagPtr);
extern int BufTableLookup(BufferTag *tagPtr, uint32 hashcode);
extern int BufTableInsert(BufferTag *tagPtr, uint32 hashcode, int buf_id);
extern void BufTableDelete(BufferTag *tagPtr, uint32 hashcode);
+extern void BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc);
/* localbuf.c */
extern bool PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount);
base-commit: e3094679b9835fed2ea5c7d7877e8ac8e7554d33
--
2.34.1
[text/x-patch] v20260128-0003-Allow-to-resize-shared-memory-without-rest.patch (151.3K, ../../CAExHW5s8s=UhjqNa_Tz1PFCRLzt3=5nvd5vD1wFdKWMQCmFySQ@mail.gmail.com/3-v20260128-0003-Allow-to-resize-shared-memory-without-rest.patch)
download | inline diff:
From 861cd9d911a28a7a44c57bfe08d14fa5d83c0069 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 14:16:55 +0200
Subject: [PATCH v20260128 3/5] Allow to resize shared memory without restart
shared_buffers is now PGC_SIGHUP instead of PGC_POSTMASTER. The value
of this GUC is saved in NBuffersPending instead of NBuffers. When the
server starts, the shared memory size is estimated and the memory is
allocated using NBuffersPending.
When a server is running, the new value of GUC (set using ALTER SYSTEM
... SET shared_buffers = ...; followed by SELECT pg_reload_conf()) does
not come into effect immediately. Instead a function
pg_resize_shared_buffers() is used to resize the buffer pool. The
function uses the current value of GUC in the backends where it is
executed. The function also coordinates the buffer resizing
synchronization across backends.
SHOW shared_buffers now shows the current size of the shared buffer pool
but it also shows pending size of shared buffers, if any.
A new GUC max_shared_buffers is introduced to control the maximum value
of shared_buffers that can be set. By default it is 0. When explicitly
set it needs to be higher than 'shared_buffers'. When max_shared_buffers
is set to 0, it assumes the same value as GUC shared_buffers. This GUC
determines the size of address space reserved for future buffer pool
sizes and the size of buffer look up table.
TBD: Describe the protocol used by pg_resize_shared_buffers() to
synchronize buffer resizing operation with other backends.
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
If a buffer being evicted is pinned, we abort the resizing operation.
There are other alternative which are not implemented in the current
patches 1. to wait for the pinned buffer to get unpinned, 2. the backend
is killed or it itself cancels the query or 3. rollback the operation.
Note that option 1 and 2 would require the pinning related local and
shared records to be accessed. But we need infrastructure to do either
of this right now.
So far the buffer pool metdata (NBuffers and the shared memory segment
address space) is saved in process local heap memory since it's static
for the life of a server. It is passed to a new backend through
Postmaster. But with buffer pool being resized while the server running,
we need Postmaster to update its buffer pool metadata as the resizing
progresses and pass it to the new backend. This has few complications:
1. Postmaster does not receive ProcSignalBarrier. So we need to signal
it separately.
2. Postmaster's local state is inherited by the new backend when
fork()ed. But we need more complex implementation to pass it to an
exec()ed backend.
3. A new backend may receive the updated state from Postmaster and also
the signal barrier which prompts the same update. Thus the proc
signal barrier code needs to be idempotent; adding further complexity
to it.
4. This task takes away Postmaster resources from it's core
functionality.
This can be avoided by following two changes:
1. The shared memory is resized only in a single backend without
requiring any changes to the memory address space.
2. Maintaining the buffer pool metadata in the shared memory instead of
process local memory. This change may affect performance so verify
that performance is not degraded.
TODO: In case the backend executing pg_resize_shared_buffers() exits
before the operation finishes, we need to make sure that the changes
made to the shared memory while resizing are cleaned up properly.
Removing the evicted buffers from buffer ring
=============================================
If the buffer pool has been shrunk, the buffers in the buffer ring may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition and masks any bug that
introduces an invalid buffer into the ring. The alternative fix is more
complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure thus accessible when processing buffer
resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work.
Author: Ashutosh Bapat
Author: Dmitrii Dolgov
Author of some tests: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
Reviewed-by: Tomas Vondra
More detailed note follow: Need to see which of those fit in the commit
message and which should be removed.
Reinitializing strategry control area
=====================================
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
2. &StrategyControl->buffer_strategy_lock need not be initialized again.
3. nextVictimBuffer, completePasses and numBufferAllocs are viewed in
the context of NBuffers. Now that NBuffers itself has changed, those
three do not make sense. Reset them as if the server has restarted
again.
Ability to delay resizing operation
===================================
This commit introduces a flag delay_shmem_resize, which postgresql
backends and workers can use to signal the coordinator to delay resizing
operation. Background writer sets this flag when its scanning buffers.
Background writer operation (needs a rethink)
===========================
Background writer is blocked when the actual resizing is in progress. It
stops a scan in progress when it sees that the resizing has begun or is
about to begin. Once the buffer resizing is finished, before resuming
the regular operation, bgwriter resets the information saved so far.
This information is viewed in the context of NBuffers and hence does not
make sense after resizing which chanegs NBuffers.
Buffer lookup table
===================
Right now there is no way to free shared memory. Even if we shrink the
buffer lookup table when shrinking the buffer pool the unused hash table
entries can not be freed. When we expand the buffer pool, more entries
can be allocated but we can not resize the hash table directory without
rehashing all the entries. Just allocating more entries will lead to
more contention. Hence we setup the buffer lookup table considering the
maximum possible size of the buffer pool which is MaxAvailableMemory
only once at the beginning. Shared buffer lookup table and
StrategyControl are not resized even if the buffer pool is resized hence
they are allocated in the main shared memory segment
BgWriter refactoring
====================
The way BgBufferSync is written today, it packs four functionalities:
setting up the buffer sync state, performing the buffer sync, resetting
the buffer sync state when bgwriter_lru_maxpages <= 0 and setting it up
again after bgwriter_lru_maxpages > 0. That makes the code hard to read.
It will be good to divide this function into 3/4 different functions
each performing one functionality. Then pack all the state (the local
variables from that function converted to static global) into a
structure, which is passed to these functions. Once that happens
BgBufferSyncReset() will call one of the functions to reset the state
when buffer pool is resized.
---
contrib/pg_buffercache/pg_buffercache_pages.c | 18 +-
doc/src/sgml/config.sgml | 45 +-
doc/src/sgml/func/func-admin.sgml | 57 +++
src/backend/access/transam/slru.c | 2 +-
src/backend/access/transam/xlog.c | 2 +-
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/optimizer/path/costsize.c | 8 +-
src/backend/port/sysv_shmem.c | 165 +++++++
src/backend/port/win32_shmem.c | 6 +
src/backend/postmaster/checkpointer.c | 23 +-
src/backend/postmaster/postmaster.c | 10 +-
src/backend/storage/aio/aio_init.c | 2 +-
src/backend/storage/buffer/Makefile | 3 +-
src/backend/storage/buffer/README | 74 +++
src/backend/storage/buffer/buf_init.c | 274 +++++++++--
src/backend/storage/buffer/buf_resize.c | 455 ++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 10 +-
src/backend/storage/buffer/bufmgr.c | 178 ++++++-
src/backend/storage/buffer/freelist.c | 103 +++-
src/backend/storage/buffer/meson.build | 1 +
src/backend/storage/ipc/ipci.c | 13 +-
src/backend/storage/ipc/procsignal.c | 56 +++
src/backend/storage/ipc/shmem.c | 101 +++-
src/backend/tcop/postgres.c | 11 +
.../utils/activity/wait_event_names.txt | 2 +
src/backend/utils/init/globals.c | 5 +-
src/backend/utils/init/postinit.c | 49 ++
src/backend/utils/misc/guc.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 15 +-
src/include/catalog/pg_proc.dat | 6 +
src/include/miscadmin.h | 8 +
src/include/storage/buf.h | 2 +-
src/include/storage/buf_internals.h | 5 +-
src/include/storage/bufmgr.h | 20 +-
src/include/storage/ipc.h | 1 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 64 ++-
src/include/storage/procsignal.h | 9 +
src/include/storage/shmem.h | 3 +
src/include/utils/guc.h | 2 +
src/test/Makefile | 3 +-
src/test/README | 3 +
src/test/buffermgr/Makefile | 30 ++
src/test/buffermgr/README | 26 +
src/test/buffermgr/buffermgr_test.conf | 11 +
src/test/buffermgr/expected/buffer_resize.out | 330 +++++++++++++
src/test/buffermgr/meson.build | 23 +
src/test/buffermgr/sql/buffer_resize.sql | 97 ++++
src/test/buffermgr/t/001_resize_buffer.pl | 143 ++++++
.../buffermgr/t/003_parallel_resize_buffer.pl | 71 +++
.../t/004_client_join_buffer_resize.pl | 243 ++++++++++
src/test/meson.build | 1 +
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 76 +++
src/tools/pgindent/typedefs.list | 1 +
54 files changed, 2748 insertions(+), 123 deletions(-)
create mode 100644 src/backend/storage/buffer/buf_resize.c
create mode 100644 src/test/buffermgr/Makefile
create mode 100644 src/test/buffermgr/README
create mode 100644 src/test/buffermgr/buffermgr_test.conf
create mode 100644 src/test/buffermgr/expected/buffer_resize.out
create mode 100644 src/test/buffermgr/meson.build
create mode 100644 src/test/buffermgr/sql/buffer_resize.sql
create mode 100644 src/test/buffermgr/t/001_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/003_parallel_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/004_client_join_buffer_resize.pl
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index f60f797a9b4..8a17319ff2a 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -125,6 +125,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
TupleDesc tupledesc;
TupleDesc expected_tupledesc;
HeapTuple tuple;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
if (SRF_IS_FIRSTCALL())
{
@@ -181,10 +182,10 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
/* Allocate NBuffers worth of BufferCachePagesRec records. */
fctx->record = (BufferCachePagesRec *)
MemoryContextAllocHuge(CurrentMemoryContext,
- sizeof(BufferCachePagesRec) * NBuffers);
+ sizeof(BufferCachePagesRec) * currentNBuffers);
/* Set max calls and remember the user function context. */
- funcctx->max_calls = NBuffers;
+ funcctx->max_calls = currentNBuffers;
funcctx->user_fctx = fctx;
/* Return to original context when allocating transient memory */
@@ -198,13 +199,24 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
* snapshot across all buffers, but we do grab the buffer header
* locks, so the information of each buffer is self-consistent.
*/
- for (i = 0; i < NBuffers; i++)
+ for (i = 0; i < currentNBuffers; i++)
{
BufferDesc *bufHdr;
uint64 buf_state;
CHECK_FOR_INTERRUPTS();
+ /*
+ * TODO: We should just scan the entire buffer descriptor array
+ * instead of relying on curent buffer pool size. But that can
+ * happen if only we setup the descriptor array large enough at
+ * the server startup time.
+ */
+ if (currentNBuffers != pg_atomic_read_u32(&ShmemCtrl->currentNBuffers))
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("number of shared buffers changed during scan of buffer cache")));
+
bufHdr = GetBufferDescriptor(i);
/* Lock each buffer header before inspecting. */
buf_state = LockBufHdr(bufHdr);
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 5560b95ee60..37bb6048cc4 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -1732,7 +1732,6 @@ include_dir 'conf.d'
that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
(Non-default values of <symbol>BLCKSZ</symbol> change the minimum
value.)
- This parameter can only be set at server start.
</para>
<para>
@@ -1755,6 +1754,50 @@ include_dir 'conf.d'
appropriate, so as to leave adequate space for the operating system.
</para>
+ <para>
+ The shared memory consumed by the buffer pool is allocated and
+ initialized according to the value of the GUC at the time of starting
+ the server. A desired new value of GUC can be loaded while the server is
+ running using <systemitem>SIGHUP</systemitem>. But the buffer pool will
+ not be resized immediately. Use
+ <function>pg_resize_shared_buffers()</function> to dynamically resize
+ the shared buffer pool (see <xref linkend="functions-admin"/> for details).
+ <command>SHOW shared_buffers</command> shows the current number of
+ shared buffers and pending number, if any. Please note that when the GUC
+ is changed, the other GUCS which use this GUCs value to set their
+ defaults will not be changed. They may still require a server restart to
+ consider new value.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-max-shared-buffers" xreflabel="max_shared_buffers">
+ <term><varname>max_shared_buffers</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>max_shared_buffers</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Sets the upper limit for the <varname>shared_buffers</varname> value.
+ The default value is <literal>0</literal>,
+ which means no explicit limit is set and <varname>max_shared_buffers</varname>
+ will be automatically set to the value of <varname>shared_buffers</varname>
+ at server startup.
+ If this value is specified without units, it is taken as blocks,
+ that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
+ This parameter can only be set at server start.
+ </para>
+
+ <para>
+ This parameter determines the amount of memory address space to reserve
+ in each backend for expanding the buffer pool in future. While the
+ memory for buffer pool is allocated on demand as it is resized, the
+ memory required to hold the buffer manager metadata is allocated
+ statically at the server start accounting for the largest buffer pool
+ size allowed by this parameter.
+ <!-- TODO: Provide a numeric example of how much extra memory say max_shared_buffers = 1GB consume. -->
+ </para>
</listitem>
</varlistentry>
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index ea42056bbc9..a67b1880932 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -99,6 +99,63 @@
<returnvalue>off</returnvalue>
</para></entry>
</row>
+
+ <row>
+ <entry role="func_table_entry"><para role="func_signature">
+ <indexterm>
+ <primary>pg_resize_shared_buffers</primary>
+ </indexterm>
+ <function>pg_resize_shared_buffers</function> ()
+ <returnvalue>boolean</returnvalue>
+ </para>
+ <para>
+ Dynamically resizes the shared buffer pool to match the current
+ value of the <varname>shared_buffers</varname> parameter. This
+ function implements a coordinated resize process that ensures all
+ backend processes acknowledge the change before completing the
+ operation. The resize happens in multiple phases to maintain
+ data consistency and system stability. Returns <literal>true</literal>
+ if the resize was successful, or raises an error if the operation
+ fails. This function can only be called by superusers.
+ </para>
+ <para>
+ To resize shared buffers, first update the <varname>shared_buffers</varname>
+ setting and reload the configuration, then verify the new value is loaded
+ before calling this function. For example:
+<programlisting>
+postgres=# ALTER SYSTEM SET shared_buffers = '256MB';
+ALTER SYSTEM
+postgres=# SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers;
+ shared_buffers
+-------------------------
+ 128MB (pending: 256MB)
+(1 row)
+
+postgres=# SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+</programlisting>
+ The <command>SHOW shared_buffers</command> step is important to verify
+ that the configuration reload was successful and the new value is
+ available to the current session before attempting the resize. The
+ output shows both the current and pending values when a change is waiting
+ to be applied.
+ </para></entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/access/transam/slru.c b/src/backend/access/transam/slru.c
index 549c7e3e64b..2f30be34731 100644
--- a/src/backend/access/transam/slru.c
+++ b/src/backend/access/transam/slru.c
@@ -232,7 +232,7 @@ SimpleLruAutotuneBuffers(int divisor, int max)
{
return Min(max - (max % SLRU_BANK_SIZE),
Max(SLRU_BANK_SIZE,
- NBuffers / divisor - (NBuffers / divisor) % SLRU_BANK_SIZE));
+ NBuffersPending / divisor - (NBuffersPending / divisor) % SLRU_BANK_SIZE));
}
/*
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 16614e152dd..c6a64a129c1 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4694,7 +4694,7 @@ XLOGChooseNumBuffers(void)
{
int xbuffers;
- xbuffers = NBuffers / 32;
+ xbuffers = NBuffersPending / 32;
if (xbuffers > (wal_segment_size / XLOG_BLCKSZ))
xbuffers = (wal_segment_size / XLOG_BLCKSZ);
if (xbuffers < 8)
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index dd57624b4f9..139823ba7e7 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -337,6 +337,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ InitializeMaxNBuffers();
+
CreateSharedMemoryAndSemaphores();
/*
diff --git a/src/backend/optimizer/path/costsize.c b/src/backend/optimizer/path/costsize.c
index 16bf1f61a0f..9eb67d1827d 100644
--- a/src/backend/optimizer/path/costsize.c
+++ b/src/backend/optimizer/path/costsize.c
@@ -19,10 +19,10 @@
* is normally considerably less than random_page_cost. (However, if the
* database is fully cached in RAM, it is reasonable to set them equal.)
*
- * We also use a rough estimate "effective_cache_size" of the number of
- * disk pages in Postgres + OS-level disk cache. (We can't simply use
- * NBuffers for this purpose because that would ignore the effects of
- * the kernel's disk cache.)
+ * We also use a rough estimate "effective_cache_size" of the number of disk
+ * pages in Postgres + OS-level disk cache. (We can't simply use size of the
+ * buffer pool for this purpose because that would ignore the effects of the
+ * kernel's disk cache.)
*
* Obviously, taking constants for these values is an oversimplification,
* but it's tough enough to get any useful estimates even at this level of
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 88fdcf854a7..0399265c4dd 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,14 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
* TODO: The first two sentences in the first paragraph below make me feel like
@@ -101,6 +106,8 @@ typedef enum
SHMSTATE_UNATTACHED, /* pertinent to DataDir, no attached PIDs */
} IpcMemoryState;
+volatile bool delay_shmem_resize = false;
+
/*
* Anonymous mapping layout we use looks like this:
*
@@ -127,6 +134,9 @@ typedef enum
* being counted against memory limits). The mapping serves as an address space
* reservation, into which shared memory segment can be extended and is
* represented by the second /memfd:main with no permissions.
+ *
+ * The reserved space for buffer manager related segments is calculated based on
+ * MaxNBuffers.
*/
PGInhShmemSeg InhShmemSegs[NUM_MEMORY_MAPPINGS];
@@ -152,6 +162,42 @@ AnonShmemSegment AnonShmemSegs[NUM_MEMORY_MAPPINGS];
*/
static bool huge_pages_on = false;
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -946,6 +992,70 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the new shared_buffers value (saved
+ * in ShmemCtrl area). The actual segment resizing is done via ftruncate, which
+ * will fail if there is not sufficient space to expand the anon file.
+ *
+ * TODO: Rename this to BufferShmemResize() or something. Only buffer manager's
+ * memory should be resized in this function.
+ *
+ * TODO: This function changes the amount of shared memory used. So it should
+ * also update the show only GUCs shared_memory_size and
+ * shared_memory_size_in_huge_pages in all backends. SetConfigOption() may be
+ * used for that. But it's not clear whether is_reload parameter is safe to use
+ * while resizing is going on; also at what stage it should be done.
+ */
+static bool
+AnonymousShmemResize(int segment_id, MemoryMappingSizes *mapping, bool expanding)
+{
+ Size hugepagesize;
+ AnonShmemSegment *anonseg = &AnonShmemSegs[segment_id];
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffers, NBuffersPending);
+
+ if (anonseg->fd == -1)
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("segment[%s]: only anonymous (mmaped) file backed segments can be resized",
+ MappingName(segment_id))));
+
+#ifndef MAP_HUGETLB
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+ GetHugePageSize(&hugepagesize, NULL, NULL);
+ round_off_mapping_sizes_for_hugepages(mapping, hugepagesize);
+ }
+#endif
+ Assert(anonseg->addr);
+
+ /*
+ * Size of the reserved address space should not change, since it depends
+ * upon MaxNBuffers, which can be changed only on restart.
+ */
+ Assert(anonseg->size == mapping->shmem_reserved);
+
+ /*
+ * Resize the backing file to resize the allocated memory, and allocate
+ * more memory on supported platforms if required.
+ */
+ if (ftruncate(anonseg->fd, mapping->shmem_req_size) == -1)
+ ereport(ERROR,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncate anonymous file segment for \"%s\": %m",
+ MappingName(segment_id))));
+ if (expanding)
+ shmem_fallocate(anonseg->fd, MappingName(segment_id), mapping->shmem_req_size, ERROR);
+
+ return true;
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1130,6 +1240,42 @@ PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping,
return anonseg->addr;
}
+bool
+PGSharedMemoryResize(int segment_id, MemoryMappingSizes *mapping)
+{
+ AnonShmemSegment *anonseg = &AnonShmemSegs[segment_id];
+ PGShmemHeader *hdr;
+
+ /* For now, we allow only mmapped memory to be resized. */
+ if (shared_memory_type != SHMEM_TYPE_MMAP || anonseg->fd == -1)
+ elog(ERROR, "only anonymous (mmaped) file backed memory can be resized");
+
+ /* Anonymous memory has header as the first chunk. */
+ hdr = (PGShmemHeader *) anonseg->addr;
+
+ /* Main shared memory segment is always static. */
+ Assert(segment_id != MAIN_SHMEM_SEGMENT);
+
+ /*
+ * We should have reserved enough address space for resizing. PANIC if
+ * that's not the case.
+ */
+ if (hdr->ReservedSize < mapping->shmem_req_size)
+ ereport(PANIC,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved")));
+
+ /* Nothing to do if size is unchanged */
+ if (hdr->totalsize == mapping->shmem_req_size)
+ return true;
+
+ AnonymousShmemResize(segment_id, mapping, mapping->shmem_req_size > hdr->totalsize);
+
+ /* Update the available size. */
+ hdr->totalsize = mapping->shmem_req_size;
+ return true;
+}
+
#ifdef EXEC_BACKEND
/*
@@ -1271,3 +1417,22 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ pg_atomic_init_u32(&ShmemCtrl->targetNBuffers, 0);
+ pg_atomic_init_u32(&ShmemCtrl->currentNBuffers, 0);
+ pg_atomic_init_flag(&ShmemCtrl->resize_in_progress);
+
+ ShmemCtrl->coordinator = 0;
+ }
+}
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 5ee36063aff..15989ad622d 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -415,6 +415,12 @@ retry:
"on" : "off", PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
}
+bool
+PGSharedMemoryResize(int segment_id, MemoryMappingSizes *mapping)
+{
+ elog(ERROR, "shared memory resizing is not supported on Windows");
+}
+
/*
* PGSharedMemoryReAttach
*
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index 6482c21b8f9..96fc8a947b1 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -660,9 +660,12 @@ CheckpointerMain(const void *startup_data, size_t startup_data_len)
static void
ProcessCheckpointerInterrupts(void)
{
- if (ProcSignalBarrierPending)
- ProcessProcSignalBarrier();
-
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
if (ConfigReloadPending)
{
ConfigReloadPending = false;
@@ -682,6 +685,9 @@ ProcessCheckpointerInterrupts(void)
UpdateSharedMemoryConfig();
}
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+
/* Perform logging of memory contexts of this process */
if (LogMemoryContextPending)
ProcessLogMemoryContextInterrupt();
@@ -959,12 +965,13 @@ CheckpointerShmemSize(void)
Size size;
/*
- * The size of the requests[] array is arbitrarily set equal to NBuffers.
- * But there is a cap of MAX_CHECKPOINT_REQUESTS to prevent accumulating
- * too many checkpoint requests in the ring buffer.
+ * The size of the requests[] array is arbitrarily set equal to the
+ * initial size of buffer pool. But there is a cap of
+ * MAX_CHECKPOINT_REQUESTS to prevent accumulating too many checkpoint
+ * requests in the ring buffer.
*/
size = offsetof(CheckpointerShmemStruct, requests);
- size = add_size(size, mul_size(Min(NBuffers,
+ size = add_size(size, mul_size(Min(NBuffersPending,
MAX_CHECKPOINT_REQUESTS),
sizeof(CheckpointerRequest)));
@@ -995,7 +1002,7 @@ CheckpointerShmemInit(void)
*/
MemSet(CheckpointerShmem, 0, size);
SpinLockInit(&CheckpointerShmem->ckpt_lck);
- CheckpointerShmem->max_requests = Min(NBuffers, MAX_CHECKPOINT_REQUESTS);
+ CheckpointerShmem->max_requests = Min(NBuffersPending, MAX_CHECKPOINT_REQUESTS);
CheckpointerShmem->head = CheckpointerShmem->tail = 0;
ConditionVariableInit(&CheckpointerShmem->start_cv);
ConditionVariableInit(&CheckpointerShmem->done_cv);
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index d6133bfebc6..4c95376648d 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -110,11 +110,15 @@
#include "replication/slotsync.h"
#include "replication/walsender.h"
#include "storage/aio_subsys.h"
+#include "storage/bufmgr.h"
#include "storage/fd.h"
#include "storage/io_worker.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "tcop/backend_startup.h"
#include "tcop/tcopprot.h"
#include "utils/datetime.h"
@@ -125,7 +129,6 @@
#ifdef EXEC_BACKEND
#include "common/file_utils.h"
-#include "storage/pg_shmem.h"
#endif
@@ -958,6 +961,11 @@ PostmasterMain(int argc, char *argv[])
*/
InitializeFastPathLocks();
+ /*
+ * Calculate MaxNBuffers for buffer pool resizing.
+ */
+ InitializeMaxNBuffers();
+
/*
* Give preloaded libraries a chance to request additional shared memory.
*/
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index d3c68d8b04c..b8dac7eb9d5 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -101,7 +101,7 @@ AioChooseMaxConcurrency(void)
/* Similar logic to LimitAdditionalPins() */
max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
- max_proportional_pins = NBuffers / max_backends;
+ max_proportional_pins = NBuffersPending / max_backends;
max_proportional_pins = Max(max_proportional_pins, 1);
diff --git a/src/backend/storage/buffer/Makefile b/src/backend/storage/buffer/Makefile
index fd7c40dcb08..3bc9aee85de 100644
--- a/src/backend/storage/buffer/Makefile
+++ b/src/backend/storage/buffer/Makefile
@@ -17,6 +17,7 @@ OBJS = \
buf_table.o \
bufmgr.o \
freelist.o \
- localbuf.o
+ localbuf.o \
+ buf_resize.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..5dd753718a0 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -261,3 +261,77 @@ As of 8.4, background writer starts during recovery mode when there is
some form of potentially extended recovery to perform. It performs an
identical service to normal processing, except that checkpoints it
writes are technically restartpoints.
+
+Resizing shared buffers at runtime
+----------------------------------
+
+Before <TODO: Add version>, the size of the shared buffer pool (i.e. the number
+of shared buffers) was given by the global variable NBuffers and was fixed at
+server start time using GUC shared_buffers. In order to change the size of the
+shared buffer pool, one needed to change the GUC and restart the server.
+Starting <TODO: add version> PostgreSQL supports resizing the buffer pool
+without a server restart. The new GUC variable max_shared_buffers defines the
+maximum size of the shared buffer pool. Existing GUC shared_buffers controls the
+size of the buffer pool at run time. See configure.sgml for more details about
+these GUCs.
+
+The buffer manager maintains following data structures in shared memory.
+1. Buffer Descriptors: An array of BufferDesc structures, one per buffer.
+2. Buffer Blocks: An array of buffer blocks, forming the buffer pool.
+3. Buffer Lookup Table: A hash table mapping a page to the buffer containing
+ that page.
+4. IO conditional variables: An array of conditional variables, one per buffer.
+5. Checkpoint buffer ids: An array of buffer ids used during checkpointing.
+
+Except for the hash table, all the above data structures are required to be
+allocated as contiguous memory chunks. The code also relies on their start
+addresses being stable throughout the server lifetime. Resizing these structures
+means stretching or shrinking their tails. This is done by a. reserving address
+spaces for each of these structures during server start up, b. allocating or
+deallocating physical memory pages as needed during resizing, and c.
+initializing the array elements in the newly allocated memory regions.
+
+TODO: Should the memory management strategy be documented here or somewhere near
+storage/pg_shmem.c or sysv_shmem.c?
+
+Following protocol is used to make each of the segments elastic in a system
+which supports anonymous shared memory segments (like Linux). For each of the
+segments:
+1. Allocate a file descriptor using memfd_create() system call.
+2. Reserve address space for the segment using mmap(), passing it the fd
+ obtained in the first step with PROT_NONE and MAP_NORESERVE flags.
+3. Change the protection of the whole address space to PROT_READ | PROT_WRITE
+ using mprotect().
+4. Allocate or deallocate physical memory pages using ftruncate() as needed.
+
+The first three steps are executed when starting the server. The last step is
+executed when starting the server and during resizing.
+
+We should ideally call mmap() with PROT_READ | PROT_WRITE and avoid calling
+mprotect() separately. However, on Linux, when using huge pages, mmap() with
+PROT_READ | PROT_WRITE allocates memory for the whole region even if
+MAP_NORESERVE is specified. Above sequence avoids this problem.
+
+We could use mmap and mremap to resize the segments. However these calls will
+need to be executed by each backend process as well as the postmaster when
+resizing. This requires additional Synchronization between the postmaster and
+the resizing coordinator and requires careful implementation to avoid missing a
+newly joining backend during resizing. Using memfd_create and ftruncate avoids
+this complexity as ftruncate can be called from only one process, usually the
+coordinator and the changes are visible to all other processes automatically.
+
+Global variables (like NBuffers before this change) are inherited by each
+backend from the postmaster. If we continue to rely on NBuffers to provide the
+size of the buffer pool at a given point in time, we need to involve the
+postmater in resizing. To avoid this complexity, we store the current size of
+the buffer pool in the shared structure ShmemCtrl (name subject to change) and
+change only that variable during resizing. We use ProcSignalBarrier to know when
+all backends have observed the new size after resizing.
+
+See CreateAnonymousSegment() and AnonymousShmemResize() for detailed
+implementation.
+
+To resize shared buffers at runtime a user performs the steps mentioned in the
+description of shared_buffers GUC variable in configure.sgml. Actual resizing
+protocol is documented in the prologue of function pg_resize_shared_buffers() in
+src/backend/storage/buffer/buf_resize.c
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 864c8268cae..a11717b6709 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,8 +17,8 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
-#include "storage/pg_shmem.h"
#include "storage/proclist.h"
+#include "utils/guc.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -59,15 +59,33 @@ CkptSortItem *CkptBufferIds;
* multiple times. Check the PrivateRefCount infrastructure in bufmgr.c.
*/
+/*
+ * Initialize a single buffer.
+ */
+static void
+InitializeBuffer(int buf_id)
+{
+ BufferDesc *buf = GetBufferDescriptor(buf_id);
+
+ ClearBufferTag(&buf->tag);
+ pg_atomic_init_u64(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ buf->buf_id = buf_id;
+ pgaio_wref_clear(&buf->io_wref);
+ proclist_init(&buf->lock_waiters);
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+}
+
/*
* Initialize shared buffer pool
*
- * This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend). Size of data structures initialized
- * here depends on NBuffers, and to be able to change NBuffers without a
- * restart we store each structure into a separate shared memory segment, which
- * could be resized on demand.
+ * This is called once during shared-memory initialization.
+ * TODO: Restore this function to it's initial form. This function should see no
+ * change in buffer resize patches, except may be use of NBuffersPending.
+ *
+ * No locks are taking in this function, it is the caller responsibility to
+ * make sure only one backend can work with new buffers.
*/
void
BufferManagerShmemInit(void)
@@ -76,24 +94,25 @@ BufferManagerShmemInit(void)
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
ShmemInitStructInSegment("Buffer Descriptors",
- NBuffers * sizeof(BufferDescPadded),
+ NBuffersPending * sizeof(BufferDescPadded),
&foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
ShmemInitStructInSegment("Buffer Blocks",
- NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ NBuffersPending * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
&foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
ShmemInitStructInSegment("Buffer IO Condition Variables",
- NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ NBuffersPending * sizeof(ConditionVariableMinimallyPadded),
&foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
@@ -105,46 +124,41 @@ BufferManagerShmemInit(void)
*/
CkptBufferIds = (CkptSortItem *)
ShmemInitStructInSegment("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ NBuffersPending * sizeof(CkptSortItem), &foundBufCkpt,
CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
+
+ /*
+ * note: this path is only taken in EXEC_BACKEND case when
+ * initializing shared memory.
+ */
}
else
{
- int i;
-
/*
* Initialize all the buffer headers.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
-
- ClearBufferTag(&buf->tag);
-
- pg_atomic_init_u64(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
-
- buf->buf_id = i;
-
- pgaio_wref_clear(&buf->io_wref);
-
- proclist_init(&buf->lock_waiters);
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ for (i = 0; i < NBuffersPending; i++)
+ InitializeBuffer(i);
}
- /* Init other shared buffer-management stuff */
+ /*
+ * Init other shared buffer-management stuff.
+ */
StrategyInitialize(!foundDescs);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
&backend_flush_after);
+
+ /* Declare the size of current buffer pool. */
+ NBuffers = NBuffersPending;
+ pg_atomic_init_u32(&ShmemCtrl->currentNBuffers, NBuffersPending);
+ pg_atomic_init_u32(&ShmemCtrl->targetNBuffers, NBuffersPending);
}
/*
@@ -155,6 +169,8 @@ BufferManagerShmemInit(void)
* shared memory segment. The main segment must not allocate anything
* related to buffers, every other segment will receive part of the
* data.
+ *
+ * Also sets the shmem_reserved field for each segment based on MaxNBuffers.
*/
Size
BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
@@ -162,31 +178,211 @@ BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
size_t size;
/* size of buffer descriptors, plus alignment padding */
- size = add_size(0, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ size = add_size(0, mul_size(NBuffersPending, sizeof(BufferDescPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
mapping_sizes[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_req_size = size;
+ size = add_size(0, mul_size(MaxNBuffers, sizeof(BufferDescPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
mapping_sizes[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_reserved = size;
/* size of data pages, plus alignment padding */
size = add_size(0, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ size = add_size(size, mul_size(NBuffersPending, BLCKSZ));
mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
+ size = add_size(0, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(MaxNBuffers, BLCKSZ));
mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_reserved = size;
- /* size of stuff controlled by freelist.c */
- mapping_sizes[STRATEGY_SHMEM_SEGMENT].shmem_req_size = StrategyShmemSize();
- mapping_sizes[STRATEGY_SHMEM_SEGMENT].shmem_reserved = StrategyShmemSize();
-
/* size of I/O condition variables, plus alignment padding */
- size = add_size(0, mul_size(NBuffers,
+ size = add_size(0, mul_size(NBuffersPending,
sizeof(ConditionVariableMinimallyPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
mapping_sizes[BUFFER_IOCV_SHMEM_SEGMENT].shmem_req_size = size;
+ size = add_size(0, mul_size(MaxNBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
+ size = add_size(size, PG_CACHE_LINE_SIZE);
mapping_sizes[BUFFER_IOCV_SHMEM_SEGMENT].shmem_reserved = size;
/* size of checkpoint sort array in bufmgr.c */
- mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
- mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_reserved = mul_size(NBuffers, sizeof(CkptSortItem));
+ mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffersPending, sizeof(CkptSortItem));
+ mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_reserved = mul_size(MaxNBuffers, sizeof(CkptSortItem));
+
+ /* Allocations in the main memory segment, at the end. */
+
+ /* size of stuff controlled by freelist.c */
+ size = add_size(0, StrategyShmemSize());
return size;
}
+
+/*
+ * Reinitialize shared buffer manager structures when resizing the buffer pool.
+ *
+ * This function is called in the backend which coordinates buffer resizing
+ * operation.
+ *
+ * TODO: Avoid code duplication with BufferManagerShmemInit() and also assess
+ * which functionality in the latter is required in this function.
+ */
+void
+BufferManagerShmemResize(int currentNBuffers, int targetNBuffers)
+{
+ bool found;
+ int i;
+ void *tmpPtr;
+
+ tmpPtr = (BufferDescPadded *)
+ ShmemResizeStructInSegment("Buffer Descriptors",
+ targetNBuffers * sizeof(BufferDescPadded),
+ &found, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+ if (BufferDescriptors != tmpPtr || !found)
+ elog(FATAL, "resizing buffer descriptors failed: expected pointer %p, got %p, found=%d",
+ BufferDescriptors, tmpPtr, found);
+
+ tmpPtr = (ConditionVariableMinimallyPadded *)
+ ShmemResizeStructInSegment("Buffer IO Condition Variables",
+ targetNBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &found, BUFFER_IOCV_SHMEM_SEGMENT);
+ if (BufferIOCVArray != tmpPtr || !found)
+ elog(FATAL, "resizing buffer IO condition variables failed: expected pointer %p, got %p, found=%d",
+ BufferIOCVArray, tmpPtr, found);
+
+ tmpPtr = (CkptSortItem *)
+ ShmemResizeStructInSegment("Checkpoint BufferIds",
+ targetNBuffers * sizeof(CkptSortItem), &found,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+ if (CkptBufferIds != tmpPtr || !found)
+ elog(FATAL, "resizing checkpoint buffer IDs failed: expected pointer %p, got %p, found=%d",
+ CkptBufferIds, tmpPtr, found);
+
+ tmpPtr = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemResizeStructInSegment("Buffer Blocks",
+ targetNBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &found, BUFFERS_SHMEM_SEGMENT));
+ if (BufferBlocks != tmpPtr || !found)
+ elog(FATAL, "resizing buffer blocks failed: expected pointer %p, got %p, found=%d",
+ BufferBlocks, tmpPtr, found);
+
+ /*
+ * Initialize the headers for new buffers. If we are shrinking the
+ * buffers, currentNBuffers >= targetNBuffers, thus this loop doesn't
+ * execute.
+ */
+ for (i = currentNBuffers; i < targetNBuffers; i++)
+ InitializeBuffer(i);
+
+ /*
+ * We do not touch StrategyControl here. Instead it is done by background
+ * writer when handling PROCSIGNAL_BARRIER_SHBUF_EXPAND or
+ * PROCSIGNAL_BARRIER_SHBUF_SHRINK barrier.
+ */
+}
+
+/*
+ * BufferManagerShmemValidate
+ * Validate that buffer manager shared memory structures have correct
+ * pointers and sizes after a resize operation.
+ *
+ * This function is called by backends during ProcessBarrierShmemResizeStruct
+ * to ensure their view of the buffer structures is consistent after memory
+ * remapping.
+ */
+void
+BufferManagerShmemValidate(int targetNBuffers)
+{
+ bool found;
+ void *tmpPtr;
+
+ /* Validate Buffer Descriptors */
+ tmpPtr = (BufferDescPadded *)
+ ShmemInitStructInSegment("Buffer Descriptors",
+ targetNBuffers * sizeof(BufferDescPadded),
+ &found, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
+ if (!found || BufferDescriptors != tmpPtr)
+ elog(FATAL, "validating buffer descriptors failed: expected pointer %p, got %p, found=%d",
+ BufferDescriptors, tmpPtr, found);
+
+ /* Validate Buffer IO Condition Variables */
+ tmpPtr = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ targetNBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &found, BUFFER_IOCV_SHMEM_SEGMENT);
+ if (!found || BufferIOCVArray != tmpPtr)
+ elog(FATAL, "validating buffer IO condition variables failed: expected pointer %p, got %p, found=%d",
+ BufferIOCVArray, tmpPtr, found);
+
+ /* Validate Checkpoint BufferIds */
+ tmpPtr = (CkptSortItem *)
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ targetNBuffers * sizeof(CkptSortItem), &found,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
+ if (!found || CkptBufferIds != tmpPtr)
+ elog(FATAL, "validating checkpoint buffer IDs failed: expected pointer %p, got %p, found=%d",
+ CkptBufferIds, tmpPtr, found);
+
+ /* Validate Buffer Blocks */
+ tmpPtr = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSegment("Buffer Blocks",
+ targetNBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &found, BUFFERS_SHMEM_SEGMENT));
+ if (!found || BufferBlocks != tmpPtr)
+ elog(FATAL, "validating buffer blocks failed: expected pointer %p, got %p, found=%d",
+ BufferBlocks, tmpPtr, found);
+}
+
+/*
+ * check_shared_buffers
+ * GUC check_hook for shared_buffers
+ *
+ * When reloading the configuration, shared_buffers should not be set to a value
+ * higher than max_shared_buffers fixed at the boot time.
+ */
+bool
+check_shared_buffers(int *newval, void **extra, GucSource source)
+{
+ if (finalMaxNBuffers && *newval > MaxNBuffers)
+ {
+ GUC_check_errdetail("\"shared_buffers\" must be less than \"max_shared_buffers\".");
+ return false;
+ }
+ return true;
+}
+
+/*
+ * show_shared_buffers
+ * GUC show_hook for shared_buffers
+ *
+ * Shows both current and pending buffer counts with proper unit formatting.
+ */
+const char *
+show_shared_buffers(void)
+{
+ static char buffer[128];
+ int64 current_value,
+ pending_value;
+ const char *current_unit,
+ *pending_unit;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+
+ if (currentNBuffers == NBuffersPending)
+ {
+ /* No buffer pool resizing pending. */
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s", current_value, current_unit);
+ }
+ else
+ {
+ /*
+ * Shared buffer pool is pending to be resized, show both current and
+ * pending sizes.
+ */
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ convert_int_from_base_unit(NBuffersPending, GUC_UNIT_BLOCKS, &pending_value, &pending_unit);
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s (pending: " INT64_FORMAT "%s)",
+ current_value, current_unit, pending_value, pending_unit);
+ }
+
+ return buffer;
+}
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
new file mode 100644
index 00000000000..7b6d46c05ba
--- /dev/null
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -0,0 +1,455 @@
+/*-------------------------------------------------------------------------
+ *
+ * buf_resize.c
+ * shared buffer pool resizing functionality
+ *
+ * This module contains the implementation of shared buffer pool resizing,
+ * including the main resize coordination function and barrier processing
+ * functions that synchronize all backends during resize operations.
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/storage/buffer/buf_resize.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "miscadmin.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
+#include "utils/injection_point.h"
+
+
+/*
+ * Prepare ShmemCtrl for resizing the shared buffer pool.
+ */
+static void
+MarkBufferResizingStart(int targetNBuffers, int currentNBuffers)
+{
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ Assert(pg_atomic_read_u32(&ShmemCtrl->currentNBuffers) == currentNBuffers);
+
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, targetNBuffers);
+ ShmemCtrl->coordinator = MyProcPid;
+}
+
+/*
+ * Reset ShmemCtrl after resizing the shared buffer pool is done.
+ */
+static void
+MarkBufferResizingEnd(int newNBuffers)
+{
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ Assert(pg_atomic_read_u32(&ShmemCtrl->currentNBuffers) == newNBuffers);
+
+ /*
+ * TODO: should we leave targetNBuffers as is? We are setting it to
+ * NBuffers in BufferManagerShmemInit().
+ */
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, 0);
+ ShmemCtrl->coordinator = -1;
+}
+
+/*
+ * Communicate given buffer pool resize barrier to all other backends and the Postmaster.
+ *
+ * ProcSignalBarrier is not sent to the Postmaster but we need the Postmaster to
+ * update its knowledge about the buffer pool so that it can be inherited by the
+ * child processes.
+ */
+static void
+SharedBufferResizeBarrier(ProcSignalBarrierType barrier, const char *barrier_name)
+{
+ WaitForProcSignalBarrier(EmitProcSignalBarrier(barrier));
+ elog(LOG, "all backends acknowledged %s barrier", barrier_name);
+
+#ifdef USE_INJECTION_POINTS
+ /* Injection point specific to this barrier type */
+ switch (barrier)
+ {
+ case PROCSIGNAL_BARRIER_SHBUF_SHRINK:
+ INJECTION_POINT("pgrsb-shrink-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM:
+ INJECTION_POINT("pgrsb-resize-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_EXPAND:
+ INJECTION_POINT("pgrsb-expand-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED:
+ /* TODO: Add an injection point here. */
+ break;
+ case PROCSIGNAL_BARRIER_SMGRRELEASE:
+ case PROCSIGNAL_BARRIER_UPDATE_XLOG_LOGICAL_INFO:
+
+ /*
+ * Not relevant in this function but it's here so that the
+ * compiler can detect any missing shared buffer resizing barrier
+ * enum here.
+ */
+ break;
+ }
+#endif /* USE_INJECTION_POINTS */
+}
+
+/*
+ * C implementation of SQL interface to update the shared buffers according to
+ * the current values of shared_buffers GUCs.
+ *
+ * The current boundaries of the buffer pool are given by two ranges.
+ *
+ * - [1, StrategyControl::activeNBuffers] is the range of buffers from which new
+ * allocations can happen at any time.
+ *
+ * - [1, ShmemCtrl::currentNBuffers] is the range of valid buffers at any given
+ * time.
+ *
+ * Let's assume that before resizing, the number of buffers in the buffer pool is
+ * NBuffersOld. After resizing it is NBuffersNew. Before resizing
+ * StrategyControl::activeNBuffers == ShmemCtrl::currentNBuffers == NBuffersOld.
+ * After the resizing finishes StrategyControl::activeNBuffers ==
+ * ShmemCtrl::currentNBuffers == NBuffersNew. Thus when no resizing happens these
+ * two ranges are same.
+ *
+ * Following steps are performed by the coordinator during resizing.
+ *
+ * 1. Marks resizing in progress to avoid multiple concurrent invocations of this
+ * function.
+ *
+ * 2. When shrinking the shared buffer pool, the coordinator sends SHBUF_SHRINK
+ * ProcSignalBarrier. In response to this barrier background writer is expected
+ * to set StrategyControl::activeNBuffers = NBuffersNew to restrict the new
+ * buffer allocations only to the new buffer pool size and also reset its
+ * internal state. Once every backend has acknowledged the barrier, the
+ * coordinator can be sure that new allocations will not happen in the buffer
+ * pool area being shrunk. Then it evicts the buffers in that area. Note that
+ * ShmemCtrl::currentNBuffers is still NBuffersOld, since backend may still
+ * access buffers allocated before the resizing started. Buffer eviction may fail
+ * if a buffer being evicted is pinned and the resizing operatino is aborted.
+ * Once the eviction is finished, the extra memory can be freed in the next step.
+ *
+ * 2. This step is executed in both cases, when expanding the buffer pool or
+ * shrinking the buffer pool. The anonymous file backing each of the shared
+ * memory segment containg the buffer pool shared data structures is resized to
+ * the amount of memory required for the new buffer pool size. When expanding the
+ * expanded portion of memory is initialized appropriately.
+ * ShmemCtrl::currentNBuffers is set to NBuffersNew to indicate new range of
+ * valid shared buffers. Every backend is sent SHBUF_RESIZE_MAP_AND_MEM barrier.
+ * All the backends validate that their pointers to the shared buffers structure
+ * are valid and have the right size. Once every backend has acknowledged the
+ * barrier, this step finishes.
+ *
+ * 3. When expanding the buffer pool, the coordinator sends SHBUF_EXPAND barrier
+ * to signal end of expansion. When expadning the background writer, in response
+ * to StrategyControl::activeNBuffers = NBufferNew so that new allocations can
+ * use expanded range of buffer pool.
+ *
+ * TODO: Handle the case when the backend executing this function dies or the
+ * query is cancelled or it hits an error while resizing.
+ */
+Datum
+pg_resize_shared_buffers(PG_FUNCTION_ARGS)
+{
+ bool result = true;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+ int targetNBuffers = NBuffersPending;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+
+ if (currentNBuffers == targetNBuffers)
+ {
+ elog(LOG, "shared buffers are already at %d, no need to resize", currentNBuffers);
+ PG_RETURN_BOOL(true);
+ }
+
+ if (!pg_atomic_test_set_flag(&ShmemCtrl->resize_in_progress))
+ {
+ elog(LOG, "shared buffer resizing already in progress");
+ PG_RETURN_BOOL(false);
+ }
+
+ /*
+ * TODO: NBuffersPending may change after it was sampled above, thus
+ * leading to wrong memory size estimates. Find a way to pass
+ * targetNBuffers value to BufferManagerShmemSize().
+ */
+ BufferManagerShmemSize(mapping_sizes);
+ /* Round it off to a multiple of a typical page size */
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ /* Structures in main memory segment are never resized. */
+ if (i == MAIN_SHMEM_SEGMENT)
+ continue;
+
+ round_off_mapping_sizes(&mapping_sizes[i]);
+ }
+
+ /*
+ * TODO: What if the NBuffersPending value seen here is not the desired
+ * one because somebody did a pg_reload_conf() between the last
+ * pg_reload_conf() and execution of this function?
+ */
+ MarkBufferResizingStart(targetNBuffers, currentNBuffers);
+ elog(LOG, "resizing shared buffers from %d to %d", currentNBuffers, targetNBuffers);
+
+ INJECTION_POINT("pg-resize-shared-buffers-flag-set", NULL);
+
+ /* Phase 1: SHBUF_SHRINK - Only for shrinking buffer pool */
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * Phase 1: Shrinking - send SHBUF_SHRINK barrier Every backend sets
+ * activeNBuffers = NewNBuffers to restrict buffer pool allocations to
+ * the new size
+ */
+ elog(LOG, "Phase 1: Shrinking buffer pool, restricting allocations to %d buffers", targetNBuffers);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_SHRINK, CppAsString(PROCSIGNAL_BARRIER_SHBUF_SHRINK));
+
+ /* Evict buffers in the area being shrunk */
+ elog(LOG, "evicting buffers %u..%u", targetNBuffers + 1, currentNBuffers);
+ if (!EvictExtraBuffers(targetNBuffers, currentNBuffers))
+ {
+ elog(WARNING, "failed to evict extra buffers during shrinking");
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED, CppAsString(PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED));
+ MarkBufferResizingEnd(currentNBuffers);
+ pg_atomic_clear_flag(&ShmemCtrl->resize_in_progress);
+ PG_RETURN_BOOL(false);
+ }
+
+ /*
+ * Shrink buffer manager structures before shrinking the shared
+ * memory.
+ */
+ BufferManagerShmemResize(currentNBuffers, targetNBuffers);
+
+ /* Update the current NBuffers. */
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, targetNBuffers);
+ }
+
+ /* Phase 2: SHBUF_RESIZE_MAP_AND_MEM - Both expanding and shrinking */
+ elog(LOG, "Phase 2: Remapping shared memory segments and updating structures");
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ /* Structures in the main memory segment are never resized. */
+ if (i == MAIN_SHMEM_SEGMENT)
+ continue;
+
+ if (!PGSharedMemoryResize(i, &mapping_sizes[i]))
+ {
+ /*
+ * This should never fail since address map should already be
+ * reserved. So the failure should be treated as PANIC.
+ */
+ elog(PANIC, "failed to resize anonymous shared memory");
+ }
+ }
+
+ INJECTION_POINT("pgrsb-after-shmem-resize", NULL);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM, CppAsString(PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM));
+
+ /* Phase 3: SHBUF_EXPAND - Only for expanding buffer pool */
+ if (targetNBuffers > currentNBuffers)
+ {
+ /* Expand buffer manager structures after expanding the shared memory. */
+ BufferManagerShmemResize(currentNBuffers, targetNBuffers);
+
+ /*
+ * Phase 3: Expanding - send SHBUF_EXPAND barrier Backends set
+ * activeNBuffers = NewNBuffers and start allocating buffers from the
+ * expanded range
+ */
+ elog(LOG, "Phase 3: Expanding buffer pool, enabling allocations up to %d buffers", targetNBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, targetNBuffers);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_EXPAND, CppAsString(PROCSIGNAL_BARRIER_SHBUF_EXPAND));
+ }
+
+ /*
+ * Reset buffer resize control area.
+ */
+ MarkBufferResizingEnd(targetNBuffers);
+
+ pg_atomic_clear_flag(&ShmemCtrl->resize_in_progress);
+
+ elog(LOG, "successfully resized shared buffers to %d", targetNBuffers);
+
+ PG_RETURN_BOOL(result);
+}
+
+bool
+ProcessBarrierShmemShrink(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * Delay adjusting the new active size of buffer pool till this process
+ * becomes ready to resize buffers.
+ */
+ if (delay_shmem_resize)
+ {
+ elog(LOG, "Phase 1: Delaying SHBUF_SHRINK barrier - restricting allocations to %d buffers, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+
+ return false;
+ }
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation
+ * statistics and the strategy control together so that background
+ * writer doesn't go out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody
+ * would reset the strategy control area. So we can't rely on
+ * background worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, targetNBuffers);
+ /* Reset strategy control to new size */
+ StrategyReset(targetNBuffers);
+ }
+
+ elog(LOG, "Phase 1: Processing SHBUF_SHRINK barrier - target buffer pool size = %d, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemResizeMapAndMem(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * If buffer pool is being shrunk, we are already working with a smaller
+ * buffer pool, so shrinking address space and shared structures should
+ * not be a problem. When expanding, expanding the address space and
+ * shared structures beyond the current boundaries is not going to be a
+ * problem since we are not accessing that memory yet. So there is no
+ * reason to delay processing this barrier.
+ */
+
+ /*
+ * Coordinator has already adjusted its address map and also updated sizes
+ * of the shared buffer structures, no further validation needed.
+ */
+ if (ShmemCtrl->coordinator == MyProcPid)
+ return true;
+
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * When shrinking, shared data structures have been resized at this
+ * point. Validate that their pointers to shared buffer structures
+ * are still valid and have the correct size after resizing.
+ *
+ * TODO: Do want to do this only in assert enabled builds?
+ */
+ BufferManagerShmemValidate(targetNBuffers);
+ elog(LOG, "Backend %d successfully validated structure pointers after resize", MyProcPid);
+ }
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemExpand(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * Delay adjusting the new active size of buffer pool till this process
+ * becomes ready to resize buffers.
+ */
+ if (delay_shmem_resize)
+ {
+ elog(LOG, "Phase 3: delaying SHBUF_EXPAND barrier - enabling allocations up to %d buffers, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+ return false;
+ }
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation
+ * statistics and the strategy control together so that background
+ * writer doesn't go out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody
+ * would reset the strategy control area. So we can't rely on
+ * background worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, targetNBuffers);
+ StrategyReset(targetNBuffers);
+ }
+
+ /*
+ * Shared data structures must have been resized by now. Validate that
+ * their pointers to shared buffer structures are still valid and have the
+ * correct size after resizing.
+ *
+ * TODO: Do want to do this only in assert enabled builds?
+ */
+ BufferManagerShmemValidate(targetNBuffers);
+ elog(LOG, "Backend %d successfully validated structure pointers after resize", MyProcPid);
+
+ elog(LOG, "Phase 3: Processing SHBUF_EXPAND barrier - targetNBuffers = %d, ShmemCtrl->coordinator = %d", targetNBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemResizeFailed(void)
+{
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation
+ * statistics and the strategy control together so that background
+ * writer doesn't go out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody
+ * would reset the strategy control area. So we can't rely on
+ * background worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, currentNBuffers);
+ /* Reset strategy control to new size */
+ StrategyReset(currentNBuffers);
+ }
+
+ elog(LOG, "received proc signal indicating failure to resize shared buffers from %d to %d, restoring to %d, coordinator is %d",
+ currentNBuffers, targetNBuffers, currentNBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
+
+/*
+ * TODO: add progress report facility if required.
+ */
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a33786a460b..0dfdffb25c8 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -41,7 +41,8 @@ static HTAB *SharedBufHash;
/*
* Estimate space needed for mapping hashtable
- * size is the desired hash table size (possibly more than NBuffers)
+ * size is the desired hash table size (possibly more than the size of buffer
+ * pool)
*/
Size
BufTableShmemSize(int size)
@@ -65,6 +66,13 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
+ /*
+ * The shared buffer look up table is set up only once with maximum
+ * possible entries considering maximum size of the buffer pool. It is not
+ * resized after that even if the buffer pool is resized. Hence it is
+ * allocated in the main shared memory segment and not in a resizeable
+ * shared memory segment.
+ */
SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
size, size,
&info,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 6f935648ae9..fb8a282391e 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/read_stream.h"
@@ -223,11 +224,11 @@ static BufferDesc *PinCountWaitBuf = NULL;
* and, if so, in what mode.
*
*
- * To avoid - as we used to - requiring an array with NBuffers entries to keep
- * track of local buffers, we use a small sequentially searched array
- * (PrivateRefCountArrayKeys, with the corresponding data stored in
- * PrivateRefCountArray) and an overflow hash table (PrivateRefCountHash) to
- * keep track of backend local pins.
+ * To avoid - as we used to - requiring an array, with as many entries as the
+ * size of buffer pool, to keep track of local buffers, we use a small
+ * sequentially searched array (PrivateRefCountArrayKeys, with the corresponding
+ * data stored in PrivateRefCountArray) and an overflow hash table
+ * (PrivateRefCountHash) to keep track of backend local pins.
*
* Until no more than REFCOUNT_ARRAY_ENTRIES buffers are pinned at once, all
* refcounts are kept track of in the array; after that, new array entries
@@ -3523,7 +3524,7 @@ BufferSync(int flags)
set_bits, 0,
0);
- /* Check for barrier events in case NBuffers is large. */
+ /* Check for barrier events in case the buffer pool is large. */
if (ProcSignalBarrierPending)
ProcessProcSignalBarrier();
}
@@ -3720,6 +3721,32 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_DONE(NBuffers, num_written, num_to_scan);
}
+/*
+ * Information saved between BgBufferSync() calls so we can determine the
+ * strategy point's advance rate and avoid scanning already-cleaned buffers. The
+ * variables are global instead of static local so that BgBufferSyncReset() can
+ * adjust it when resizing shared buffers.
+ */
+static bool saved_info_valid = false;
+static int prev_strategy_buf_id;
+static uint32 prev_strategy_passes;
+static int next_to_clean;
+static uint32 next_passes;
+
+/* Moving averages of allocation rate and clean-buffer density */
+static float smoothed_alloc = 0;
+static float smoothed_density = 10.0;
+
+void
+BgBufferSyncReset(int currentNBuffers, int targetNBuffers)
+{
+ saved_info_valid = false;
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
+ currentNBuffers, targetNBuffers);
+#endif
+}
+
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
@@ -3739,20 +3766,6 @@ BgBufferSync(WritebackContext *wb_context)
uint32 strategy_passes;
uint32 recent_alloc;
- /*
- * Information saved between calls so we can determine the strategy
- * point's advance rate and avoid scanning already-cleaned buffers.
- */
- static bool saved_info_valid = false;
- static int prev_strategy_buf_id;
- static uint32 prev_strategy_passes;
- static int next_to_clean;
- static uint32 next_passes;
-
- /* Moving averages of allocation rate and clean-buffer density */
- static float smoothed_alloc = 0;
- static float smoothed_density = 10.0;
-
/* Potentially these could be tunables, but for now, not */
float smoothing_samples = 16;
float scan_whole_pool_milliseconds = 120000.0;
@@ -3775,6 +3788,25 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ /*
+ * If buffer pool is being shrunk the buffer being written out may not
+ * remain valid. If the buffer pool is being expanded, more buffers will
+ * become available without even this function writing out any. Hence wait
+ * till buffer resizing finishes i.e. go into hibernation mode.
+ *
+ * TODO: We may not need this synchronization if background worker itself
+ * becomes the coordinator.
+ */
+ if (!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress))
+ return true;
+
+ /*
+ * Resizing shared buffers while this function is performing an LRU scan
+ * on them may lead to wrong results. Indicate that the resizing should
+ * wait for the LRU scan to complete.
+ */
+ delay_shmem_resize = true;
+
/*
* Find out where the clock-sweep currently is, and how many buffer
* allocations have happened since our last call.
@@ -3792,6 +3824,7 @@ BgBufferSync(WritebackContext *wb_context)
if (bgwriter_lru_maxpages <= 0)
{
saved_info_valid = false;
+ delay_shmem_resize = false;
return true;
}
@@ -3951,8 +3984,17 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
- while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
+ /*
+ * Execute the LRU scan.
+ *
+ * If buffer pool is being shrunk, the buffer being written may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing any. Hence stop what we
+ * are doing. This also unblocks other processes that are waiting for
+ * buffer resizing to finish.
+ */
+ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
+ !pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress))
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -4011,6 +4053,9 @@ BgBufferSync(WritebackContext *wb_context)
#endif
}
+ /* Let the resizing commence. */
+ delay_shmem_resize = false;
+
/* Return true if OK to hibernate */
return (bufs_to_lap == 0 && recent_alloc == 0);
}
@@ -4341,7 +4386,29 @@ DebugPrintBufferRefcount(Buffer buffer)
void
CheckPointBuffers(int flags)
{
+ /*
+ * Mark that buffer sync is in progress - delay any shared memory
+ * resizing.
+ */
+ /*
+ * TODO: We need to assess whether we should allow checkpoint and buffer
+ * resizing to run in parallel. When expanding buffers it may be fine to
+ * let the checkpointer run in RESIZE_MAP_AND_MEM phase but delay phase
+ * EXPAND phase till the checkpoint finishes, at the same time not allow
+ * checkpoint to run during expansion phase. When shrinking the buffers,
+ * we should delay SHRINK phase till checkpoint finishes and not allow to
+ * start checkpoint till SHRINK phase is done, but allow it to run in
+ * RESIZE_MAP_AND_MEM phase. This needs careful analysis and testing.
+ */
+ delay_shmem_resize = true;
+
BufferSync(flags);
+
+ /*
+ * Mark that buffer sync is no longer in progress - allow shared memory
+ * resizing
+ */
+ delay_shmem_resize = false;
}
/*
@@ -8528,3 +8595,70 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
+{
+ bool result = true;
+
+ Assert(targetNBuffers < currentNBuffers);
+
+ /*
+ * If the buffer being evicated is locked, this function will need to
+ * wait. This function should not be called from a Postmaster since it can
+ * not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = targetNBuffers + 1; buf <= currentNBuffers; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint64 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u64(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is going one
+ * hence unlocked precheck should be safe and saves some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ /*
+ * XXX: Looks like CurrentResourceOwner can be NULL here, find another
+ * one in that case?
+ */
+ if (CurrentResourceOwner)
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 13e701ee4a1..6221e035024 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -33,10 +33,16 @@ typedef struct
/* Spinlock: protects the values below */
slock_t buffer_strategy_lock;
+ /*
+ * Number of active buffers that can be allocated. During buffer resizing,
+ * this may be different from the actual size of the buffer pool.
+ */
+ pg_atomic_uint32 activeNBuffers;
+
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo NBuffers.
+ * get an actual buffer, it needs to be used modulo activeNBuffers.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -101,21 +107,27 @@ static inline uint32
ClockSweepTick(void)
{
uint32 victim;
+ int activeBuffers;
/*
* Atomically move hand ahead one buffer - if there's several processes
* doing this, this can lead to buffers being returned slightly out of
- * apparent order.
+ * apparent order. We need to read both the current position of hand and
+ * the current buffer allocation limit together consistently. They may be
+ * reset by concurrent resize.
*/
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
victim =
pg_atomic_fetch_add_u32(&StrategyControl->nextVictimBuffer, 1);
+ activeBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
- if (victim >= NBuffers)
+ if (victim >= activeBuffers)
{
uint32 originalVictim = victim;
/* always wrap what we look up in BufferDescriptors */
- victim = victim % NBuffers;
+ victim = victim % activeBuffers;
/*
* If we're the one that just caused a wraparound, force
@@ -143,7 +155,7 @@ ClockSweepTick(void)
*/
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
- wrapped = expected % NBuffers;
+ wrapped = expected % activeBuffers;
success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
&expected, wrapped);
@@ -228,7 +240,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
pg_atomic_fetch_add_u32(&StrategyControl->numBufferAllocs, 1);
/* Use the "clock sweep" algorithm to find a free buffer */
- trycounter = NBuffers;
+ trycounter = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+
for (;;)
{
uint64 old_buf_state;
@@ -281,7 +294,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
local_buf_state))
{
- trycounter = NBuffers;
+ trycounter = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
break;
}
}
@@ -323,10 +336,12 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
{
uint32 nextVictimBuffer;
int result;
+ uint32 activeNBuffers;
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
- result = nextVictimBuffer % NBuffers;
+ activeNBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+ result = nextVictimBuffer % activeNBuffers;
if (complete_passes)
{
@@ -336,7 +351,7 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
* Additionally add the number of wraparounds that happened before
* completePasses could be incremented. C.f. ClockSweepTick().
*/
- *complete_passes += nextVictimBuffer / NBuffers;
+ *complete_passes += nextVictimBuffer / activeNBuffers;
}
if (num_buf_alloc)
@@ -383,7 +398,7 @@ StrategyShmemSize(void)
Size size = 0;
/* size of lookup hash table ... see comment in StrategyInitialize */
- size = add_size(size, BufTableShmemSize(NBuffers + NUM_BUFFER_PARTITIONS));
+ size = add_size(size, BufTableShmemSize(MaxNBuffers + NUM_BUFFER_PARTITIONS));
/* size of the shared replacement strategy control block */
size = add_size(size, MAXALIGN(sizeof(BufferStrategyControl)));
@@ -391,6 +406,32 @@ StrategyShmemSize(void)
return size;
}
+void
+StrategyReset(int activeNBuffers)
+{
+ Assert(StrategyControl);
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ /* Update the active buffer count for the strategy */
+ pg_atomic_write_u32(&StrategyControl->activeNBuffers, activeNBuffers);
+
+ /* Reset the clock-sweep pointer to start from beginning */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The statistics is viewed in the context of the number of shared
+ * buffers. Reset it as the size of active number of shared buffers
+ * changes.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_write_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* TODO: Do we need to seset background writer notifications? */
+ StrategyControl->bgwprocno = -1;
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+}
+
/*
* StrategyInitialize -- initialize the buffer cache replacement
* strategy.
@@ -408,12 +449,21 @@ StrategyInitialize(bool init)
*
* Since we can't tolerate running out of lookup table entries, we must be
* sure to specify an adequate table size here. The maximum steady-state
- * usage is of course NBuffers entries, but BufferAlloc() tries to insert
- * a new entry before deleting the old. In principle this could be
- * happening in each partition concurrently, so we could need as many as
- * NBuffers + NUM_BUFFER_PARTITIONS entries.
+ * usage is of course is as many number of entries as the number of
+ * buffers in the buffer pool. Right now there is no way to free shared
+ * memory. Even if we shrink the buffer lookup table when shrinking the
+ * buffer pool the unused hash table entries can not be freed. When we
+ * expand the buffer pool, more entries can be allocated but we can not
+ * resize the hash table directory without rehashing all the entries. Just
+ * allocating more entries will lead to more contention. Hence we setup
+ * the buffer lookup table considering the maximum possible size of the
+ * buffer pool which is MaxNBuffers.
+ *
+ * Additionally BufferAlloc() tries to insert a new entry before deleting
+ * the old. In principle this could be happening in each partition
+ * concurrently, so we need extra NUM_BUFFER_PARTITIONS entries.
*/
- InitBufTable(NBuffers + NUM_BUFFER_PARTITIONS);
+ InitBufTable(MaxNBuffers + NUM_BUFFER_PARTITIONS);
/*
* Get or create the shared strategy control block
@@ -421,7 +471,7 @@ StrategyInitialize(bool init)
StrategyControl = (BufferStrategyControl *)
ShmemInitStructInSegment("Buffer Strategy Status",
sizeof(BufferStrategyControl),
- &found, STRATEGY_SHMEM_SEGMENT);
+ &found, MAIN_SHMEM_SEGMENT);
if (!found)
{
@@ -432,6 +482,8 @@ StrategyInitialize(bool init)
SpinLockInit(&StrategyControl->buffer_strategy_lock);
+ /* Initialize the active buffer count */
+ pg_atomic_init_u32(&StrategyControl->activeNBuffers, NBuffersPending);
/* Initialize the clock-sweep pointer */
pg_atomic_init_u32(&StrategyControl->nextVictimBuffer, 0);
@@ -669,12 +721,25 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
+ * If the slot hasn't been filled yet or the buffer in the slot has been
+ * invalidated when buffer pool was shrunk, tell the caller to allocate a
+ * new buffer with the normal allocation strategy. He will then fill this
* slot by calling AddBufferToRing with the new buffer.
+ *
+ * TODO: buffer ids in the ring will never be greater than the size of
+ * buffer pool, except maybe the first time ring is accessed after
+ * shrinking the buffer pool. Checking the upper bound on buffer id always
+ * may mask a bug bugs that introduces buffer ids higher than the size of
+ * buffer pool in the ring. But performing that check only once after
+ * shrinking seems impossible. The BufferAccessStrategy objects are not
+ * accessible outside the ScanState. Hence we can not purge the buffers
+ * while evicting the buffers. After the resizing is finished, it's not
+ * possible to notice when we touch the first of those objects and the
+ * last of objects. See if this can fixed.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer ||
+ bufnum > pg_atomic_read_u32(&StrategyControl->activeNBuffers))
return NULL;
buf = GetBufferDescriptor(bufnum - 1);
diff --git a/src/backend/storage/buffer/meson.build b/src/backend/storage/buffer/meson.build
index ed84bf08971..f219e29d5ef 100644
--- a/src/backend/storage/buffer/meson.build
+++ b/src/backend/storage/buffer/meson.build
@@ -6,4 +6,5 @@ backend_sources += files(
'bufmgr.c',
'freelist.c',
'localbuf.c',
+ 'buf_resize.c',
)
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 12457e1bbf3..1e92b0bcc5e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -156,6 +156,14 @@ CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
size = add_size(size, WaitLSNShmemSize());
size = add_size(size, LogicalDecodingCtlShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -170,8 +178,7 @@ CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
/* might as well round it off to a multiple of a typical page size */
for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
{
- mapping_sizes[segment].shmem_req_size = add_size(mapping_sizes[segment].shmem_req_size, 8192 - (mapping_sizes[segment].shmem_req_size % 8192));
- mapping_sizes[segment].shmem_reserved = add_size(mapping_sizes[segment].shmem_reserved, 8192 - (mapping_sizes[segment].shmem_reserved % 8192));
+ round_off_mapping_sizes(&mapping_sizes[segment]);
/* Compute the total size of all segments */
size = size + mapping_sizes[segment].shmem_req_size;
}
@@ -327,6 +334,8 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
+ /* TODO: This should be part of BufferManagerShmemInit() */
+ ShmemControlInit();
BufferManagerShmemInit();
/*
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 8e56922dcea..df9bb7a0f1a 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -24,9 +24,11 @@
#include "port/pg_bitutils.h"
#include "replication/logicalworker.h"
#include "replication/walsender.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -109,6 +111,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -170,6 +176,44 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -579,6 +623,18 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_UPDATE_XLOG_LOGICAL_INFO:
processed = ProcessBarrierUpdateXLogLogicalInfo();
break;
+ case PROCSIGNAL_BARRIER_SHBUF_SHRINK:
+ processed = ProcessBarrierShmemShrink();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM:
+ processed = ProcessBarrierShmemResizeMapAndMem();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_EXPAND:
+ processed = ProcessBarrierShmemExpand();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED:
+ processed = ProcessBarrierShmemResizeFailed();
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 2e365261c09..5a8ea67b311 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -82,11 +82,19 @@
#include "funcapi.h"
#include "miscadmin.h"
#include "port/pg_numa.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
#include "storage/shmem.h"
#include "storage/spin.h"
#include "utils/builtins.h"
+#include "utils/injection_point.h"
+#include "utils/wait_event.h"
/* Structure managing one shared memory segment. */
typedef struct ShmemSegment
@@ -541,8 +549,7 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, int segmen
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. The size better be the same as the size we are trying to
*/
if (result->size != size)
{
@@ -552,6 +559,7 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, int segmen
" \"%s\": expected %zu, actual %zu",
name, size, result->size)));
}
+
structPtr = result->location;
}
else
@@ -586,6 +594,95 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, int segmen
return structPtr;
}
+/*
+ * ShmemResizeStructInSegment -- Resize the given structure in shared memory.
+ *
+ * This function resizes an existing shared memory structure while preserving
+ * the existing memory location.
+ *
+ * Returns: pointer to the existing structure location, if the resize is
+ * successful, otherwise NULL.
+ */
+void *
+ShmemResizeStructInSegment(const char *name, Size size, bool *foundPtr,
+ int segment_id)
+{
+ ShmemIndexEnt *result;
+ void *structPtr;
+ ShmemSegment *segment;
+ PGShmemHeader *shmhdr;
+ Size allocated_size;
+ Size newFree;
+
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+ Assert(segment_id != MAIN_SHMEM_SEGMENT); /* main segment structures not
+ * resizable */
+ Assert(ShmemIndex);
+ Assert(size > 0);
+ segment = &Segments[segment_id];
+ shmhdr = segment->ShmemSegHdr;
+ Assert(shmhdr != NULL);
+
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+ /* Look up the structure in the shmem index */
+ result = (ShmemIndexEnt *)
+ hash_search(ShmemIndex, name, HASH_FIND, foundPtr);
+
+ Assert(*foundPtr);
+ Assert(result);
+ Assert(result->segment_id == segment_id);
+
+ /* Save the existing structure pointer to be returned. */
+ structPtr = result->location;
+
+ /* Cachealign new size */
+ allocated_size = CACHELINEALIGN(size);
+
+ if (allocated_size == result->allocated_size)
+ {
+ result->size = size;
+ /* No need to resize if the existing allocated size is sufficient */
+ LWLockRelease(ShmemIndexLock);
+ return structPtr;
+ }
+
+ SpinLockAcquire(segment->ShmemLock);
+
+ /*
+ * The resizable structures are placed in their own segment after the
+ * header and the spinlock. Hence the memory location where they end are
+ * same as the start of free memory in that segment.
+ */
+ Assert((char *) segment->ShmemBase + shmhdr->freeoffset == (char *) result->location + result->allocated_size);
+ newFree = shmhdr->freeoffset + (allocated_size - result->allocated_size);
+ if (newFree > shmhdr->totalsize)
+ {
+ structPtr = NULL;
+ }
+ else
+ {
+ shmhdr->freeoffset = newFree;
+ result->size = size;
+ result->allocated_size = allocated_size;
+ }
+
+ /*
+ * End of the structure should still be same as the start of free memory
+ * in the segment
+ */
+ Assert((char *) segment->ShmemBase + shmhdr->freeoffset == (char *) result->location + result->allocated_size);
+
+ SpinLockRelease(segment->ShmemLock);
+ LWLockRelease(ShmemIndexLock);
+
+ /* note this assert is okay with structPtr == NULL */
+ Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
+
+ return structPtr;
+}
+
+
+
/*
* Add two Size values, checking for overflow
*/
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index e54bf1e760f..79884f060ed 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -63,6 +63,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4128,6 +4129,9 @@ PostgresSingleUserMain(int argc, char *argv[],
/* Initialize size of fast-path lock cache. */
InitializeFastPathLocks();
+ /* Initialize MaxNBuffers for buffer pool resizing. */
+ InitializeMaxNBuffers();
+
/*
* Give preloaded libraries a chance to request additional shared memory.
*/
@@ -4318,6 +4322,13 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /*
+ * TODO: The new backend should fetch the shared buffers status. If the
+ * resizing is going on, it should bring itself upto speed with it. If
+ * not, simply fetch the latest pointers are sizes. Is this the right
+ * place to do that?
+ */
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 5537a2d2530..e1c2873a139 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -163,6 +163,7 @@ WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
WAL_RECEIVER_WAIT_START "Waiting for startup process to send initial data for streaming replication."
WAL_SUMMARY_READY "Waiting for a new WAL summary to be generated."
XACT_GROUP_UPDATE "Waiting for the group leader to update transaction status at transaction end."
+PM_BUFFER_RESIZE_WAIT "Waiting for the postmaster to complete shared buffer pool resize operations."
ABI_compatibility:
@@ -364,6 +365,7 @@ SerialControl "Waiting to read or update shared <filename>pg_serial</filename> s
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index 36ad708b360..be9d3909c0f 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -139,7 +139,10 @@ int max_parallel_maintenance_workers = 2;
* MaxBackends is computed by PostmasterMain after modules have had a chance to
* register background workers.
*/
-int NBuffers = 16384;
+int NBuffers = 0;
+int NBuffersPending = 16384;
+bool finalMaxNBuffers = false;
+int MaxNBuffers = 0;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 3f401faf3de..3869ed5201c 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -595,6 +595,55 @@ InitializeFastPathLocks(void)
pg_nextpower2_32(FastPathLockGroupsPerBackend));
}
+/*
+ * Initialize MaxNBuffers variable with validation.
+ *
+ * This must be called after GUCs have been loaded but before shared memory size
+ * is determined.
+ *
+ * Since MaxNBuffers limits the size of the buffer pool, it must be at least as
+ * much as NBuffersPending. If MaxNBuffers is 0 (default), set it to
+ * NBuffersPending. Otherwise, validate that MaxNBuffers is not less than
+ * NBuffersPending.
+ */
+void
+InitializeMaxNBuffers(void)
+{
+ if (MaxNBuffers == 0) /* default/boot value */
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", NBuffersPending);
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+
+ /*
+ * We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
+ * However, if the DBA explicitly set max_shared_buffers = 0 in the
+ * config file, then PGC_S_DYNAMIC_DEFAULT will fail to override that
+ * and we must force the matter with PGC_S_OVERRIDE.
+ */
+ if (MaxNBuffers == 0) /* failed to apply it? */
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+ else
+ {
+ if (MaxNBuffers < NBuffersPending)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("max_shared_buffers (%d) cannot be less than current shared_buffers (%d)",
+ MaxNBuffers, NBuffersPending),
+ errhint("Increase max_shared_buffers or decrease shared_buffers.")));
+ }
+ }
+
+ Assert(MaxNBuffers > 0);
+ Assert(!finalMaxNBuffers);
+ finalMaxNBuffers = true;
+}
+
/*
* Early initialization of a backend (either standalone or under postmaster).
* This happens even before InitPostgres.
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index ae9d5f3fb70..6bc35a1132b 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -2599,7 +2599,7 @@ convert_to_base_unit(double value, const char *unit,
* the value without loss. For example, if the base unit is GUC_UNIT_KB, 1024
* is converted to 1 MB, but 1025 is represented as 1025 kB.
*/
-static void
+void
convert_int_from_base_unit(int64 base_value, int base_unit,
int64 *value, const char **unit)
{
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index f0260e6e412..f017b4f22d6 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2023,6 +2023,15 @@
max => 'MAX_BACKENDS /* XXX? */',
},
+{ name => "max_shared_buffers", type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+ short_desc => 'Sets the upper limit for the shared_buffers value.',
+ flags => 'GUC_UNIT_BLOCKS',
+ variable => 'MaxNBuffers',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX / 2',
+},
+
{ name => 'max_slot_wal_keep_size', type => 'int', context => 'PGC_SIGHUP', group => 'REPLICATION_SENDING',
short_desc => 'Sets the maximum WAL size that can be reserved by replication slots.',
long_desc => 'Replication slots will be marked as failed, and segments released for deletion or recycling, if this much space is occupied by WAL on disk. -1 means no maximum.',
@@ -2591,13 +2600,15 @@
# We sometimes multiply the number of shared buffers by two without
# checking for overflow, so we mustn't allow more than INT_MAX / 2.
-{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
- variable => 'NBuffers',
+ variable => 'NBuffersPending',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
+ check_hook => 'check_shared_buffers',
+ show_hook => 'show_shared_buffers',
},
{ name => 'shared_memory_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 52de299c2d8..84c636407bf 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12703,4 +12703,10 @@
proname => 'hashoid8extended', prorettype => 'int8',
proargtypes => 'oid8 int8', prosrc => 'hashoid8extended' },
+{ oid => '9999', descr => 'resize shared buffers according to the value of GUC `shared_buffers`',
+ proname => 'pg_resize_shared_buffers',
+ provolatile => 'v',
+ prorettype => 'bool',
+ proargtypes => '',
+ prosrc => 'pg_resize_shared_buffers'},
]
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index db559b39c4d..df6d3d8f4dd 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,7 +173,14 @@ extern PGDLLIMPORT bool ExitOnAnyError;
extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
+/*
+ * TODO: This is no more a GUC variable and does not track the size of the shared
+ * buffer pool; should be removed.
+ */
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersPending;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
@@ -502,6 +509,7 @@ extern PGDLLIMPORT ProcessingMode Mode;
extern void pg_split_opts(char **argv, int *argcp, const char *optstr);
extern void InitializeMaxBackends(void);
extern void InitializeFastPathLocks(void);
+extern void InitializeMaxNBuffers(void);
extern void InitPostgres(const char *in_dbname, Oid dboid,
const char *username, Oid useroid,
bits32 flags,
diff --git a/src/include/storage/buf.h b/src/include/storage/buf.h
index b21445522b1..a12d6b9082a 100644
--- a/src/include/storage/buf.h
+++ b/src/include/storage/buf.h
@@ -17,7 +17,7 @@
/*
* Buffer identifiers.
*
- * Zero is invalid, positive is the index of a shared buffer (1..NBuffers),
+ * Zero is invalid, positive is the index of a shared buffer (1..{size of shared buffer pool}),
* negative is the index of a local buffer (-1 .. -NLocBuffer).
*/
typedef int Buffer;
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index f4e9e703b8b..91e28ce19d2 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -137,8 +137,8 @@ StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2),
/*
* The maximum allowed value of usage_count represents a tradeoff between
- * accuracy and speed of the clock-sweep buffer management algorithm. A
- * large value (comparable to NBuffers) would approximate LRU semantics.
+ * accuracy and speed of the clock-sweep buffer management algorithm. A large
+ * value (comparable to the size of buffer pool) would approximate LRU semantics.
* But it can take as many as BM_MAX_USAGE_COUNT+1 complete cycles of the
* clock-sweep hand to find a free buffer, so in practice we don't want the
* value to be very large.
@@ -574,6 +574,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyReset(int activeNBuffers);
/* buf_table.c */
extern Size BufTableShmemSize(int size);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 93348a34378..649a1a35105 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -21,6 +21,7 @@
#include "storage/bufpage.h"
#include "storage/pg_shmem.h"
#include "storage/relfilelocator.h"
+#include "utils/guc.h"
#include "utils/relcache.h"
#include "utils/snapmgr.h"
@@ -159,6 +160,7 @@ typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersPending;
/* in bufmgr.c */
extern PGDLLIMPORT bool zero_damaged_pages;
@@ -221,6 +223,11 @@ typedef enum BufferLockMode
BUFFER_LOCK_EXCLUSIVE,
} BufferLockMode;
+/*
+ * prototypes for functions in buf_init.c
+ */
+extern const char *show_shared_buffers(void);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
/*
* prototypes for functions in bufmgr.c
@@ -341,6 +348,7 @@ extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
extern bool BgBufferSync(WritebackContext *wb_context);
+extern void BgBufferSyncReset(int currentNBuffers, int targetNBuffers);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
@@ -365,10 +373,13 @@ extern void MarkDirtyRelUnpinnedBuffers(Relation rel,
extern void MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
int32 *buffers_already_dirty,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int targetNBuffers, int currentNBuffers);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
extern Size BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes);
+extern void BufferManagerShmemResize(int currentNBuffers, int targetNBuffers);
+extern void BufferManagerShmemValidate(int targetNBuffers);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
@@ -417,7 +428,7 @@ extern void FreeAccessStrategy(BufferAccessStrategy strategy);
static inline bool
BufferIsValid(Buffer bufnum)
{
- Assert(bufnum <= NBuffers);
+ Assert(bufnum <= (Buffer) pg_atomic_read_u32(&ShmemCtrl->currentNBuffers));
Assert(bufnum >= -NLocBuffer);
return bufnum != InvalidBuffer;
@@ -471,4 +482,11 @@ BufferGetPage(Buffer buffer)
#endif /* FRONTEND */
+/* buf_resize.c */
+extern Datum pg_resize_shared_buffers(PG_FUNCTION_ARGS);
+extern bool ProcessBarrierShmemShrink(void);
+extern bool ProcessBarrierShmemResizeMapAndMem(void);
+extern bool ProcessBarrierShmemExpand(void);
+extern bool ProcessBarrierShmemResizeFailed(void);
+
#endif /* BUFMGR_H */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index f1d0802d048..23003412c9a 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -66,6 +66,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index e94ebce95b9..cb454f4c81d 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -87,6 +87,7 @@ PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
+PG_LWLOCK(56, ShmemResize)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 8bd78a89d38..ac679259787 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,8 +24,13 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "port/atomics.h"
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
+#include "storage/procsignal.h"
#include "storage/spin.h"
+#include "storage/shmem.h"
+#include "utils/guc.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
@@ -97,7 +102,7 @@ typedef struct PGInhShmemSeg
#define STRATEGY_SHMEM_SEGMENT 5
/* Number of available segments for anonymous memory mappings */
-#define NUM_MEMORY_MAPPINGS 6
+#define NUM_MEMORY_MAPPINGS 5
/*
* Structure to hold required sizes of each shared memory segment as calculated
@@ -114,11 +119,39 @@ typedef struct MemoryMappingSizes
extern PGDLLIMPORT PGInhShmemSeg InhShmemSegs[NUM_MEMORY_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ *
+ * TODO: I think we need a lock to protect this structure. If we do so, do we
+ * need to use atomic integers?
+ *
+ * TODO: Merge this structure into StrategyControl?
+ */
+typedef struct
+{
+ pg_atomic_flag resize_in_progress; /* true if resizing is in progress.
+ * false otherwise. */
+ pg_atomic_uint32 currentNBuffers; /* Original NBuffers value before
+ * resize started */
+ pg_atomic_uint32 targetNBuffers;
+ pid_t coordinator;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -150,13 +183,15 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping_sizes,
- PGShmemHeader **shim);
-extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
-extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
- int *memfd_flags);
-extern void PrepareHugePages(void);
+/*
+ * round off mapping size to a multiple of a typical page size.
+ */
+static inline void
+round_off_mapping_sizes(MemoryMappingSizes *mapping_sizes)
+{
+ mapping_sizes->shmem_req_size = add_size(mapping_sizes->shmem_req_size, 8192 - (mapping_sizes->shmem_req_size % 8192));
+ mapping_sizes->shmem_reserved = add_size(mapping_sizes->shmem_reserved, 8192 - (mapping_sizes->shmem_reserved % 8192));
+}
static inline const char *
MappingName(int segment_id)
@@ -180,5 +215,18 @@ MappingName(int segment_id)
}
}
+extern PGShmemHeader *PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping_sizes,
+ PGShmemHeader **shim);
+extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
+extern void PGSharedMemoryDetach(void);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+extern bool PGSharedMemoryResize(int segment_id, MemoryMappingSizes *mapping_sizes);
+
+extern void PrepareHugePages(void);
+extern const char *show_shared_buffers(void);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
+extern void ShmemControlInit(void);
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index e52b8eb7697..9dc9819a72b 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -56,6 +56,15 @@ typedef enum
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
PROCSIGNAL_BARRIER_UPDATE_XLOG_LOGICAL_INFO, /* ask to update
* XLogLogicalInfo */
+ PROCSIGNAL_BARRIER_SHBUF_SHRINK, /* shrink buffer pool - restrict
+ * allocations to new size */
+ PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM, /* remap shared memory
+ * segments and update
+ * structure pointers */
+ PROCSIGNAL_BARRIER_SHBUF_EXPAND, /* expand buffer pool - enable
+ * allocations in new range */
+ PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED, /* signal backends that the shared
+ * buffer resizing failed. */
} ProcSignalBarrierType;
/*
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index e227119f59d..da55034e5cc 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -40,11 +40,14 @@ extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
extern void *ShmemInitStructInSegment(const char *name, Size size,
bool *foundPtr, int segment_id);
+extern void *ShmemResizeStructInSegment(const char *name, Size size,
+ bool *foundPtr, int segment_id);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
extern PGDLLIMPORT Size pg_get_shmem_pagesize(void);
+
/* ipci.c */
extern void RequestAddinShmemSpace(Size size);
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index bf39878c43e..72d8ba9e59e 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -459,6 +459,8 @@ extern config_handle *get_config_handle(const char *name);
extern void AlterSystemSetConfigFile(AlterSystemStmt *altersysstmt);
extern char *GetConfigOptionByName(const char *name, const char **varname,
bool missing_ok);
+extern void convert_int_from_base_unit(int64 base_value, int base_unit,
+ int64 *value, const char **unit);
extern void TransformGUCArray(ArrayType *array, List **names,
List **values);
diff --git a/src/test/Makefile b/src/test/Makefile
index 3eb0a06abb4..7a0d74086c1 100644
--- a/src/test/Makefile
+++ b/src/test/Makefile
@@ -20,7 +20,8 @@ SUBDIRS = \
postmaster \
recovery \
regress \
- subscription
+ subscription \
+ buffermgr
ifeq ($(with_icu),yes)
SUBDIRS += icu
diff --git a/src/test/README b/src/test/README
index afdc7676519..77f11607ff7 100644
--- a/src/test/README
+++ b/src/test/README
@@ -15,6 +15,9 @@ examples/
Demonstration programs for libpq that double as regression tests via
"make check"
+buffermgr/
+ Tests for resizing buffer pool without restarting the server
+
isolation/
Tests for concurrent behavior at the SQL level
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
new file mode 100644
index 00000000000..eb275027fa6
--- /dev/null
+++ b/src/test/buffermgr/Makefile
@@ -0,0 +1,30 @@
+#-------------------------------------------------------------------------
+#
+# Makefile for src/test/buffermgr
+#
+# Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+# Portions Copyright (c) 1994, Regents of the University of California
+#
+# src/test/buffermgr/Makefile
+#
+#-------------------------------------------------------------------------
+
+EXTRA_INSTALL = contrib/pg_buffercache
+
+REGRESS = buffer_resize
+
+# Custom configuration for buffer manager tests
+TEMP_CONFIG = $(srcdir)/buffermgr_test.conf
+
+subdir = src/test/buffermgr
+top_builddir = ../../..
+include $(top_builddir)/src/Makefile.global
+
+check:
+ $(prove_check)
+
+installcheck:
+ $(prove_installcheck)
+
+clean distclean:
+ rm -rf tmp_check
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
new file mode 100644
index 00000000000..c375ad80989
--- /dev/null
+++ b/src/test/buffermgr/README
@@ -0,0 +1,26 @@
+src/test/buffermgr/README
+
+Regression tests for buffer manager
+===================================
+
+This directory contains a test suite for resizing buffer manager without restarting the server.
+
+
+Running the tests
+=================
+
+NOTE: You must have given the --enable-tap-tests argument to configure.
+
+Run
+ make check
+or
+ make installcheck
+You can use "make installcheck" if you previously did "make install".
+In that case, the code in the installation tree is tested. With
+"make check", a temporary installation tree is built from the current
+sources and then tested.
+
+Either way, this test initializes, starts, and stops a test Postgres
+cluster.
+
+See src/test/perl/README for more info about running these tests.
diff --git a/src/test/buffermgr/buffermgr_test.conf b/src/test/buffermgr/buffermgr_test.conf
new file mode 100644
index 00000000000..a15f3e442a5
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.conf
@@ -0,0 +1,11 @@
+# Configuration for buffer manager regression tests
+
+# Even if max_shared_buffers is set multiple times only the last one is used to
+# as the limit on shared_buffers.
+max_shared_buffers = 128kB
+# Set initial shared_buffers as expected by test
+shared_buffers = 128MB
+# Set a larger value for max_shared_buffers to allow testing resize operations
+max_shared_buffers = 300MB
+# Turn huge pages off, since that affects the size of memory segments
+huge_pages = off
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
new file mode 100644
index 00000000000..739c2560da8
--- /dev/null
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -0,0 +1,330 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+-- TODO: The test sets shared_buffers values in MBs. Instead it could use values
+-- in kBs so that the test runs on very small machines.
+-- TODO: test the actual memory allocated in the shared memory segments.
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | descriptors | 1048576 | 1048576
+ Buffer IO Condition Variables | iocv | 262144 | 262144
+ Checkpoint BufferIds | checkpoint | 327680 | 327680
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+-------------+-----------+---------------
+ buffers | 134225920 | 314580992
+ checkpoint | 335872 | 770048
+ descriptors | 1056768 | 2465792
+ iocv | 270336 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | descriptors | 1048576 | 1048576
+ Buffer IO Condition Variables | iocv | 262144 | 262144
+ Checkpoint BufferIds | checkpoint | 327680 | 327680
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+-------------+-----------+---------------
+ buffers | 134225920 | 314580992
+ checkpoint | 335872 | 770048
+ descriptors | 1056768 | 2465792
+ iocv | 270336 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 128MB (pending: 64MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 64MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+----------+----------------
+ Buffer Blocks | buffers | 67112960 | 67112960
+ Buffer Descriptors | descriptors | 524288 | 524288
+ Buffer IO Condition Variables | iocv | 131072 | 131072
+ Checkpoint BufferIds | checkpoint | 163840 | 163840
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+-------------+----------+---------------
+ buffers | 67117056 | 314580992
+ checkpoint | 172032 | 770048
+ descriptors | 532480 | 2465792
+ iocv | 139264 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 8192
+(1 row)
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 64MB (pending: 256MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 268439552 | 268439552
+ Buffer Descriptors | descriptors | 2097152 | 2097152
+ Buffer IO Condition Variables | iocv | 524288 | 524288
+ Checkpoint BufferIds | checkpoint | 655360 | 655360
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+-------------+-----------+---------------
+ buffers | 268443648 | 314580992
+ checkpoint | 663552 | 770048
+ descriptors | 2105344 | 2465792
+ iocv | 532480 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 32768
+(1 row)
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 256MB (pending: 100MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 100MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+-----------+----------------
+ Buffer Blocks | buffers | 104861696 | 104861696
+ Buffer Descriptors | descriptors | 819200 | 819200
+ Buffer IO Condition Variables | iocv | 204800 | 204800
+ Checkpoint BufferIds | checkpoint | 256000 | 256000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+-------------+-----------+---------------
+ buffers | 104865792 | 314580992
+ checkpoint | 262144 | 770048
+ descriptors | 827392 | 2465792
+ iocv | 212992 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 12800
+(1 row)
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 100MB (pending: 128kB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+-------------+--------+----------------
+ Buffer Blocks | buffers | 135168 | 135168
+ Buffer Descriptors | descriptors | 1024 | 1024
+ Buffer IO Condition Variables | iocv | 256 | 256
+ Checkpoint BufferIds | checkpoint | 320 | 384
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+-------------+--------+---------------
+ buffers | 139264 | 314580992
+ checkpoint | 8192 | 770048
+ descriptors | 8192 | 2465792
+ iocv | 8192 | 622592
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16
+(1 row)
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+ERROR: invalid value for parameter "shared_buffers": 51200
+DETAIL: "shared_buffers" must be less than "max_shared_buffers".
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
new file mode 100644
index 00000000000..c24bff721e6
--- /dev/null
+++ b/src/test/buffermgr/meson.build
@@ -0,0 +1,23 @@
+# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+
+tests += {
+ 'name': 'buffermgr',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'buffer_resize',
+ ],
+ 'regress_args': ['--temp-config', files('buffermgr_test.conf')],
+ },
+ 'tap': {
+ 'env': {
+ 'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
+ },
+ 'tests': [
+ 't/001_resize_buffer.pl',
+ 't/003_parallel_resize_buffer.pl',
+ 't/004_client_join_buffer_resize.pl',
+ ],
+ },
+}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
new file mode 100644
index 00000000000..fc27522f097
--- /dev/null
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -0,0 +1,97 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+-- TODO: The test sets shared_buffers values in MBs. Instead it could use values
+-- in kBs so that the test runs on very small machines.
+
+-- TODO: test the actual memory allocated in the shared memory segments.
+
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+SHOW max_shared_buffers;
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
new file mode 100644
index 00000000000..30e8bfea9cf
--- /dev/null
+++ b/src/test/buffermgr/t/001_resize_buffer.pl
@@ -0,0 +1,143 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Minimal test testing shared_buffer resizing under load
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Function to resize buffer pool and verify the change.
+sub apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ # Use the new pg_resize_shared_buffers() interface which handles everything synchronously
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ # If resize function fails, try a few times before giving up
+ my $max_retries = 5;
+ my $retry_delay = 1; # seconds
+ my $success = 0;
+ for my $attempt (1..$max_retries) {
+ my $result = $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()");
+ if ($result eq 't') {
+ $success = 1;
+ last;
+ }
+
+ # If not the last attempt, wait before retrying
+ if ($attempt < $max_retries) {
+ note "Resizing buffer pool to $new_size, attempt $attempt failed, retrying after $retry_delay seconds...";
+ sleep($retry_delay);
+ }
+ }
+
+ is($success, 1, 'resizing to ' . $new_size . ' succeeded after retries');
+ is($node->safe_psql('postgres', "SHOW shared_buffers"), $new_size,
+ 'SHOW after resizing to '. $new_size . ' succeeded');
+}
+
+# Initialize a cluster and start pgbench in the background for concurrent load.
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Permit resizing up to 1GB for this test and let the server start with 128MB.
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 160
+shared_buffers = 16
+log_statement = none
+});
+
+$node->start;
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+my $pgb_scale = 1;
+my $pgb_duration = 120;
+my $pgb_num_clients = 3;
+$node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$pgb_scale --quiet",
+ 0,
+ [qr{^$}],
+ [ # stderr patterns to verify initialization stages
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$pgb_scale)"
+);
+my ($pgbench_stdin, $pgbench_stdout, $pgbench_stderr) = ('', '', '');
+# Use --exit-on-abort so that the test stops on the first server crash or error,
+# thus making it easy to debug the failure. Use -C to increase the chances of a
+# new backend being created while resizing the buffer pool.
+my $pgbench_process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-T', $pgb_duration,
+ '-c', $pgb_num_clients,
+ '-C',
+ '--exit-on-abort',
+ 'postgres'
+ ],
+ '<' => \$pgbench_stdin,
+ '>' => \$pgbench_stdout,
+ '2>' => \$pgbench_stderr
+);
+
+ok($pgbench_process, "pgbench started successfully");
+
+# Allow pgbench to establish connections and start generating load.
+#
+# TODO: When creating new backends is known to work well with buffer pool
+# resizing, this wait should be removed.
+sleep(1);
+
+# Resize buffer pool to various sizes while pgbench is running in the
+# background. We use smaller sizes to induce frequent buffer eviction and
+# allocation. Also smaller buffer pool means frequent wraparound in background
+# writer, default buffer allocation strategy and checkpointer.
+#
+# TODO: These are pseudo-randomly picked sizes, but we can do better.
+my $tests_completed = 0;
+my @buffer_sizes = (32, 24, 29, 40, 29, 20, 16, 24);
+for my $target_size (@buffer_sizes)
+{
+ # Convert number of buffers to a string that will be reported by SHOW
+ # shared_buffers. This simple calculation works for sizes smaller than 128
+ # beyond which the unit changes to MB.
+ $target_size = $target_size * 8;
+ $target_size = $target_size . 'kB';
+
+ # Verify workload generator is still running
+ if (!$pgbench_process->pumpable) {
+ ok(0, "pgbench is still running");
+ last;
+ }
+
+ apply_and_verify_buffer_change($node, $target_size);
+ $tests_completed++;
+
+ # Wait for the resized buffer pool to stabilize. If the resized buffer pool
+ # is utilized fully, it might hit any wrongly initialized areas of shared
+ # memory.
+ sleep(2);
+}
+is($tests_completed, scalar(@buffer_sizes), "All buffer sizes were tested");
+
+# Make sure that pgbench can end normally.
+$pgbench_process->signal('TERM');
+IPC::Run::finish $pgbench_process;
+ok(grep { $pgbench_process->result == $_ } (0, 15), "pgbench exited gracefully");
+
+# Log any error output from pgbench for debugging
+diag("pgbench stderr:\n$pgbench_stderr");
+diag("pgbench stdout:\n$pgbench_stdout");
+
+# Ensure database is still functional after all the buffer changes
+$node->connect_ok("dbname=postgres",
+ "Database remains accessible after $tests_completed buffer resize operations");
+
+done_testing();
diff --git a/src/test/buffermgr/t/003_parallel_resize_buffer.pl b/src/test/buffermgr/t/003_parallel_resize_buffer.pl
new file mode 100644
index 00000000000..40d4bfde437
--- /dev/null
+++ b/src/test/buffermgr/t/003_parallel_resize_buffer.pl
@@ -0,0 +1,71 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test that only one pg_resize_shared_buffers() call succeeds when multiple
+# sessions attempt to resize buffers concurrently
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Initialize a cluster
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', 'shared_buffers = 128kB');
+$node->append_conf('postgresql.conf', 'max_shared_buffers = 256kB');
+$node->start;
+
+# Load injection points extension for test coordination
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Test 1: Two concurrent pg_resize_shared_buffers() calls
+# Set up injection point to pause the first resize call
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('pg-resize-shared-buffers-flag-set', 'wait')");
+
+# Change shared_buffers for the resize operation
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '144kB'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+# Start first resize session (will pause at injection point)
+my $session1 = $node->background_psql('postgres');
+$session1->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+);
+
+# Wait until session actually reaches the injection point
+$node->wait_for_event('client backend', 'pg-resize-shared-buffers-flag-set');
+
+# Start second resize session (should fail immediately since resize is in progress)
+my $result2 = $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()");
+
+# The second call should return false (already in progress)
+is($result2, 'f', 'Second concurrent resize call returns false');
+
+# Wake up the first session
+$node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('pg-resize-shared-buffers-flag-set')");
+
+# The pg_resize_shared_buffers() in session1 should now complete successfully
+# We can't easily capture the return value from query_until, but we can
+# verify the session completes without error and the resize actually happened
+$session1->quit;
+
+# Detach injection point
+$node->safe_psql('postgres',
+ "SELECT injection_points_detach('pg-resize-shared-buffers-flag-set')");
+
+done_testing();
diff --git a/src/test/buffermgr/t/004_client_join_buffer_resize.pl b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
new file mode 100644
index 00000000000..072eee535f6
--- /dev/null
+++ b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
@@ -0,0 +1,243 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test shared_buffer resizing coordination with client connections joining using injection points
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+use Time::HiRes qw(sleep);
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Function to calculate the size of test table required to fill up maximum
+# buffer pool when populating it.
+sub calculate_test_sizes
+{
+ my ($node, $block_size) = @_;
+
+ # Get the maximum buffer pool size from configuration
+ my $max_shared_buffers = $node->safe_psql('postgres', "SHOW max_shared_buffers");
+ my ($max_val, $max_unit) = ($max_shared_buffers =~ /(\d+)(\w+)/);
+ my $max_size_bytes;
+ if (lc($max_unit) eq 'kb') {
+ $max_size_bytes = $max_val * 1024;
+ } elsif (lc($max_unit) eq 'mb') {
+ $max_size_bytes = $max_val * 1024 * 1024;
+ } elsif (lc($max_unit) eq 'gb') {
+ $max_size_bytes = $max_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $max_size_bytes = $max_val * 1024;
+ }
+
+ # Fill more pages than minimally required to increase the chances of pages
+ # from the test table filling the buffer cache.
+ $max_size_bytes = $max_size_bytes;
+ my $pages_needed = int($max_size_bytes / $block_size) + 10; # Add some extra to ensure buffers are filled
+ my $rows_to_insert = $pages_needed * 100; # Assuming roughly 100 rows per page for our table structure
+ return ($max_size_bytes, $pages_needed, $rows_to_insert);
+}
+
+# Function to calculate expected buffer count from size string
+sub calculate_buffer_count
+{
+ my ($size_string, $block_size) = @_;
+ # Parse size and convert to bytes
+ my ($size_val, $unit) = ($size_string =~ /(\d+)(\w+)/);
+ my $size_bytes;
+ if (lc($unit) eq 'kb') {
+ $size_bytes = $size_val * 1024;
+ } elsif (lc($unit) eq 'mb') {
+ $size_bytes = $size_val * 1024 * 1024;
+ } elsif (lc($unit) eq 'gb') {
+ $size_bytes = $size_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $size_bytes = $size_val * 1024;
+ }
+ return int($size_bytes / $block_size);
+}
+
+# Initialize cluster with very small buffer sizes for testing
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Configure for buffer resizing with very small buffer pool sizes for faster tests.
+# TODO: for some reason parallel workers try to load default number of shared_buffers which doesn't work with lower max_shared_buffers. We need to fix that - somewhere it's picking default value of shared buffers. For now disable parallelism
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 512kB
+shared_buffers = 320kB
+max_parallel_workers_per_gather = 0
+});
+$node->start;
+
+# Enable injection points
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Get the block size (this is fixed for the binary)
+my $block_size = $node->safe_psql('postgres', "SHOW block_size");
+
+# Try to create pg_buffercache extension for buffer analysis
+eval {
+ $node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+};
+if ($@) {
+ $node->stop;
+ plan skip_all => 'pg_buffercache extension not available - cannot verify buffer usage';
+}
+
+# Create a small test table, and fetch its properties for later reference if required.
+$node->safe_psql('postgres', qq{
+ CREATE TABLE client_test (c1 int, data char(50));
+});
+my $table_oid = $node->safe_psql('postgres', "SELECT oid FROM pg_class WHERE relname = 'client_test'");
+my $table_relfilenode = $node->safe_psql('postgres', "SELECT relfilenode FROM pg_class WHERE relname = 'client_test'");
+note("Test table client_test: OID = $table_oid, relfilenode = $table_relfilenode");
+my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $block_size);
+
+# Create dedicated sessions for injection point handling and test queries,
+# so that we don't create new backends for test operations after starting
+# resize operation. Only one backend, which tests new backend synchronization
+# with resizing operation, should start after resizing has commenced.
+my $injection_session = $node->background_psql('postgres');
+my $query_session = $node->background_psql('postgres');
+my $resize_session = $node->background_psql('postgres');
+
+# Function to run a single injection point test
+sub run_injection_point_test
+{
+ my ($test_name, $injection_point, $target_size, $operation_type) = @_;
+
+ # Silence the logging of the statements we run to avoid
+ # unnecessarily bloating the test logs. This runs before the
+ # upgrade we're testing, so the details should not be very
+ # interesting for debugging. But if needed, you can make it more
+ # verbose by setting this.
+ my $verbose = 0;
+
+ note("Test with $test_name ($operation_type)");
+
+ # Calculate test parameters before starting resize
+ my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $target_size, $block_size);
+
+ # Update buffer pool size and wait for it to reflect pending state
+ $resize_session->query_safe("ALTER SYSTEM SET shared_buffers = '$target_size'", verbose => $verbose);
+ $resize_session->query_safe("SELECT pg_reload_conf()", verbose => $verbose);
+ my $pending_size_str = "pending: $target_size";
+ $resize_session->poll_query_until("SELECT substring(current_setting('shared_buffers'), '$pending_size_str')", $pending_size_str, verbose => $verbose);
+
+ # Set up injection point in injection session
+ $injection_session->query_safe("SELECT injection_points_attach('$injection_point', 'wait')", verbose => $verbose);
+
+ # Trigger resize
+ $resize_session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+ );
+
+ # Wait until resize actually reaches the injection point using the query session
+ $query_session->wait_for_event('client backend', $injection_point, verbose => $verbose);
+
+ # Start a client while resize is paused
+ my $client = $node->background_psql('postgres');
+ note("Background client backend PID: " . $client->query_safe("SELECT pg_backend_pid()", verbose => $verbose));
+
+ # Wake up the injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_wakeup('$injection_point')", verbose => $verbose);
+
+ # Test buffer functionality immediately after waking up injection point
+ # Insert data to test buffer pool functionality during/after resize
+ $client->query_safe("INSERT INTO client_test SELECT i, 'test_data_' || i FROM generate_series(1, $rows_to_insert) i", verbose => $verbose);
+ # Verify the data was inserted correctly and can be read back
+ is($client->query_safe("SELECT COUNT(*) FROM client_test", verbose => $verbose), $rows_to_insert, "inserted $rows_to_insert during $test_name ($operation_type) successful");
+
+ # Verify table size is reasonable (should be substantial for testing)
+ ok($query_session->query_safe("SELECT pg_total_relation_size('client_test')", verbose => $verbose) >= $max_size_bytes,"table size is large enough to overflow buffer pool in test $test_name ($operation_type)");
+
+ # Wait for the resize operation to complete. There is no direct way to do so
+ # in background_psql. Hence fire a psql command and wait for it to finish
+ $resize_session->query(q(\echo 'done'), verbose => $verbose);
+
+ # Detach injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_detach('$injection_point')", verbose => $verbose);
+
+ # Verify resize completed successfully
+ is($query_session->query_safe("SELECT current_setting('shared_buffers')", verbose => $verbose), $target_size,
+ "resize completed successfully to $target_size");
+
+ # Check buffer pool size using pg_buffercache after resize completion
+ is($query_session->query_safe("SELECT COUNT(*) FROM pg_buffercache", verbose => $verbose), calculate_buffer_count($target_size, $block_size), "all buffers in the buffer pool used in $test_name ($operation_type)");
+
+ # Wait for client to complete
+ ok($client->quit, "client succeeded during $test_name ($operation_type)");
+
+ # Clean up for next test
+ $query_session->query_safe("DELETE FROM client_test", verbose => $verbose);
+}
+
+# Test injection points during buffer resize with client connections
+my @common_injection_tests = (
+ {
+ name => 'flag setting phase',
+ injection_point => 'pg-resize-shared-buffers-flag-set',
+ },
+ {
+ name => 'memory remap phase',
+ injection_point => 'pgrsb-after-shmem-resize',
+ },
+ {
+ name => 'resize map barrier complete',
+ injection_point => 'pgrsb-resize-barrier-sent',
+ },
+);
+
+# Test common injection points for both shrinking and expanding
+foreach my $test (@common_injection_tests)
+{
+ # Test shrinking scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '272kB', 'shrinking');
+
+ # Test expanding scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '400kB', 'expanding');
+}
+
+my @shrink_only_tests = (
+ {
+ name => 'shrink barrier complete',
+ injection_point => 'pgrsb-shrink-barrier-sent',
+ size => '200kB',
+ }
+);
+foreach my $test (@shrink_only_tests)
+{
+ run_injection_point_test($test->{name}, $test->{injection_point}, $test->{size}, 'shrinking only');
+}
+
+my @expand_only_tests = (
+ {
+ name => 'expand barrier complete',
+ injection_point => 'pgrsb-expand-barrier-sent',
+ size => '416kB',
+ }
+);
+
+foreach my $test (@expand_only_tests)
+{
+ run_injection_point_test($test->{name}, $test->{injection_point}, $test->{size}, 'expanding only');
+}
+
+$injection_session->quit;
+$query_session->quit;
+$resize_session->quit;
+
+done_testing();
diff --git a/src/test/meson.build b/src/test/meson.build
index cd45cbf57fb..e9550933063 100644
--- a/src/test/meson.build
+++ b/src/test/meson.build
@@ -4,6 +4,7 @@ subdir('regress')
subdir('isolation')
subdir('authentication')
+subdir('buffermgr')
subdir('postmaster')
subdir('recovery')
subdir('subscription')
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index 5bd41a278dd..0482def8db8 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -61,6 +61,7 @@ use Config;
use IPC::Run;
use PostgreSQL::Test::Utils qw(pump_until);
use Test::More;
+use Time::HiRes qw(usleep);
=pod
@@ -376,4 +377,79 @@ sub set_query_timer_restart
return $self->{query_timer_restart};
}
+=pod
+
+=item $session->poll_query_until($query [, $expected ])
+
+Run B<$query> repeatedly in this background session, until it returns the
+B<$expected> result ('t', or SQL boolean true, by default).
+Continues polling if the query returns an error result.
+Times out after a reasonable number of attempts.
+Returns 1 if successful, 0 if timed out.
+
+=cut
+
+sub poll_query_until
+{
+ my ($self, $query, $expected, %params) = @_;
+
+ $expected = 't' unless defined($expected); # default value
+
+ my $max_attempts = 10 * $PostgreSQL::Test::Utils::timeout_default;
+ my $attempts = 0;
+ my ($stdout, $stderr_flag);
+
+ while ($attempts < $max_attempts)
+ {
+ ($stdout, $stderr_flag) = $self->query($query, %params);
+
+ chomp($stdout);
+
+ # If query succeeded and returned expected result
+ if (!$stderr_flag && $stdout eq $expected)
+ {
+ return 1;
+ }
+
+ # Wait 0.1 second before retrying.
+ usleep(100_000);
+
+ $attempts++;
+ }
+
+ # Give up. Print the output from the last attempt, hopefully that's useful
+ # for debugging.
+ my $stderr_output = $stderr_flag ? $self->{stderr} : '';
+ diag qq(poll_query_until timed out executing this query:
+$query
+expecting this output:
+$expected
+last actual query output:
+$stdout
+with stderr:
+$stderr_output);
+ return 0;
+}
+
+=item $session->wait_for_event(backend_type, wait_event_name)
+
+Poll pg_stat_activity until backend_type reaches wait_event_name using this
+background session.
+
+=cut
+
+sub wait_for_event
+{
+ my ($self, $backend_type, $wait_event_name, %params) = @_;
+
+ $self->poll_query_until(qq[
+ SELECT count(*) > 0 FROM pg_stat_activity
+ WHERE backend_type = '$backend_type' AND wait_event = '$wait_event_name'
+ ], undef, %params)
+ or die
+ qq(timed out when waiting for $backend_type to reach wait event '$wait_event_name');
+
+ return;
+}
+
1;
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index d1d4f2b41b4..a01e95a0fb2 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2808,6 +2808,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShmemSegment
ShutdownForeignScan_function
ShutdownInformation
--
2.34.1
[text/x-patch] v20260128-0002-Memory-and-address-space-management-for-bu.patch (101.6K, ../../CAExHW5s8s=UhjqNa_Tz1PFCRLzt3=5nvd5vD1wFdKWMQCmFySQ@mail.gmail.com/4-v20260128-0002-Memory-and-address-space-management-for-bu.patch)
download | inline diff:
From c38ff20e91dcc504ca05c51d9743e1a9c5c6c2bb Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Fri, 28 Feb 2025 19:54:47 +0100
Subject: [PATCH v20260128 2/5] Memory and address space management for buffer
resizing
This has three changes
1. Allow to use multiple shared memory mappings
============================================
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
A new shared memory API is introduced, extended with a segment as a new
parameter. As a path of least resistance, the original API is kept in
place, utilizing the main shared memory segment.
Modifies pg_shmem_allocations to report shared memory segment as well.
Adds pg_shmem_segments to report shared memory segment information.
2. Address space reservation for shared memory
============================================
Currently the shared memory layout is designed to pack everything tight
together, leaving no space between mappings for resizing. Here is how it
looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
...
Make the layout more dynamic via splitting every shared memory segment
into two parts:
* An anonymous file, which actually contains shared memory content.
Such an anonymous file is created via memfd_create, it lives in
memory, behaves like a regular file and semantically equivalent to an
anonymous memory allocated via mmap with MAP_ANONYMOUS.
* A reservation mapping, which size is much larger than required shared
segment size. This mapping is created with flag MAP_NORESERVE (to not
count the reserved space against memory limits). The anonymous file is
mapped into this reservation mapping.
If we have to change the address maps while resizing the shared buffer
pool, it is needed to be done in Postmaster too, so that the new
backends will inherit the resized address space from the Postmaster.
However, Postmaster is not invovled in ProcSignalBarrier mechanism and
we don't want it to spend time in things other than its core
functionality. To achive that, maximum required address space maps are
setup upfront with read and write access when starting the server. When
resizing the buffer pool only the backing file object is resized from
the coordinator. This also makes the ProcSignalBarrier handling code
light for backends other than the coordinator.
The resulting layout looks like this:
00400000-00490000 /path/bin/postgres
...
3f526000-3f590000 rw-p [heap]
7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
To resize a shared memory segment in this layout it's possible to use
ftruncate on the memory mapped file.
This approach also do not impact the actual memory usage as reported by
the kernel.
TODO: Verify that Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched within
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
# postgres from the master branch has being successfully launched
# from that shell
$ cat memory.current
17465344 (~16.6 MB)
# stop postgres
$ echo $PATCH_PID_SHELL > cgroup.procs
# postgres from the patch has being successfully launched from that shell
$ cat memory.current
20770816 (~19.8 MB)
There are also few unrelated advantages of using memory mapped files:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps.
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process. - Some hackers have expressed concerns over it.
The downside is that memfd_create is Linux specific.
3. Refactor CalculateShmemSize()
=============================
This function calls many functions which return the amount of shared
memory required for different shared memory data structures. Up until
now, the returned total of these sizes was used to create a single
shared memory segment. With this change, CalculateShmemSize() needs to
estimate memory requirements for each of the segments. It now takes an
array of MemoryMappingSizes, containing as many elements as the number
of segments, as an argument. The sizes returned by all the function it
calls, except BufferManagerShmemSize(), are added and saved in the first
element (index 0) of the array. BufferManagerShmemSize() is modified to
save the amount of memory required for buffer manager related segments
in the corresponding array element. Additionally it also saves the
amount of reserved space. For now, the amount of reserved address space
is same as the amount of required memory but that is expected to change
with the next commit which implements buffer pool resize.
CalculateShmemSize() now returns the total of sizes corresponding to all
the sizes.
Author: Dmitrii Dolgov and Ashutosh Bapat
Reviewed-by: Tomas Vondra
---
doc/src/sgml/system-views.sgml | 9 +
src/backend/catalog/system_views.sql | 7 +
src/backend/port/posix_sema.c | 2 +-
src/backend/port/sysv_sema.c | 2 +-
src/backend/port/sysv_shmem.c | 552 ++++++++++++++++------
src/backend/port/win32_sema.c | 2 +-
src/backend/port/win32_shmem.c | 291 +++++++-----
src/backend/postmaster/launch_backend.c | 36 +-
src/backend/storage/buffer/buf_init.c | 64 ++-
src/backend/storage/buffer/buf_table.c | 1 +
src/backend/storage/buffer/freelist.c | 7 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 106 ++++-
src/backend/storage/ipc/shmem.c | 339 +++++++++----
src/backend/storage/lmgr/lwlock.c | 17 +-
src/backend/storage/lmgr/predicate.c | 3 +-
src/backend/utils/activity/pgstat_shmem.c | 3 +-
src/include/catalog/pg_proc.dat | 12 +-
src/include/storage/bufmgr.h | 3 +-
src/include/storage/ipc.h | 4 +-
src/include/storage/pg_shmem.h | 112 ++++-
src/include/storage/shmem.h | 13 +-
src/test/regress/expected/rules.out | 9 +-
src/tools/pgindent/typedefs.list | 4 +
24 files changed, 1138 insertions(+), 464 deletions(-)
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index c5683068470..6fa47e3c63d 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -4305,6 +4305,15 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
</para></entry>
</row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>segment</structfield> <type>text</type>
+ </para>
+ <para>
+ The name of the shared memory segment concerning the allocation.
+ </para></entry>
+ </row>
+
<row>
<entry role="catalog_table_entry"><para role="column_definition">
<structfield>off</structfield> <type>int8</type>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 7553f31fef0..bc11589aeab 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -668,6 +668,13 @@ GRANT SELECT ON pg_shmem_allocations TO pg_read_all_stats;
REVOKE EXECUTE ON FUNCTION pg_get_shmem_allocations() FROM PUBLIC;
GRANT EXECUTE ON FUNCTION pg_get_shmem_allocations() TO pg_read_all_stats;
+CREATE VIEW pg_shmem_segments AS
+ SELECT * FROM pg_get_shmem_segments();
+
+REVOKE ALL ON pg_shmem_segments FROM PUBLIC;
+GRANT SELECT ON pg_shmem_segments TO pg_read_all_stats;
+REVOKE EXECUTE ON FUNCTION pg_get_shmem_segments() FROM PUBLIC;
+GRANT EXECUTE ON FUNCTION pg_get_shmem_segments() TO pg_read_all_stats;
CREATE VIEW pg_shmem_allocations_numa AS
SELECT * FROM pg_get_shmem_allocations_numa();
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index e368e5ee7ed..5ad50c79dcd 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -216,7 +216,7 @@ PGReserveSemaphores(int maxSemas)
#else
sharedSemas = (PGSemaphore)
- ShmemAlloc(PGSemaphoreShmemSize(maxSemas));
+ ShmemAlloc(MAIN_SHMEM_SEGMENT, PGSemaphoreShmemSize(maxSemas));
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 86c4d359ef7..f0c7b064ffb 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -344,7 +344,7 @@ PGReserveSemaphores(int maxSemas)
DataDir)));
sharedSemas = (PGSemaphore)
- ShmemAlloc(PGSemaphoreShmemSize(maxSemas));
+ ShmemAlloc(MAIN_SHMEM_SEGMENT, PGSemaphoreShmemSize(maxSemas));
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 3cd3544fa2b..88fdcf854a7 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -39,7 +39,17 @@
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
-
+/*
+ * TODO: The first two sentences in the first paragraph below make me feel like
+ * we should have only one SysV segment. Is that true? Needs investigation.
+ */
+/*
+ * TODO: third paragraph should mention that we use memfd_create to create
+ * shared memory segment, and possibly there's a way to share that segment
+ * between two processes using the file descriptor instead of going through SysV
+ * shared memory segment. So one day EXEC_BACKEND can also use anonymous shared
+ * memory.
+ */
/*
* As of PostgreSQL 9.3, we normally allocate only a very small amount of
* System V shared memory, and only for the purposes of providing an
@@ -91,12 +101,56 @@ typedef enum
SHMSTATE_UNATTACHED, /* pertinent to DataDir, no attached PIDs */
} IpcMemoryState;
+/*
+ * Anonymous mapping layout we use looks like this:
+ *
+ * 00400000-00c2a000 r-xp /bin/postgres
+ * ...
+ * 3f526000-3f590000 rw-p [heap]
+ * 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted)
+ * 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted)
+ * 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
+ * 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
+ * ...
+ *
+ * We need to place shared memory mappings in such a way, that there will be
+ * gaps between them in the address space. Those gaps have to be large enough
+ * to resize the mapping up to certain size, without counting towards the total
+ * memory consumption.
+ *
+ * To achieve this, for each shared memory segment we first create an anonymous
+ * file of specified size using memfd_create, which will accomodate actual
+ * shared memory mapping content. It is represented by the first /memfd:main
+ * with rw permissions. Then we create a mapping for this file using mmap, with
+ * size much larger than required and flags PROT_NONE (allows to make sure the
+ * reserved space will not be used) and MAP_NORESERVE (prevents the space from
+ * being counted against memory limits). The mapping serves as an address space
+ * reservation, into which shared memory segment can be extended and is
+ * represented by the second /memfd:main with no permissions.
+ */
+
+PGInhShmemSeg InhShmemSegs[NUM_MEMORY_MAPPINGS];
+
+ /*
+ * Structure to hold anonymous shared memory segment properties.
+ */
+typedef struct AnonShmemSegment
+{
+ int fd; /* fd for the backing anon file */
+ void *addr; /* Pointer to the start of the mapped memory */
+ Size size; /* Size of the mapped memory */
-unsigned long UsedShmemSegID = 0;
-void *UsedShmemSegAddr = NULL;
+} AnonShmemSegment;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+AnonShmemSegment AnonShmemSegs[NUM_MEMORY_MAPPINGS];
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -471,19 +525,20 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* hugepage sizes, we might want to think about more invasive strategies,
* such as increasing shared_buffers to absorb the extra space.
*
- * Returns the (real, assumed or config provided) page size into
- * *hugepagesize, and the hugepage-related mmap flags to use into
- * *mmap_flags if requested by the caller. If huge pages are not supported,
- * *hugepagesize and *mmap_flags are set to 0.
+ * Returns the (real, assumed or config provided) page size into *hugepagesize,
+ * the hugepage-related mmap and memfd flags to use into *mmap_flags and
+ * *memfd_flags if requested by the caller. If huge pages are not supported,
+ * *hugepagesize, *mmap_flags and *memfd_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -542,6 +597,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -556,11 +612,22 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
#endif
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MFD_HUGE_MASK) << MFD_HUGE_SHIFT;
+ }
+#endif
+
/* assign the results found */
if (mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -568,6 +635,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -589,84 +658,266 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
return true;
}
+/*
+ * Wrapper around posix_fallocate() to allocate memory for a given shared memory
+ * segment.
+ *
+ * Performs retry on EINTR, and raises error upon failure.
+ */
+static void
+shmem_fallocate(int fd, const char *mapping_name, Size size, int elevel)
+{
+#if defined(HAVE_POSIX_FALLOCATE) && defined(__linux__)
+ int ret;
+
+
+ /*
+ * If there is not enough memory, trying to access a hole in address space
+ * will cause SIGBUS. If supported, avoid that by allocating memory
+ * upfront.
+ *
+ * We still use a traditional EINTR retry loop to handle SIGCONT.
+ * posix_fallocate() doesn't restart automatically, and we don't want this
+ * to fail if you attach a debugger.
+ */
+ do
+ {
+ ret = posix_fallocate(fd, 0, size);
+ } while (ret == EINTR);
+
+ if (ret != 0)
+ {
+ ereport(elevel,
+ (errmsg("segment[%s]: could not allocate space for anonymous file: %s",
+ mapping_name, strerror(ret)),
+ (ret == ENOMEM) ?
+ errhint("This error usually means that PostgreSQL's request "
+ "for a shared memory segment exceeded available memory, "
+ "swap space, or huge pages. To reduce the request size "
+ "(currently %zu bytes), reduce PostgreSQL's shared "
+ "memory usage, perhaps by reducing \"shared_buffers\" or "
+ "\"max_connections\".",
+ size) : 0));
+ }
+#endif
+}
+
+/*
+ * Round up the required amount of memory and the amount of required reserved
+ * address space to the nearest huge page size.
+ */
+static inline void
+round_off_mapping_sizes_for_hugepages(MemoryMappingSizes *mapping, int hugepagesize)
+{
+ if (hugepagesize == 0)
+ return;
+
+ if (mapping->shmem_req_size % hugepagesize != 0)
+ mapping->shmem_req_size += add_size(mapping->shmem_req_size,
+ hugepagesize - (mapping->shmem_req_size % hugepagesize));
+
+ if (mapping->shmem_reserved % hugepagesize != 0)
+ mapping->shmem_reserved = add_size(mapping->shmem_reserved,
+ hugepagesize - (mapping->shmem_reserved % hugepagesize));
+}
+
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested. If needed,
+ * it also rounds up the mapping reserved size to be a multiple of huge page
+ * size.
+ *
+ * Note that we do not fallback from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
+ *
+ * TODO: Update the prologue to be consistent with the code.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(int segment_id, MemoryMappingSizes *mapping)
{
- Size allocsize = *size;
void *ptr = MAP_FAILED;
- int mmap_errno = 0;
- int mmap_flags = MAP_SHARED | MAP_ANONYMOUS | MAP_HASSEMAPHORE;
+ int mmap_flags = MAP_SHARED | MAP_HASSEMAPHORE | MAP_NORESERVE;
+ AnonShmemSegment *anonseg = &AnonShmemSegs[segment_id];
+ const char *segname = MappingName(segment_id);
+ int memfd_flags = 0;
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
int huge_mmap_flags;
+ int huge_memfd_flags;
- GetHugePageSize(&hugepagesize, &huge_mmap_flags);
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
- if (allocsize % hugepagesize != 0)
- allocsize = add_size(allocsize, hugepagesize - (allocsize % hugepagesize));
+ /* Round up the request size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &huge_mmap_flags, &huge_memfd_flags);
+ round_off_mapping_sizes_for_hugepages(mapping, hugepagesize);
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | huge_mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ /* Verify that the new size is withing the reserved boundaries */
+ Assert(mapping->shmem_reserved >= mapping->shmem_req_size);
+
+ mmap_flags = mmap_flags | huge_mmap_flags;
+ memfd_flags = memfd_flags | huge_memfd_flags;
}
#endif
/*
- * Report whether huge pages are in use. This needs to be tracked before
- * the second mmap() call if attempting to use huge pages failed
- * previously.
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ *
+ * TODO: Need a configuration test for memfd_create.
+ *
+ * TODO: Earlier releases did not use file backed shared memory segments.
+ * By setting bit 1 in /proc/<PID>/coredump_filter, those shared memory
+ * segments could be dumped to the core file. But dumping file backed
+ * shared memory segments requires bit 3 to be set. We need to document
+ * this change in the release notes.
*/
- SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
- PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ anonseg->fd = memfd_create(segname, memfd_flags);
+ if (anonseg->fd == -1)
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not create anonymous shared memory file: %m",
+ segname)));
- if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
- {
- /*
- * Use the original size, not the rounded-up value, when falling back
- * to non-huge pages.
- */
- allocsize = *size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags, -1, 0);
- mmap_errno = errno;
- }
+ elog(DEBUG1, "segment[%s]: mmap(%zu)", segname, mapping->shmem_req_size);
+ /*
+ * Reserve maximum required address space for future expansion of this
+ * memory segment. The whole address space will be setup for read/write
+ * access, so that memory allocated to this address space can be read or
+ * written to even if it is resized in the future using just ftruncate.
+ * MAP_NORESERVE alone should ensure that no memory is allocated. But when
+ * using huge pages, the memory is allocated at mmap time if PROT_WRITE |
+ * PROT_READ is used. Hence we create the mapping with PROT_NONE first and
+ * then use mprotect to set the required permissions.
+ */
+ ptr = mmap(NULL, mapping->shmem_reserved, PROT_NONE,
+ mmap_flags, anonseg->fd, 0);
if (ptr == MAP_FAILED)
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ segname)));
+
+ if (mprotect(ptr, mapping->shmem_reserved, PROT_READ | PROT_WRITE) == -1)
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not update anonymous shared memory permissions: %m",
+ segname)));
+
+
+ /*
+ * Resize the backing file to the required size. On platforms where it is
+ * supported, we also allocate the required memory upfront. On other
+ * platform the memory upto the size of file will be allocated on demand.
+ */
+ if (ftruncate(anonseg->fd, mapping->shmem_req_size) == -1)
{
- errno = mmap_errno;
+ int save_errno = errno;
+
+ close(anonseg->fd);
+ anonseg->fd = -1;
+
+ errno = save_errno;
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
- (mmap_errno == ENOMEM) ?
+ (errmsg("segment[%s]: could not truncate anonymous file to size %zu: %m",
+ segname, mapping->shmem_req_size),
+ (save_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
"swap space, or huge pages. To reduce the request size "
"(currently %zu bytes), reduce PostgreSQL's shared "
"memory usage, perhaps by reducing \"shared_buffers\" or "
"\"max_connections\".",
- allocsize) : 0));
+ mapping->shmem_req_size) : 0));
}
+ shmem_fallocate(anonseg->fd, segname, mapping->shmem_req_size, FATAL);
- *size = allocsize;
- return ptr;
+ anonseg->addr = ptr;
+ anonseg->size = mapping->shmem_reserved;
+}
+
+/*
+ * PrepareHugePages
+ *
+ * Figure out if there are enough huge pages to allocate all shared memory
+ * segments, and report that information via huge_pages_status and
+ * huge_pages_on. It needs to be called before creating shared memory segments.
+ *
+ * It is necessary to maintain the same semantic (simple on/off) for
+ * huge_pages_status, even if there are multiple shared memory segments: all
+ * segments either use huge pages or not, there is no mix of segments with
+ * different page size. The latter might be actually beneficial, in particular
+ * because only some segments may require large amount of memory, but for now
+ * we go with a simple solution.
+ */
+void
+PrepareHugePages()
+{
+ void *ptr = MAP_FAILED;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+ int mmap_flags = (MAP_SHARED | MAP_HASSEMAPHORE);
+
+ CalculateShmemSize(mapping_sizes);
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize,
+ total_size = 0;
+ int huge_mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &huge_mmap_flags, NULL);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory
+ * and decide whether we will allocate huge pages or not.
+ */
+ for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ Size segment_size = mapping_sizes[segment].shmem_req_size;
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ mmap_flags | MAP_ANONYMOUS | huge_mmap_flags, -1, 0);
+ }
+#endif
+
+ /*
+ * Report whether huge pages are in use. This needs to be tracked before
+ * creating shared memory segments.
+ */
+ SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
}
/*
@@ -676,20 +927,29 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonShmemSegment *segment = &AnonShmemSegs[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (segment->addr != NULL)
+ {
+ Assert(segment->fd != -1);
+
+ if (munmap(segment->addr, segment->size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ segment->addr, segment->size);
+ segment->addr = NULL;
+ close(segment->fd);
+ segment->fd = -1;
+ }
}
}
/*
* PGSharedMemoryCreate
*
- * Create a shared memory segment of the given size and initialize its
+ * Create a shared memory segment for the given mapping and initialize its
* standard header. Also, register an on_shmem_exit callback to release
* the storage.
*
@@ -699,7 +959,7 @@ AnonymousShmemDetach(int status, Datum arg)
* postmaster or backend.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping,
PGShmemHeader **shim)
{
IpcMemoryKey NextShmemSegID;
@@ -707,6 +967,8 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonShmemSegment *anonseg = &AnonShmemSegs[segment_id];
+ PGInhShmemSeg *inhseg = &InhShmemSegs[segment_id];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -719,14 +981,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -734,12 +988,12 @@ PGSharedMemoryCreate(Size size,
errmsg("huge pages not supported with the current \"shared_memory_type\" setting")));
/* Room for a header? */
- Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ Assert(mapping->shmem_req_size > MAXALIGN(sizeof(PGShmemHeader)));
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(segment_id, mapping);
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -749,7 +1003,7 @@ PGSharedMemoryCreate(Size size,
}
else
{
- sysvsize = size;
+ sysvsize = mapping->shmem_req_size;
/* huge pages are only available with mmap */
SetConfigOption("huge_pages_status", "off",
@@ -762,7 +1016,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + inhseg->UsedShmemSegID;
for (;;)
{
@@ -800,6 +1054,8 @@ PGSharedMemoryCreate(Size size,
errmsg("pre-existing shared memory block (key %lu, ID %lu) is still in use",
(unsigned long) NextShmemSegID,
(unsigned long) shmid),
+ errdetail("when trying to create shared memory block for segment \"%s\"",
+ MappingName(segment_id)),
errhint("Terminate any old server processes associated with data directory \"%s\".",
DataDir)));
break;
@@ -854,24 +1110,24 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_req_size;
+ hdr->ReservedSize = mapping->shmem_reserved;
hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ inhseg->UsedShmemSegAddr = memAddress;
+ inhseg->UsedShmemSegID = (unsigned long) NextShmemSegID;
/*
- * If AnonymousShmem is NULL here, then we're not using anonymous shared
- * memory, and should return a pointer to the System V shared memory
- * block. Otherwise, the System V shared memory block is only a shim, and
- * we must return a pointer to the real block.
+ * If we're not using anonymous shared memory, return a pointer to the
+ * System V shared memory block. Otherwise, the System V shared memory
+ * block is only a shim, and we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (anonseg->addr == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(anonseg->addr, hdr, sizeof(PGShmemHeader));
+ return anonseg->addr;
}
#ifdef EXEC_BACKEND
@@ -884,9 +1140,9 @@ PGSharedMemoryCreate(Size size,
* EXEC_BACKEND case; otherwise postmaster children inherit the shared memory
* segment attachment via fork().
*
- * UsedShmemSegID and UsedShmemSegAddr are implicit parameters to this
- * routine. The caller must have already restored them to the postmaster's
- * values.
+ * Segments array is an implicit parameter to this
+ * routine. The caller must have already restored it to the postmaster's
+ * state.
*/
void
PGSharedMemoryReAttach(void)
@@ -894,32 +1150,42 @@ PGSharedMemoryReAttach(void)
IpcMemoryId shmid;
PGShmemHeader *hdr;
IpcMemoryState state;
- void *origUsedShmemSegAddr = UsedShmemSegAddr;
+ void *origUsedShmemSegAddr;
- Assert(UsedShmemSegAddr != NULL);
- Assert(IsUnderPostmaster);
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGInhShmemSeg *inhseg = &InhShmemSegs[i];
+
+ origUsedShmemSegAddr = inhseg->UsedShmemSegAddr;
+
+ Assert(inhseg->UsedShmemSegAddr != NULL);
+ Assert(IsUnderPostmaster);
#ifdef __CYGWIN__
- /* cygipc (currently) appears to not detach on exec. */
- PGSharedMemoryDetach();
- UsedShmemSegAddr = origUsedShmemSegAddr;
+ /* cygipc (currently) appears to not detach on exec. */
+ PGSharedMemoryDetach();
+ inhseg->UsedShmemSegAddr = origUsedShmemSegAddr;
#endif
- elog(DEBUG3, "attaching to %p", UsedShmemSegAddr);
- shmid = shmget(UsedShmemSegID, sizeof(PGShmemHeader), 0);
- if (shmid < 0)
- state = SHMSTATE_FOREIGN;
- else
- state = PGSharedMemoryAttach(shmid, UsedShmemSegAddr, &hdr);
- if (state != SHMSTATE_ATTACHED)
- elog(FATAL, "could not reattach to shared memory (key=%d, addr=%p): %m",
- (int) UsedShmemSegID, UsedShmemSegAddr);
- if (hdr != origUsedShmemSegAddr)
- elog(FATAL, "reattaching to shared memory returned unexpected address (got %p, expected %p)",
- hdr, origUsedShmemSegAddr);
- dsm_set_control_handle(hdr->dsm_control);
-
- UsedShmemSegAddr = hdr; /* probably redundant */
+ elog(DEBUG3, "attaching to %p", inhseg->UsedShmemSegAddr);
+ shmid = shmget(inhseg->UsedShmemSegID, sizeof(PGShmemHeader), 0);
+ if (shmid < 0)
+ state = SHMSTATE_FOREIGN;
+ else
+ state = PGSharedMemoryAttach(shmid, inhseg->UsedShmemSegAddr, &hdr);
+ if (state != SHMSTATE_ATTACHED)
+ elog(FATAL, "could not reattach to shared memory (key=%d, addr=%p): %m",
+ (int) inhseg->UsedShmemSegID, inhseg->UsedShmemSegAddr);
+ if (hdr != origUsedShmemSegAddr)
+ elog(FATAL, "reattaching to shared memory returned unexpected address (got %p, expected %p)",
+ hdr, origUsedShmemSegAddr);
+
+ /* Re-establish dsm_control mapping, if any */
+ if (hdr->dsm_control != 0)
+ dsm_set_control_handle(hdr->dsm_control);
+
+ inhseg->UsedShmemSegAddr = hdr; /* probably redundant */
+ }
}
/*
@@ -933,14 +1199,13 @@ PGSharedMemoryReAttach(void)
* The child process startup logic might or might not call PGSharedMemoryDetach
* after this; make sure that it will be a no-op if called.
*
- * UsedShmemSegID and UsedShmemSegAddr are implicit parameters to this
- * routine. The caller must have already restored them to the postmaster's
- * values.
+ * Segments array is an implicit parameter to this
+ * routine. The caller must have already restored it to the postmaster's
+ * state.
*/
void
PGSharedMemoryNoReAttach(void)
{
- Assert(UsedShmemSegAddr != NULL);
Assert(IsUnderPostmaster);
#ifdef __CYGWIN__
@@ -948,10 +1213,16 @@ PGSharedMemoryNoReAttach(void)
PGSharedMemoryDetach();
#endif
- /* For cleanliness, reset UsedShmemSegAddr to show we're not attached. */
- UsedShmemSegAddr = NULL;
- /* And the same for UsedShmemSegID. */
- UsedShmemSegID = 0;
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGInhShmemSeg *inhseg = &InhShmemSegs[i];
+
+ Assert(inhseg->UsedShmemSegAddr != NULL);
+ /* For cleanliness, reset UsedShmemSegAddr to show we're not attached. */
+ inhseg->UsedShmemSegAddr = NULL;
+ /* And the same for UsedShmemSegID. */
+ inhseg->UsedShmemSegID = 0;
+ }
}
#endif /* EXEC_BACKEND */
@@ -959,35 +1230,44 @@ PGSharedMemoryNoReAttach(void)
/*
* PGSharedMemoryDetach
*
- * Detach from the shared memory segment, if still attached. This is not
+ * Detach from the shared memory segments, if still attached. This is not
* intended to be called explicitly by the process that originally created the
- * segment (it will have on_shmem_exit callback(s) registered to do that).
+ * segments (it will have on_shmem_exit callback(s) registered to do that).
* Rather, this is for subprocesses that have inherited an attachment and want
* to get rid of it.
*
- * UsedShmemSegID and UsedShmemSegAddr are implicit parameters to this
- * routine, also AnonymousShmem and AnonymousShmemSize.
+ * PGInhShmemSeg::UsedShmemSegID and PGInhShmemSeg::UsedShmemSegAddr are
+ * implicit parameters to this routine obtained from entries in InhShmemSegs
+ * array.
*/
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ PGInhShmemSeg *inhseg = &InhShmemSegs[i];
+ AnonShmemSegment *anonseg = &AnonShmemSegs[i];
+
+ if (inhseg->UsedShmemSegAddr != NULL)
+ {
+ if ((shmdt(inhseg->UsedShmemSegAddr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", inhseg->UsedShmemSegAddr);
+ inhseg->UsedShmemSegAddr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (anonseg->addr != NULL)
+ {
+ if (munmap(anonseg->addr, anonseg->size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ anonseg->addr, anonseg->size);
+ anonseg->addr = NULL;
+ close(anonseg->fd);
+ anonseg->fd = -1;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index ba97c9b2d64..4683736415b 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 7cb8b4c9b60..5ee36063aff 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -39,15 +39,14 @@
* address space and is negligible relative to the 64-bit address space.
*/
#define PROTECTIVE_REGION_SIZE (10 * WIN32_STACK_RLIMIT)
-void *ShmemProtectiveRegion = NULL;
-
-HANDLE UsedShmemSegID = INVALID_HANDLE_VALUE;
-void *UsedShmemSegAddr = NULL;
-static Size UsedShmemSegSize = 0;
static bool EnableLockPagesPrivilege(int elevel);
static void pgwin32_SharedMemoryDelete(int status, Datum shmId);
+PGInhShmemSeg InhShmemSegs[NUM_MEMORY_MAPPINGS];
+
+static Size UsedShmemSegSizes[NUM_MEMORY_MAPPINGS] = {0};
+
/*
* Generate shared memory segment name. Expand the data directory, to generate
* an identifier unique for this data directory. Then replace all backslashes
@@ -202,9 +201,11 @@ EnableLockPagesPrivilege(int elevel)
*
* Create a shared memory segment of the given size and initialize its
* standard header.
+ *
+ * TODO: Check that the segment_id is a valid one before indexing corresponding arrays.
*/
-PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+void
+PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping_sizes,
PGShmemHeader **shim)
{
void *memAddress;
@@ -216,13 +217,14 @@ PGSharedMemoryCreate(Size size,
DWORD size_high;
DWORD size_low;
SIZE_T largePageSize = 0;
- Size orig_size = size;
+ Size size = mapping_sizes->shmem_req_size;
DWORD flProtect = PAGE_READWRITE;
DWORD desiredAccess;
+ PGInhShmemSeg *inhseg = &InhShmemSegs[segment_id];
- ShmemProtectiveRegion = VirtualAlloc(NULL, PROTECTIVE_REGION_SIZE,
- MEM_RESERVE, PAGE_NOACCESS);
- if (ShmemProtectiveRegion == NULL)
+ inhseg->ShmemProtectiveRegion = VirtualAlloc(NULL, PROTECTIVE_REGION_SIZE,
+ MEM_RESERVE, PAGE_NOACCESS);
+ if (inhseg->ShmemProtectiveRegion == NULL)
elog(FATAL, "could not reserve memory region: error code %lu",
GetLastError());
@@ -231,8 +233,12 @@ PGSharedMemoryCreate(Size size,
szShareMem = GetSharedMemName();
- UsedShmemSegAddr = NULL;
+ inhseg->UsedShmemSegAddr = NULL;
+ /*
+ * TODO: We don't need to perform this as many times as the number of
+ * segments. Instead do something similar to sysv_shmem.c
+ */
if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
{
/* Does the processor support large pages? */
@@ -304,7 +310,7 @@ retry:
* Use the original size, not the rounded-up value, when
* falling back to non-huge pages.
*/
- size = orig_size;
+ size = mapping_sizes->shmem_req_size;
flProtect = PAGE_READWRITE;
goto retry;
}
@@ -337,6 +343,8 @@ retry:
if (!hmap)
ereport(FATAL,
(errmsg("pre-existing shared memory block is still in use"),
+ errdetail("when trying to create shared memory block for segment \"%s\"",
+ PGShmemSegmentName(segment)),
errhint("Check if there are any old server processes still running, and terminate them.")));
free(szShareMem);
@@ -393,9 +401,9 @@ retry:
hdr->dsm_control = 0;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegSize = size;
- UsedShmemSegID = hmap2;
+ inhseg->UsedShmemSegAddr = memAddress;
+ UsedShmemSegSizes[segment_id] = size;
+ inhseg->UsedShmemSegID = (unsigned long) hmap2;
/* Register on-exit routine to delete the new segment */
on_shmem_exit(pgwin32_SharedMemoryDelete, PointerGetDatum(hmap2));
@@ -405,8 +413,6 @@ retry:
/* Report whether huge pages are in use */
SetConfigOption("huge_pages_status", (flProtect & SEC_LARGE_PAGES) ?
"on" : "off", PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
-
- return hdr;
}
/*
@@ -416,42 +422,52 @@ retry:
* an already existing shared memory segment, using the handle inherited from
* the postmaster.
*
- * ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr are implicit
- * parameters to this routine. The caller must have already restored them to
- * the postmaster's values.
+ * Segments is an implicit parameters to this routine. The caller must have
+ * already restored ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr
+ * in each Segment to the postmaster's values.
*/
void
PGSharedMemoryReAttach(void)
{
PGShmemHeader *hdr;
- void *origUsedShmemSegAddr = UsedShmemSegAddr;
+ void *origUsedShmemSegAddr;
- Assert(ShmemProtectiveRegion != NULL);
- Assert(UsedShmemSegAddr != NULL);
Assert(IsUnderPostmaster);
- /*
- * Release memory region reservations made by the postmaster
- */
- if (VirtualFree(ShmemProtectiveRegion, 0, MEM_RELEASE) == 0)
- elog(FATAL, "failed to release reserved memory region (addr=%p): error code %lu",
- ShmemProtectiveRegion, GetLastError());
- if (VirtualFree(UsedShmemSegAddr, 0, MEM_RELEASE) == 0)
- elog(FATAL, "failed to release reserved memory region (addr=%p): error code %lu",
- UsedShmemSegAddr, GetLastError());
-
- hdr = (PGShmemHeader *) MapViewOfFileEx(UsedShmemSegID, FILE_MAP_READ | FILE_MAP_WRITE, 0, 0, 0, UsedShmemSegAddr);
- if (!hdr)
- elog(FATAL, "could not reattach to shared memory (key=%p, addr=%p): error code %lu",
- UsedShmemSegID, UsedShmemSegAddr, GetLastError());
- if (hdr != origUsedShmemSegAddr)
- elog(FATAL, "reattaching to shared memory returned unexpected address (got %p, expected %p)",
- hdr, origUsedShmemSegAddr);
- if (hdr->magic != PGShmemMagic)
- elog(FATAL, "reattaching to shared memory returned non-PostgreSQL memory");
- dsm_set_control_handle(hdr->dsm_control);
-
- UsedShmemSegAddr = hdr; /* probably redundant */
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGInhShmemSeg *inhseg = &InhShmemSegs[i];
+
+ Assert(inhseg->ShmemProtectiveRegion != NULL);
+ Assert(inhseg->UsedShmemSegAddr != NULL);
+
+ origUsedShmemSegAddr = inhseg->UsedShmemSegAddr;
+
+ /*
+ * Release memory region reservations made by the postmaster
+ */
+ if (VirtualFree(inhseg->ShmemProtectiveRegion, 0, MEM_RELEASE) == 0)
+ elog(FATAL, "failed to release reserved memory region (addr=%p): error code %lu",
+ inhseg->ShmemProtectiveRegion, GetLastError());
+ if (VirtualFree(inhseg->UsedShmemSegAddr, 0, MEM_RELEASE) == 0)
+ elog(FATAL, "failed to release reserved memory region (addr=%p): error code %lu",
+ inhseg->UsedShmemSegAddr, GetLastError());
+
+ hdr = (PGShmemHeader *) MapViewOfFileEx(inhseg->UsedShmemSegID, FILE_MAP_READ | FILE_MAP_WRITE, 0, 0, 0, inhseg->UsedShmemSegAddr);
+ if (!hdr)
+ elog(FATAL, "could not reattach to shared memory (key=%p, addr=%p): error code %lu",
+ inhseg->UsedShmemSegID, inhseg->UsedShmemSegAddr, GetLastError());
+ if (hdr != origUsedShmemSegAddr)
+ elog(FATAL, "reattaching to shared memory returned unexpected address (got %p, expected %p)",
+ hdr, origUsedShmemSegAddr);
+ if (hdr->magic != PGShmemMagic)
+ elog(FATAL, "reattaching to shared memory returned non-PostgreSQL memory");
+ /* Re-establish dsm_control mapping, if any */
+ if (hdr->dsm_control != 0)
+ dsm_set_control_handle(hdr->dsm_control);
+
+ inhseg->UsedShmemSegAddr = hdr; /* probably redundant */
+ }
}
/*
@@ -464,22 +480,28 @@ PGSharedMemoryReAttach(void)
* The child process startup logic might or might not call PGSharedMemoryDetach
* after this; make sure that it will be a no-op if called.
*
- * ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr are implicit
- * parameters to this routine. The caller must have already restored them to
- * the postmaster's values.
+ * Segments is an implicit parameters to this routine. The caller must have
+ * already restored ShmemProtectiveRegion and UsedShmemSegAddr
+ * in each Segment to the postmaster's values.
*/
void
PGSharedMemoryNoReAttach(void)
{
- Assert(ShmemProtectiveRegion != NULL);
- Assert(UsedShmemSegAddr != NULL);
Assert(IsUnderPostmaster);
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGInhShmemSeg *segment = &InhShmemSegs[i];
- /*
- * Under Windows we will not have mapped the segment, so we don't need to
- * un-map it. Just reset UsedShmemSegAddr to show we're not attached.
- */
- UsedShmemSegAddr = NULL;
+ Assert(segment->ShmemProtectiveRegion != NULL);
+ Assert(segment->UsedShmemSegAddr != NULL);
+
+ /*
+ * Under Windows we will not have mapped the segment, so we don't need
+ * to un-map it. Just reset UsedShmemSegAddr to show we're not
+ * attached.
+ */
+ segment->UsedShmemSegAddr = NULL;
+ }
/*
* We *must* close the inherited shmem segment handle, else Windows will
@@ -492,49 +514,55 @@ PGSharedMemoryNoReAttach(void)
/*
* PGSharedMemoryDetach
*
- * Detach from the shared memory segment, if still attached. This is not
+ * Detach from the shared memory segments, if still attached. This is not
* intended to be called explicitly by the process that originally created the
- * segment (it will have an on_shmem_exit callback registered to do that).
- * Rather, this is for subprocesses that have inherited an attachment and want
- * to get rid of it.
+ * segments (it will have an on_shmem_exit callback registered to do that).
+ * Rather, this is for subprocesses that have inherited an attachment and want to
+ * get rid of it.
*
- * ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr are implicit
- * parameters to this routine.
+ * InhShmemSegs is an implicit parameters to this routine. The caller must have
+ * already restored ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr in
+ * each Segment to the postmaster's values.
*/
void
PGSharedMemoryDetach(void)
{
- /*
- * Releasing the protective region liberates an unimportant quantity of
- * address space, but be tidy.
- */
- if (ShmemProtectiveRegion != NULL)
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- if (VirtualFree(ShmemProtectiveRegion, 0, MEM_RELEASE) == 0)
- elog(LOG, "failed to release reserved memory region (addr=%p): error code %lu",
- ShmemProtectiveRegion, GetLastError());
+ PGInhShmemSeg *segment = &InhShmemSegs[i];
- ShmemProtectiveRegion = NULL;
- }
+ /*
+ * Releasing the protective region liberates an unimportant quantity
+ * of address space, but be tidy.
+ */
+ if (segment->ShmemProtectiveRegion != NULL)
+ {
+ if (VirtualFree(segment->ShmemProtectiveRegion, 0, MEM_RELEASE) == 0)
+ elog(LOG, "failed to release reserved memory region (addr=%p): error code %lu",
+ segment->ShmemProtectiveRegion, GetLastError());
- /* Unmap the view, if it's mapped */
- if (UsedShmemSegAddr != NULL)
- {
- if (!UnmapViewOfFile(UsedShmemSegAddr))
- elog(LOG, "could not unmap view of shared memory: error code %lu",
- GetLastError());
+ segment->ShmemProtectiveRegion = NULL;
+ }
- UsedShmemSegAddr = NULL;
- }
+ /* Unmap the view, if it's mapped */
+ if (segment->UsedShmemSegAddr != NULL)
+ {
+ if (!UnmapViewOfFile(segment->UsedShmemSegAddr))
+ elog(LOG, "could not unmap view of shared memory: error code %lu",
+ GetLastError());
- /* And close the shmem handle, if we have one */
- if (UsedShmemSegID != INVALID_HANDLE_VALUE)
- {
- if (!CloseHandle(UsedShmemSegID))
- elog(LOG, "could not close handle to shared memory: error code %lu",
- GetLastError());
+ segment->UsedShmemSegAddr = NULL;
+ }
- UsedShmemSegID = INVALID_HANDLE_VALUE;
+ /* And close the shmem handle, if we have one */
+ if (segment->UsedShmemSegID != INVALID_HANDLE_VALUE)
+ {
+ if (!CloseHandle(segment->UsedShmemSegID))
+ elog(LOG, "could not close handle to shared memory: error code %lu",
+ GetLastError());
+
+ segment->UsedShmemSegID = INVALID_HANDLE_VALUE;
+ }
}
}
@@ -574,50 +602,55 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
{
void *address;
- Assert(ShmemProtectiveRegion != NULL);
- Assert(UsedShmemSegAddr != NULL);
- Assert(UsedShmemSegSize != 0);
-
- /* ShmemProtectiveRegion */
- address = VirtualAllocEx(hChild, ShmemProtectiveRegion,
- PROTECTIVE_REGION_SIZE,
- MEM_RESERVE, PAGE_NOACCESS);
- if (address == NULL)
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- /* Don't use FATAL since we're running in the postmaster */
- elog(LOG, "could not reserve shared memory region (addr=%p) for child %p: error code %lu",
- ShmemProtectiveRegion, hChild, GetLastError());
- return false;
- }
- if (address != ShmemProtectiveRegion)
- {
- /*
- * Should never happen - in theory if allocation granularity causes
- * strange effects it could, so check just in case.
- *
- * Don't use FATAL since we're running in the postmaster.
- */
- elog(LOG, "reserved shared memory region got incorrect address %p, expected %p",
- address, ShmemProtectiveRegion);
- return false;
- }
+ PGInhShmemSeg *segment = &InhShmemSegs[i];
- /* UsedShmemSegAddr */
- address = VirtualAllocEx(hChild, UsedShmemSegAddr, UsedShmemSegSize,
- MEM_RESERVE, PAGE_READWRITE);
- if (address == NULL)
- {
- elog(LOG, "could not reserve shared memory region (addr=%p) for child %p: error code %lu",
- UsedShmemSegAddr, hChild, GetLastError());
- return false;
- }
- if (address != UsedShmemSegAddr)
- {
- elog(LOG, "reserved shared memory region got incorrect address %p, expected %p",
- address, UsedShmemSegAddr);
- return false;
- }
+ Assert(segment->ShmemProtectiveRegion != NULL);
+ Assert(segment->UsedShmemSegAddr != NULL);
+ Assert(UsedShmemSegSizes[i] != 0);
+
+ /* ShmemProtectiveRegion */
+ address = VirtualAllocEx(hChild, segment->ShmemProtectiveRegion,
+ PROTECTIVE_REGION_SIZE,
+ MEM_RESERVE, PAGE_NOACCESS);
+ if (address == NULL)
+ {
+ /* Don't use FATAL since we're running in the postmaster */
+ elog(LOG, "could not reserve shared memory region (addr=%p) for child %p: error code %lu",
+ segment->ShmemProtectiveRegion, hChild, GetLastError());
+ return false;
+ }
+ if (address != segment->ShmemProtectiveRegion)
+ {
+ /*
+ * Should never happen - in theory if allocation granularity
+ * causes strange effects it could, so check just in case.
+ *
+ * Don't use FATAL since we're running in the postmaster.
+ */
+ elog(LOG, "reserved shared memory region got incorrect address %p, expected %p",
+ address, segment->ShmemProtectiveRegion);
+ return false;
+ }
+
+ /* UsedShmemSegAddr */
+ address = VirtualAllocEx(hChild, segment->UsedShmemSegAddr, UsedShmemSegSizes[i],
+ MEM_RESERVE, PAGE_READWRITE);
+ if (address == NULL)
+ {
+ elog(LOG, "could not reserve shared memory region (addr=%p) for child %p: error code %lu",
+ segment->UsedShmemSegAddr, hChild, GetLastError());
+ return false;
+ }
+ if (address != segment->UsedShmemSegAddr)
+ {
+ elog(LOG, "reserved shared memory region got incorrect address %p, expected %p",
+ address, segment->UsedShmemSegAddr);
+ return false;
+ }
+ }
return true;
}
@@ -627,7 +660,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index cea229ad6a4..85da8ac381a 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -89,14 +89,7 @@ typedef int InheritableSocket;
typedef struct
{
char DataDir[MAXPGPATH];
-#ifndef WIN32
- unsigned long UsedShmemSegID;
-#else
- void *ShmemProtectiveRegion;
- HANDLE UsedShmemSegID;
-#endif
- void *UsedShmemSegAddr;
- slock_t *ShmemLock;
+ PGInhShmemSeg InhShmemSegs[NUM_MEMORY_MAPPINGS];
#ifdef USE_INJECTION_POINTS
struct InjectionPointsCtl *ActiveInjectionPoints;
#endif
@@ -675,8 +668,13 @@ SubPostmasterMain(int argc, char *argv[])
process_shared_preload_libraries();
/* Restore basic shared memory pointers */
- if (UsedShmemSegAddr != NULL)
- InitShmemAccess(UsedShmemSegAddr);
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGInhShmemSeg *inhseg = &InhShmemSegs[i];
+
+ if (inhseg->UsedShmemSegAddr != NULL)
+ InitShmemAccess(i, inhseg->UsedShmemSegAddr, inhseg->ShmemLock);
+ }
/*
* Run the appropriate Main function
@@ -717,14 +715,7 @@ save_backend_variables(BackendParameters *param,
strlcpy(param->DataDir, DataDir, MAXPGPATH);
param->MyPMChildSlot = child_slot;
-
-#ifdef WIN32
- param->ShmemProtectiveRegion = ShmemProtectiveRegion;
-#endif
- param->UsedShmemSegID = UsedShmemSegID;
- param->UsedShmemSegAddr = UsedShmemSegAddr;
-
- param->ShmemLock = ShmemLock;
+ memcpy(param->InhShmemSegs, InhShmemSegs, sizeof(InhShmemSegs));
#ifdef USE_INJECTION_POINTS
param->ActiveInjectionPoints = ActiveInjectionPoints;
@@ -979,14 +970,7 @@ restore_backend_variables(BackendParameters *param)
SetDataDir(param->DataDir);
MyPMChildSlot = param->MyPMChildSlot;
-
-#ifdef WIN32
- ShmemProtectiveRegion = param->ShmemProtectiveRegion;
-#endif
- UsedShmemSegID = param->UsedShmemSegID;
- UsedShmemSegAddr = param->UsedShmemSegAddr;
-
- ShmemLock = param->ShmemLock;
+ memcpy(InhShmemSegs, param->InhShmemSegs, sizeof(InhShmemSegs));
#ifdef USE_INJECTION_POINTS
ActiveInjectionPoints = param->ActiveInjectionPoints;
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index c0c223b2e32..864c8268cae 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proclist.h"
BufferDescPadded *BufferDescriptors;
@@ -63,7 +64,10 @@ CkptSortItem *CkptBufferIds;
* Initialize shared buffer pool
*
* This is called once during shared-memory initialization (either in the
- * postmaster, or in a standalone backend).
+ * postmaster, or in a standalone backend). Size of data structures initialized
+ * here depends on NBuffers, and to be able to change NBuffers without a
+ * restart we store each structure into a separate shared memory segment, which
+ * could be resized on demand.
*/
void
BufferManagerShmemInit(void)
@@ -75,22 +79,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
- NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ ShmemInitStructInSegment("Buffer Descriptors",
+ NBuffers * sizeof(BufferDescPadded),
+ &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
- NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ ShmemInitStructInSegment("Buffer Blocks",
+ NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
- NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -100,8 +104,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ CHECKPOINT_BUFFERS_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -146,33 +151,42 @@ BufferManagerShmemInit(void)
* BufferManagerShmemSize
*
* compute the size of shared memory for the buffer pool including
- * data pages, buffer descriptors, hash tables, etc.
+ * data pages, buffer descriptors, hash tables, etc. based on the
+ * shared memory segment. The main segment must not allocate anything
+ * related to buffers, every other segment will receive part of the
+ * data.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
{
- Size size = 0;
+ size_t size;
- /* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
- /* to allow aligning buffer descriptors */
+ /* size of buffer descriptors, plus alignment padding */
+ size = add_size(0, mul_size(NBuffers, sizeof(BufferDescPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
+ mapping_sizes[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[BUFFER_DESCRIPTORS_SHMEM_SEGMENT].shmem_reserved = size;
/* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
+ size = add_size(0, PG_IO_ALIGN_SIZE);
size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_reserved = size;
/* size of stuff controlled by freelist.c */
- size = add_size(size, StrategyShmemSize());
+ mapping_sizes[STRATEGY_SHMEM_SEGMENT].shmem_req_size = StrategyShmemSize();
+ mapping_sizes[STRATEGY_SHMEM_SEGMENT].shmem_reserved = StrategyShmemSize();
- /* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
- sizeof(ConditionVariableMinimallyPadded)));
- /* to allow aligning the above */
+ /* size of I/O condition variables, plus alignment padding */
+ size = add_size(0, mul_size(NBuffers,
+ sizeof(ConditionVariableMinimallyPadded)));
size = add_size(size, PG_CACHE_LINE_SIZE);
+ mapping_sizes[BUFFER_IOCV_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[BUFFER_IOCV_SHMEM_SEGMENT].shmem_reserved = size;
/* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_req_size = mul_size(NBuffers, sizeof(CkptSortItem));
+ mapping_sizes[CHECKPOINT_BUFFERS_SHMEM_SEGMENT].shmem_reserved = mul_size(NBuffers, sizeof(CkptSortItem));
return size;
}
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 5089c7322f3..a33786a460b 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -25,6 +25,7 @@
#include "funcapi.h"
#include "storage/buf_internals.h"
#include "storage/lwlock.h"
+#include "storage/pg_shmem.h"
#include "utils/rel.h"
#include "utils/builtins.h"
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index b7687836188..13e701ee4a1 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -418,9 +419,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
- sizeof(BufferStrategyControl),
- &found);
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, STRATEGY_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index cb944edd8df..4af7782795e 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -62,6 +62,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -69,7 +71,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 2a3dfedf7e9..12457e1bbf3 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -50,6 +50,7 @@
#include "storage/procarray.h"
#include "storage/procsignal.h"
#include "storage/sinvaladt.h"
+#include "utils/builtins.h"
#include "utils/guc.h"
#include "utils/injection_point.h"
@@ -81,10 +82,17 @@ RequestAddinShmemSpace(Size size)
/*
* CalculateShmemSize
- * Calculates the amount of shared memory needed.
+ * Calculates the amount of shared memory needed.
+ *
+ * The amount of shared memory required per segment is saved in mapping_sizes,
+ * which is expected to be an array of size NUM_MEMORY_MAPPINGS. The total
+ * amount of memory needed across all the segments is returned. For the memory
+ * mappings which reserve address space for future expansion, the required
+ * amount of reserved space is saved in mapping_sizes of those segments.
+ * This memory is not included in the returned value.
*/
Size
-CalculateShmemSize(void)
+CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
{
Size size;
@@ -102,7 +110,13 @@ CalculateShmemSize(void)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+
+ /*
+ * Buffer manager adds estimates for memory requirements for every shared
+ * memory segment that it uses in the corresponding AnonymousMappings.
+ * Consider size required from only the main shared memory segment here.
+ */
+ size = add_size(size, BufferManagerShmemSize(mapping_sizes));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
@@ -145,8 +159,22 @@ CalculateShmemSize(void)
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
+ /*
+ * All the shared memory allocations considered so far happen in the main
+ * shared memory segment.
+ */
+ mapping_sizes[MAIN_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[MAIN_SHMEM_SEGMENT].shmem_reserved = size;
+
+ size = 0;
/* might as well round it off to a multiple of a typical page size */
- size = add_size(size, 8192 - (size % 8192));
+ for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ mapping_sizes[segment].shmem_req_size = add_size(mapping_sizes[segment].shmem_req_size, 8192 - (mapping_sizes[segment].shmem_req_size % 8192));
+ mapping_sizes[segment].shmem_reserved = add_size(mapping_sizes[segment].shmem_reserved, 8192 - (mapping_sizes[segment].shmem_reserved % 8192));
+ /* Compute the total size of all segments */
+ size = size + mapping_sizes[segment].shmem_req_size;
+ }
return size;
}
@@ -185,25 +213,21 @@ AttachSharedMemoryStructs(void)
/*
* CreateSharedMemoryAndSemaphores
- * Creates and initializes shared memory and semaphores.
+ * Creates shared memory segments and initializes shared memory structures
+ * and semaphores.
*/
void
CreateSharedMemoryAndSemaphores(void)
{
- PGShmemHeader *shim;
- PGShmemHeader *seghdr;
- Size size;
+ PGShmemHeader *main_seg_shim = NULL;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize();
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+ CalculateShmemSize(mapping_sizes);
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
+ /* Decide if we use huge pages or regular size pages */
+ PrepareHugePages();
/*
* Make sure that huge pages are never reported as "unknown" while the
@@ -212,18 +236,46 @@ CreateSharedMemoryAndSemaphores(void)
Assert(strcmp("unknown",
GetConfigOption("huge_pages_status", false, false)) != 0);
- InitShmemAccess(seghdr);
-
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocation();
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ MemoryMappingSizes *mapping = &mapping_sizes[i];
+ PGInhShmemSeg *inhseg = &InhShmemSegs[i];
+ PGShmemHeader *shim;
+ PGShmemHeader *seghdr;
+
+ /*
+ * Set seed shmem identifier which will be changed to the final one
+ * when creating the shared memory segment.
+ */
+ inhseg->UsedShmemSegID = i;
+
+ /* Compute the size of the shared-memory block */
+ elog(DEBUG3, "invoking IpcMemoryCreate(segment %s, size=%zu, reserved address space=%zu)",
+ MappingName(i), mapping->shmem_req_size, mapping->shmem_reserved);
+
+ /*
+ * Create the shmem segment.
+ *
+ * XXX: Do multiple shims are needed, one per segment?
+ */
+ seghdr = PGSharedMemoryCreate(i, mapping, &shim);
+
+ InitShmemAccess(i, seghdr, NULL);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ inhseg->ShmemLock = InitShmemAllocation(i);
+
+ if (i == MAIN_SHMEM_SEGMENT)
+ main_seg_shim = shim;
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
/* Initialize dynamic shared memory facilities. */
- dsm_postmaster_startup(shim);
+ dsm_postmaster_startup(main_seg_shim);
/*
* Now give loadable modules a chance to set up their shmem allocations
@@ -336,7 +388,9 @@ CreateOrAttachShmemStructs(void)
* InitializeShmemGUCs
*
* This function initializes runtime-computed GUCs related to the amount of
- * shared memory required for the current configuration.
+ * shared memory required for the current configuration. It assumes that the
+ * memory required by the shared memory segments is already calculated and is
+ * available in AnonymousMappings.
*/
void
InitializeShmemGUCs(void)
@@ -345,11 +399,13 @@ InitializeShmemGUCs(void)
Size size_b;
Size size_mb;
Size hp_size;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize();
+ size_b = CalculateShmemSize(mapping_sizes);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
@@ -358,7 +414,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 1b536363152..2e365261c09 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -63,6 +63,18 @@
* unnecessary.
*/
+ /*
+ * TODO: Most of the functions here can take PGShmemSegment as argument
+ * instead of segment_id except for ShmemAlloc, ShmemInitStructInSegment, and
+ * ShmemAddrIsValid. The first one is used in lwlock.c. We need to check
+ * whether we can use ShmemAllocInternal() there and expose Segements.
+ * Exposing Segments where the third one is used seems even harder. The
+ * second one can not replace ShmemInitStruct since the latter is used in
+ * many places. Further we need ShmemInitStructInSegment to accept segment_id
+ * so that we can avoid exposing Segments in all the places where the
+ * function is used.
+ */
+
#include "postgres.h"
#include "common/int.h"
@@ -76,21 +88,27 @@
#include "storage/spin.h"
#include "utils/builtins.h"
-static void *ShmemAllocRaw(Size size, Size *allocated_size);
-static void *ShmemAllocUnlocked(Size size);
-
-/* shared memory global variables */
-
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
-
-static void *ShmemBase; /* start address of shared memory */
+/* Structure managing one shared memory segment. */
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+ void *ShmemBase; /* start address of shared memory */
+ const char *ShmemSegmentName; /* name of the segment for logging */
+} ShmemSegment;
-static void *ShmemEnd; /* end+1 address of shared memory */
+ShmemSegment Segments[NUM_MEMORY_MAPPINGS];
-slock_t *ShmemLock; /* spinlock for shared memory and LWLock
- * allocation */
+static void *ShmemAllocRaw(ShmemSegment *segment, Size size, Size *allocated_size);
+static void *ShmemAllocUnlocked(ShmemSegment *segment, Size size);
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -98,36 +116,69 @@ static bool firstNumaTouch = true;
Datum pg_numa_available(PG_FUNCTION_ARGS);
/*
- * InitShmemAccess() --- set up basic pointers to shared memory.
+ * InitShmemAccess() --- set up basic pointers in the given shared memory segment.
+ *
+ * These addresses are expected to be stable throughout the life of the process
+ * even if the underlying segments get resized.
*/
void
-InitShmemAccess(PGShmemHeader *seghdr)
+InitShmemAccess(int segment_id, PGShmemHeader *seghdr, slock_t *ShmemLock)
{
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ ShmemSegment *segment;
+
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+
+ /*
+ * When called from Postmaster code after creating shared memory segment
+ * ShmemLock is expected to be NULL; it will be created later. But a
+ * backend initialized under EXEC_BACKEND inherits already initialized
+ * lock.
+ */
+ Assert((!IsUnderPostmaster && !ShmemLock) || (IsUnderPostmaster && ShmemLock));
+
+ segment = &Segments[segment_id];
+
+ segment->ShmemSegHdr = seghdr;
+ segment->ShmemBase = (void *) seghdr;
+ segment->ShmemLock = ShmemLock;
+ segment->ShmemSegmentName = MappingName(segment_id);
+
}
/*
* InitShmemAllocation() --- set up shared-memory space allocation.
*
* This should be called only in the postmaster or a standalone backend.
+ *
+ * The function initializes the ShmemLock spinlock in the given segment, and
+ * returns it.
*/
-void
-InitShmemAllocation(void)
+slock_t *
+InitShmemAllocation(int segment_id)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ ShmemSegment *segment;
+ PGShmemHeader *shmhdr;
char *aligned;
+ Assert(!IsUnderPostmaster);
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+
+ segment = &Segments[segment_id];
+ shmhdr = segment->ShmemSegHdr;
+
+ /* This function should be called only once for every segment. */
Assert(shmhdr != NULL);
+ Assert(!segment->ShmemLock);
/*
* Initialize the spinlock used by ShmemAlloc. We must use
* ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet.
+ * Pass it back to the caller through inhseg, so that it can be shared
+ * with backends.
*/
- ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t));
+ segment->ShmemLock = (slock_t *) ShmemAllocUnlocked(segment, sizeof(slock_t));
- SpinLockInit(ShmemLock);
+ SpinLockInit(segment->ShmemLock);
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -140,42 +191,60 @@ InitShmemAllocation(void)
/* ShmemIndex can't be set up yet (need LWLocks first) */
shmhdr->index = NULL;
- ShmemIndex = (HTAB *) NULL;
+ Assert(!ShmemIndex);
+
+ return segment->ShmemLock;
}
/*
- * ShmemAlloc -- allocate max-aligned chunk from shared memory
+ * ShmemAlloc --
+ * allocate max-aligned chunk from given shared memory segment
*
* Throws error if request cannot be satisfied.
*
- * Assumes ShmemLock and ShmemSegHdr are initialized.
+ * Assumes ShmemLock and ShmemSegHdr in the given segment are initialized.
*/
-void *
-ShmemAlloc(Size size)
+
+static void *
+ShmemAllocInternal(ShmemSegment *segment, Size size)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRaw(segment, size, &allocated_size);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ segment->ShmemSegmentName, size)));
return newSpace;
}
+void *
+ShmemAlloc(int segment_id, Size size)
+{
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+
+ return ShmemAllocInternal(&Segments[segment_id], size);
+}
+
/*
* ShmemAllocNoError -- allocate max-aligned chunk from shared memory
*
* As ShmemAlloc, but returns NULL if out of space, rather than erroring.
+ *
+ * This is used as a memory allocation callback for hash tables created using
+ * dynahash.c APIs. It's a bit of work to make the callback specify the segment
+ * where to allocate the memory. For now, there is not need to create shared
+ * memory hash tables in shared memory segments other than main memory segment.
+ * Hence we do not support segment_id parameter here.
*/
void *
ShmemAllocNoError(Size size)
{
Size allocated_size;
- return ShmemAllocRaw(size, &allocated_size);
+ return ShmemAllocRaw(&Segments[MAIN_SHMEM_SEGMENT], size, &allocated_size);
}
/*
@@ -185,11 +254,12 @@ ShmemAllocNoError(Size size)
* be equal to the number requested plus any padding we choose to add.
*/
static void *
-ShmemAllocRaw(Size size, Size *allocated_size)
+ShmemAllocRaw(ShmemSegment *segment, Size size, Size *allocated_size)
{
Size newStart;
Size newFree;
void *newSpace;
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
/*
* Ensure all space is adequately aligned. We used to only MAXALIGN this
@@ -205,22 +275,21 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
-
- SpinLockAcquire(ShmemLock);
+ Assert(shmhdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ SpinLockAcquire(segment->ShmemLock);
+ newStart = shmhdr->freeoffset;
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= shmhdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
- ShmemSegHdr->freeoffset = newFree;
+ newSpace = (char *) segment->ShmemBase + newStart;
+ shmhdr->freeoffset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(segment->ShmemLock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -229,7 +298,8 @@ ShmemAllocRaw(Size size, Size *allocated_size)
}
/*
- * ShmemAllocUnlocked -- allocate max-aligned chunk from shared memory
+ * ShmemAllocUnlocked
+ * allocate max-aligned chunk from given shared memory segment
*
* Allocate space without locking ShmemLock. This should be used for,
* and only for, allocations that must happen before ShmemLock is ready.
@@ -237,30 +307,31 @@ ShmemAllocRaw(Size size, Size *allocated_size)
* We consider maxalign, rather than cachealign, sufficient here.
*/
static void *
-ShmemAllocUnlocked(Size size)
+ShmemAllocUnlocked(ShmemSegment *segment, Size size)
{
Size newStart;
Size newFree;
void *newSpace;
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
/*
* Ensure allocated space is adequately aligned.
*/
size = MAXALIGN(size);
- Assert(ShmemSegHdr != NULL);
+ Assert(shmhdr != NULL);
- newStart = ShmemSegHdr->freeoffset;
+ newStart = shmhdr->freeoffset;
newFree = newStart + size;
- if (newFree > ShmemSegHdr->totalsize)
+ if (newFree > shmhdr->totalsize)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
- ShmemSegHdr->freeoffset = newFree;
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ segment->ShmemSegmentName, size)));
+ shmhdr->freeoffset = newFree;
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) segment->ShmemBase + newStart;
Assert(newSpace == (void *) MAXALIGN(newSpace));
@@ -268,14 +339,23 @@ ShmemAllocUnlocked(Size size)
}
/*
- * ShmemAddrIsValid -- test if an address refers to shared memory
+ * ShmemAddrIsValid
+ * test if an address refers to the given shared memory segment.
*
* Returns true if the pointer points within the shared memory segment.
*/
bool
-ShmemAddrIsValid(const void *addr)
+ShmemAddrIsValid(int segment_id, const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ ShmemSegment *segment;
+ void *shmemEnd;
+
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+
+ segment = &Segments[segment_id];
+ shmemEnd = (char *) segment->ShmemBase + segment->ShmemSegHdr->totalsize;
+
+ return (addr >= segment->ShmemBase) && (addr < shmemEnd);
}
/*
@@ -329,6 +409,9 @@ InitShmemIndex(void)
* Note: before Postgres 9.0, this function returned NULL for some failure
* cases. Now, it always throws error instead, so callers need not check
* for NULL.
+ *
+ * See prologue of ShmemAllocNoError for explanation about lack of segment_id
+ * parameter.
*/
HTAB *
ShmemInitHash(const char *name, /* table string name for shmem index */
@@ -352,9 +435,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
- hash_get_shared_size(infoP, hash_flags),
- &found);
+ location = ShmemInitStructInSegment(name,
+ hash_get_shared_size(infoP, hash_flags),
+ &found, MAIN_SHMEM_SEGMENT);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -387,24 +470,39 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, int segment_id)
{
ShmemIndexEnt *result;
void *structPtr;
+ ShmemSegment *segment;
+
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+
+ segment = &Segments[segment_id];
LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
if (!ShmemIndex)
{
- PGShmemHeader *shmemseghdr = ShmemSegHdr;
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
- /* Must be trying to create/attach to ShmemIndex itself */
+ /*
+ * Must be trying to create/attach to ShmemIndex itself in the main
+ * shared memory segment.
+ */
+ Assert(segment_id == MAIN_SHMEM_SEGMENT);
Assert(strcmp(name, "ShmemIndex") == 0);
if (IsUnderPostmaster)
{
/* Must be initializing a (non-standalone) backend */
- Assert(shmemseghdr->index != NULL);
- structPtr = shmemseghdr->index;
+ Assert(shmhdr->index != NULL);
+ structPtr = shmhdr->index;
*foundPtr = true;
}
else
@@ -417,9 +515,9 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* index has been initialized. This should be OK because no other
* process can be accessing shared memory yet.
*/
- Assert(shmemseghdr->index == NULL);
- structPtr = ShmemAlloc(size);
- shmemseghdr->index = structPtr;
+ Assert(shmhdr->index == NULL);
+ structPtr = ShmemAllocInternal(segment, size);
+ shmhdr->index = structPtr;
*foundPtr = false;
}
LWLockRelease(ShmemIndexLock);
@@ -435,8 +533,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, segment_id)));
}
if (*foundPtr)
@@ -461,7 +559,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRaw(segment, size, &allocated_size);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -476,18 +574,18 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
result->size = size;
result->allocated_size = allocated_size;
result->location = structPtr;
+ result->segment_id = segment_id;
}
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValid(segment_id, structPtr));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -522,13 +620,14 @@ mul_size(Size s1, Size s2)
Datum
pg_get_shmem_allocations(PG_FUNCTION_ARGS)
{
-#define PG_GET_SHMEM_SIZES_COLS 4
+#define PG_GET_SHMEM_SIZES_COLS 5
ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
HASH_SEQ_STATUS hstat;
ShmemIndexEnt *ent;
- Size named_allocated = 0;
+ Size named_allocated[NUM_MEMORY_MAPPINGS] = {0};
Datum values[PG_GET_SHMEM_SIZES_COLS];
bool nulls[PG_GET_SHMEM_SIZES_COLS];
+ int i;
InitMaterializedSRF(fcinfo, 0);
@@ -540,30 +639,49 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
memset(nulls, 0, sizeof(nulls));
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
+ ShmemSegment *segment = &Segments[ent->segment_id];
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
- values[2] = Int64GetDatum(ent->size);
- values[3] = Int64GetDatum(ent->allocated_size);
- named_allocated += ent->allocated_size;
+ values[1] = CStringGetTextDatum(segment->ShmemSegmentName);
+ values[2] = Int64GetDatum((char *) ent->location - (char *) shmhdr);
+ values[3] = Int64GetDatum(ent->size);
+ values[4] = Int64GetDatum(ent->allocated_size);
+ named_allocated[ent->segment_id] += ent->allocated_size;
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
values, nulls);
}
/* output shared memory allocated but not counted via the shmem index */
- values[0] = CStringGetTextDatum("<anonymous>");
- nulls[1] = true;
- values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+
+ values[0] = CStringGetTextDatum("<anonymous>");
+ values[1] = CStringGetTextDatum(segment->ShmemSegmentName);
+ nulls[2] = true;
+ values[3] = Int64GetDatum(shmhdr->freeoffset - named_allocated[i]);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
/* output as-of-yet unused shared memory */
- nulls[0] = true;
- values[1] = Int64GetDatum(ShmemSegHdr->freeoffset);
- nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ memset(nulls, 0, sizeof(nulls));
+
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+
+ nulls[0] = true;
+ values[1] = CStringGetTextDatum(segment->ShmemSegmentName);
+ values[2] = Int64GetDatum(shmhdr->freeoffset);
+ values[3] = Int64GetDatum(shmhdr->totalsize - shmhdr->freeoffset);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
LWLockRelease(ShmemIndexLock);
@@ -588,7 +706,7 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size os_page_size;
void **page_ptrs;
int *pages_status;
- uint64 shm_total_page_count,
+ uint64 shm_total_page_count = 0,
shm_ent_page_count,
max_nodes;
Size *nodes;
@@ -623,7 +741,13 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ PGShmemHeader *shmhdr = Segments[segment].ShmemSegHdr;
+
+ shm_total_page_count += (shmhdr->totalsize / os_page_size) + 1;
+ }
+
page_ptrs = palloc0_array(void *, shm_total_page_count);
pages_status = palloc_array(int, shm_total_page_count);
@@ -764,7 +888,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
@@ -774,3 +898,44 @@ pg_numa_available(PG_FUNCTION_ARGS)
{
PG_RETURN_BOOL(pg_numa_init() != -1);
}
+
+/* SQL SRF showing shared memory segments */
+Datum
+pg_get_shmem_segments(PG_FUNCTION_ARGS)
+{
+#define PG_GET_SHMEM_SEGS_COLS 5
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ Datum values[PG_GET_SHMEM_SEGS_COLS];
+ bool nulls[PG_GET_SHMEM_SEGS_COLS];
+ int i;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* output all allocated entries */
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+ int j;
+
+ if (shmhdr == NULL)
+ {
+ for (j = 0; j < PG_GET_SHMEM_SEGS_COLS; j++)
+ nulls[j] = true;
+ }
+ else
+ {
+ memset(nulls, 0, sizeof(nulls));
+ values[0] = Int32GetDatum(i);
+ values[1] = CStringGetTextDatum(segment->ShmemSegmentName);
+ values[2] = Int64GetDatum(shmhdr->totalsize);
+ values[3] = Int64GetDatum(shmhdr->freeoffset);
+ values[4] = Int64GetDatum(shmhdr->ReservedSize);
+ }
+
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
+
+ return (Datum) 0;
+}
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 517c55375b4..0f115dc89f3 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -446,7 +448,7 @@ CreateLWLocks(void)
char *ptr;
/* Allocate space */
- ptr = (char *) ShmemAlloc(spaceLocks);
+ ptr = (char *) ShmemAlloc(MAIN_SHMEM_SEGMENT, spaceLocks);
/* Initialize the dynamic-allocation counter for tranches */
LWLockCounter = (int *) ptr;
@@ -612,12 +614,15 @@ LWLockNewTrancheId(const char *name)
/*
* We use the ShmemLock spinlock to protect LWLockCounter and
* LWLockTrancheNames.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
*/
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(InhShmemSegs[MAIN_SHMEM_SEGMENT].ShmemLock);
if (*LWLockCounter - LWTRANCHE_FIRST_USER_DEFINED >= MAX_NAMED_TRANCHES)
{
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(InhShmemSegs[MAIN_SHMEM_SEGMENT].ShmemLock);
ereport(ERROR,
(errmsg("maximum number of tranches already registered"),
errdetail("No more than %d tranches may be registered.",
@@ -628,7 +633,7 @@ LWLockNewTrancheId(const char *name)
LocalLWLockCounter = *LWLockCounter;
strlcpy(LWLockTrancheNames[result - LWTRANCHE_FIRST_USER_DEFINED], name, NAMEDATALEN);
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(InhShmemSegs[MAIN_SHMEM_SEGMENT].ShmemLock);
return result;
}
@@ -750,9 +755,9 @@ GetLWTrancheName(uint16 trancheId)
*/
if (trancheId >= LocalLWLockCounter)
{
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(InhShmemSegs[MAIN_SHMEM_SEGMENT].ShmemLock);
LocalLWLockCounter = *LWLockCounter;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(InhShmemSegs[MAIN_SHMEM_SEGMENT].ShmemLock);
if (trancheId >= LocalLWLockCounter)
elog(ERROR, "tranche %d is not registered", trancheId);
diff --git a/src/backend/storage/lmgr/predicate.c b/src/backend/storage/lmgr/predicate.c
index fe75ead3501..9aab75f54d6 100644
--- a/src/backend/storage/lmgr/predicate.c
+++ b/src/backend/storage/lmgr/predicate.c
@@ -207,6 +207,7 @@
#include "miscadmin.h"
#include "pgstat.h"
#include "port/pg_lfind.h"
+#include "storage/pg_shmem.h"
#include "storage/predicate.h"
#include "storage/predicate_internals.h"
#include "storage/proc.h"
@@ -595,7 +596,7 @@ CreatePredXact(void)
static void
ReleasePredXact(SERIALIZABLEXACT *sxact)
{
- Assert(ShmemAddrIsValid(sxact));
+ Assert(ShmemAddrIsValid(MAIN_SHMEM_SEGMENT, sxact));
dlist_delete(&sxact->xactLink);
dlist_push_tail(&PredXact->availableList, &sxact->xactLink);
diff --git a/src/backend/utils/activity/pgstat_shmem.c b/src/backend/utils/activity/pgstat_shmem.c
index 33fbdca9609..c6d9157f417 100644
--- a/src/backend/utils/activity/pgstat_shmem.c
+++ b/src/backend/utils/activity/pgstat_shmem.c
@@ -13,6 +13,7 @@
#include "postgres.h"
#include "pgstat.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "utils/memutils.h"
#include "utils/pgstat_internal.h"
@@ -233,7 +234,7 @@ StatsShmemInit(void)
int idx = kind - PGSTAT_KIND_CUSTOM_MIN;
Assert(kind_info->shared_size != 0);
- ctl->custom_data[idx] = ShmemAlloc(kind_info->shared_size);
+ ctl->custom_data[idx] = ShmemAlloc(MAIN_SHMEM_SEGMENT, kind_info->shared_size);
ptr = ctl->custom_data[idx];
}
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 5e5e33f64fc..52de299c2d8 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8592,8 +8592,8 @@
{ oid => '5052', descr => 'allocations from the main shared memory segment',
proname => 'pg_get_shmem_allocations', prorows => '50', proretset => 't',
provolatile => 'v', prorettype => 'record', proargtypes => '',
- proallargtypes => '{text,int8,int8,int8}', proargmodes => '{o,o,o,o}',
- proargnames => '{name,off,size,allocated_size}',
+ proallargtypes => '{text,text,int8,int8,int8}', proargmodes => '{o,o,o,o,o}',
+ proargnames => '{name,segment,off,size,allocated_size}',
prosrc => 'pg_get_shmem_allocations' },
{ oid => '4099', descr => 'Is NUMA support available?',
@@ -8616,6 +8616,14 @@
proargmodes => '{o,o,o}', proargnames => '{name,type,size}',
prosrc => 'pg_get_dsm_registry_allocations' },
+# shared memory segments
+{ oid => '5101', descr => 'shared memory segments',
+ proname => 'pg_get_shmem_segments', prorows => '6', proretset => 't',
+ provolatile => 'v', prorettype => 'record', proargtypes => '',
+ proallargtypes => '{int4,text,int8,int8,int8}', proargmodes => '{o,o,o,o,o}',
+ proargnames => '{id,name,size,freeoffset,reserved_size}',
+ prosrc => 'pg_get_shmem_segments' },
+
# memory context of local backend
{ oid => '2282',
descr => 'information about all memory contexts of local backend',
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index a40adf6b2a8..93348a34378 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -19,6 +19,7 @@
#include "storage/block.h"
#include "storage/buf.h"
#include "storage/bufpage.h"
+#include "storage/pg_shmem.h"
#include "storage/relfilelocator.h"
#include "utils/relcache.h"
#include "utils/snapmgr.h"
@@ -367,7 +368,7 @@ extern void MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index da32787ab51..f1d0802d048 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -18,6 +18,8 @@
#ifndef IPC_H
#define IPC_H
+#include "storage/pg_shmem.h"
+
typedef void (*pg_on_exit_callback) (int code, Datum arg);
typedef void (*shmem_startup_hook_type) (void);
@@ -77,7 +79,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(void);
+extern Size CalculateShmemSize(MemoryMappingSizes *mapping_sizes);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 3aeada554b2..8bd78a89d38 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -25,14 +25,24 @@
#define PG_SHMEM_H
#include "storage/dsm_impl.h"
+#include "storage/spin.h"
+
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
int32 magic; /* magic # to identify Postgres segments */
#define PGShmemMagic 679834894
pid_t creatorPID; /* PID of creating process (set but unread) */
- Size totalsize; /* total size of segment */
+
Size freeoffset; /* offset to first free space */
+
+ /*
+ * TODO: We might have to rename these fields to allocSize (for amount of
+ * memory allocated currently in this segment), maxSize (for maximum size
+ * the segment can grow to.)
+ */
+ Size totalsize; /* total size of segment */
+ Size ReservedSize; /* Size of the reserved mapping */
dsm_handle dsm_control; /* ID of dynamic shared memory control seg */
void *index; /* pointer to ShmemIndex table */
#ifndef WIN32 /* Windows doesn't have useful inode#s */
@@ -41,6 +51,69 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+/*
+ * Information about the shared memory segment that is required to be passed
+ * from the Postmaster to each backend.
+ */
+typedef struct PGInhShmemSeg
+{
+ slock_t *ShmemLock; /* spinlock for shared memory and LWLock
+ * allocation */
+ void *UsedShmemSegAddr; /* SysV shared memory for the header */
+#ifndef WIN32
+ unsigned long UsedShmemSegID; /* IPC key */
+#else
+ void *ShmemProtectiveRegion; /* Protective region for Windows
+ * shared memory */
+ HANDLE UsedShmemSegID;
+#endif
+} PGInhShmemSeg;
+
+/*
+ * To be able to dynamically resize largest parts of the data stored in shared
+ * memory, we split it into multiple shared memory mappings segments. Each
+ * segment contains only certain part of the data, whose size depends on
+ * the size of buffer pool.
+ *
+ * TODO: convert this to enum?
+ */
+
+/* The main segment, contains everything except buffer blocks and related data. */
+#define MAIN_SHMEM_SEGMENT 0
+
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Buffer descriptors */
+#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2
+
+/* Condition variables for buffers */
+#define BUFFER_IOCV_SHMEM_SEGMENT 3
+
+/* Checkpoint BufferIds */
+#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4
+
+/* Buffer strategy status */
+#define STRATEGY_SHMEM_SEGMENT 5
+
+/* Number of available segments for anonymous memory mappings */
+#define NUM_MEMORY_MAPPINGS 6
+
+/*
+ * Structure to hold required sizes of each shared memory segment as calculated
+ * by CalculateShmemSize().
+ *
+ * TODO: Does ShmemMappingSizes sound better?
+ */
+typedef struct MemoryMappingSizes
+{
+ Size shmem_req_size; /* Required size of the segment */
+ Size shmem_reserved; /* Required size of the reserved address
+ * space. */
+} MemoryMappingSizes;
+
+extern PGDLLIMPORT PGInhShmemSeg InhShmemSegs[NUM_MEMORY_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -64,14 +137,6 @@ typedef enum
SHMEM_TYPE_MMAP,
} PGShmemType;
-#ifndef WIN32
-extern PGDLLIMPORT unsigned long UsedShmemSegID;
-#else
-extern PGDLLIMPORT HANDLE UsedShmemSegID;
-extern PGDLLIMPORT void *ShmemProtectiveRegion;
-#endif
-extern PGDLLIMPORT void *UsedShmemSegAddr;
-
#if !defined(WIN32) && !defined(EXEC_BACKEND)
#define DEFAULT_SHARED_MEMORY_TYPE SHMEM_TYPE_MMAP
#elif !defined(WIN32)
@@ -85,10 +150,35 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(Size size,
+extern PGShmemHeader *PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping_sizes,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+extern void PrepareHugePages(void);
+
+static inline const char *
+MappingName(int segment_id)
+{
+ switch (segment_id)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ case BUFFER_DESCRIPTORS_SHMEM_SEGMENT:
+ return "descriptors";
+ case BUFFER_IOCV_SHMEM_SEGMENT:
+ return "iocv";
+ case CHECKPOINT_BUFFERS_SHMEM_SEGMENT:
+ return "checkpoint";
+ case STRATEGY_SHMEM_SEGMENT:
+ return "strategy";
+ default:
+ return "unknown";
+ }
+}
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index e71a51dfe84..e227119f59d 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -26,18 +26,20 @@
/* shmem.c */
-extern PGDLLIMPORT slock_t *ShmemLock;
typedef struct PGShmemHeader PGShmemHeader; /* avoid including
* storage/pg_shmem.h here */
-extern void InitShmemAccess(PGShmemHeader *seghdr);
-extern void InitShmemAllocation(void);
-extern void *ShmemAlloc(Size size);
+
+extern void InitShmemAccess(int segment_id, PGShmemHeader *seghdr, slock_t *ShmemLock);
+extern slock_t *InitShmemAllocation(int segment_id);
+extern void *ShmemAlloc(int segment_id, Size size);
extern void *ShmemAllocNoError(Size size);
-extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValid(int segment_id, const void *addr);
extern void InitShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
HASHCTL *infoP, int hash_flags);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int segment_id);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
@@ -59,6 +61,7 @@ typedef struct
void *location; /* location in shared mem */
Size size; /* # bytes requested for the structure */
Size allocated_size; /* # bytes actually allocated */
+ int segment_id; /* segment in which the structure is allocated */
} ShmemIndexEnt;
#endif /* SHMEM_H */
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index f4ee2bd7459..1e1bd1eb8b4 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1768,14 +1768,21 @@ pg_shadow| SELECT pg_authid.rolname AS usename,
LEFT JOIN pg_db_role_setting s ON (((pg_authid.oid = s.setrole) AND (s.setdatabase = (0)::oid))))
WHERE pg_authid.rolcanlogin;
pg_shmem_allocations| SELECT name,
+ segment,
off,
size,
allocated_size
- FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size);
+ FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, segment, off, size, allocated_size);
pg_shmem_allocations_numa| SELECT name,
numa_node,
size
FROM pg_get_shmem_allocations_numa() pg_get_shmem_allocations_numa(name, numa_node, size);
+pg_shmem_segments| SELECT id,
+ name,
+ size,
+ freeoffset,
+ reserved_size
+ FROM pg_get_shmem_segments() pg_get_shmem_segments(id, name, size, freeoffset, reserved_size);
pg_stat_activity| SELECT s.datid,
d.datname,
s.pid,
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index ddbe4c64971..d1d4f2b41b4 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -120,6 +120,7 @@ AmcheckOptions
AnalyzeAttrComputeStatsFunc
AnalyzeAttrFetchFunc
AnalyzeForeignTable_function
+AnonShmemSegment
AnlExprData
AnlIndexData
AnyArrayType
@@ -1685,6 +1686,7 @@ MVNDistinct
MVNDistinctItem
ManyTestResource
ManyTestResourceKind
+MemoryMappingSizes
Material
MaterialPath
MaterialState
@@ -1887,6 +1889,7 @@ PGFInfoFunction
PGFileType
PGFunction
PGIOAlignedBlock
+PGInhShmemSeg
PGLZ_HistEntry
PGLZ_Strategy
PGLoadBalanceType
@@ -2805,6 +2808,7 @@ ShellTypeInfo
ShippableCacheEntry
ShippableCacheKey
ShmemIndexEnt
+ShmemSegment
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-06 09:25 ` Bowen Shi <zxwsbg12138@gmail.com>
2026-02-06 10:00 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2 siblings, 1 reply; 167+ messages in thread
From: Bowen Shi @ 2026-02-06 09:25 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
Hi Ashutosh,
I tried applying the v20260128 patches but encountered conflicts.
Could you let me know the base commit they were developed against?
Thanks
On Fri, Feb 6, 2026 at 5:21 PM Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
wrote:
> On Fri, Nov 14, 2025 at 5:23 PM Ashutosh Bapat
> <ashutosh.bapat.oss@gmail.com> wrote:
> >
> > Hi,
> > PFA new patchset with some TODOs from previous email addressed:
> >
> > On Mon, Oct 13, 2025 at 9:28 PM Ashutosh Bapat
> > <ashutosh.bapat.oss@gmail.com> wrote:
> > > 1. New backends join while the synchronization is going on.
> >
> > Done. Explained the solution below in detail.
> >
> > > An existing backend exiting.
> >
> > Not tested specifically, but should work.
> >
> > > 2. Failure or crash in the backend which is executing
> pg_resize_buffer_pool()
> >
> > still a TODO
> >
> > > 3. Fix crashes in the tests.
> >
> > core regression passes, pg_buffercache regression tests pass and the
> > tests for buffer resizing pass most of the time. So far I have seen
> > two issues
> > 1. An assertion from AIO worker - which happened only once and I
> > couldn't reproduce again. Need to study interaction of AIO worker with
> > buffer resizing.
> > 2. checkpointer crashes - which is one of the TODOs listed below.
> > 3. Also there's an shared memory id related failure, which I don't
> > understand but happen more frequently than the first one. Need to look
> > into that.
> >
> > > go through Tomas's detailed comments and address those
> > > which still apply.
> >
> > Still a TODO. But since many of those patches are revised heavily, I
> > think many of the comments may have been addressed, some may not apply
> > anymore.
>
> Thanks a lot Tomas for these review comments and summary of the
> discussion as of your response. Sorry, it took so long to fully
> respond to this email, but I wanted the patches to be in a reasonable
> shape before responding. It's been a while since I have posted the
> last patchset. Hence attaching patchset along with responses to your
> comments as well as the earlier comments that you mentioned in your
> response. Dmitry has responded to some parts of these emails already
> but I am responding to all the comments in the context of the latest
> implementation which has been revised heavily since his responses. The
> latest design and UI is described in details in [1]. Please note that
> that email also responds to some earlier comments from Andres and
> others. I have not repeated those responses here.
>
> The earlier patches did not build in EXEC_BACKEND mode. The attached
> patches build in EXEC_BACKEND mode and also the regression tests pass
> in that build. I haven't tried running whole test suite though.
>
> >
> > I agree it'd be useful to be able to resize shared buffers, without
> > having to restart the instance (which is obviously very disruptive). So
> > if we can make this work reliably, with reasonable trade offs (both on
> > the backends, and also the risks/complexity introduced by the feature).
>
> Agreed. That's the goal - to make it work reliably with reasonable
> trade-offs. Adding some details to this goal as follows:
>
> 1. Performance
> --------------
> a. The impact of resizing on the performance of the concurrent
> database operations should be reasonable and acceptable.
> b. The code changes should not cause significant performance
> degradation during normal operations
>
> Palak Chaturvedi has already done some benchmarking. I am listing her
> findings here
> 1. The code changes do not have any noticeable performance degradation
> during normal operations.
> 2. Resizing itself is also reasonably fast and with very minimal,
> almost unnoticeable, performance impact on the concurrent
> transactions.
>
> We will share the benchmarks once the code is cleaner and more complete
> stable.
>
> 2. Reliability
> --------------
> We are adding tests to make sure that resizing does not introduce any
> crashes, hazards or data corruption. Some tests are already part of
> the patch set.
>
> >
> > I'm far from an expert on mmap() and similar low-level stuff, but the
> > current appproach (reserving a big chunk of shared memory and slicing
> > it by mmap() into smaller segments) seems reasonable.
>
> The latest patches, we reserve address space by mmap and manage memory
> within that address space using file backed memory.
>
> >
> > But I'm getting a bit lost in how exactly this interacts with things
> > like overcommit, system memory accounting / OOM killer and this sort of
> > stuff. I went through the thread and it seems to me the reserve+map
> > approach works OK in this regard (and the messages on linux-mm seem to
> > confirm this). But this information is scattered over many messages and
> > it's hard to say for sure, because some of this might be relevant for
> > an earlier approach, or a subtly different variant of it.
>
> Looking at the documentation and the messages on linux-mm, it seems
> that the mmap with MAP_NORESERVE approach works well with overcommit,
> system memory accounting and OOM killer. I plan to add tests verify
> these aspects, so that any changes in the future do not cause any
> regressions. I am looking for ways to write a small test program for
> the same, which we can add to our test battery. But no success so far.
> Any idea?
>
> >
> > A similar question is portability. The comments and commit messages
> > seem to suggest most of this is linux-specific, and other platforms just
> > don't have these capabilities. But there's a bunch of messages (mostly
> > by Thomas Munro) that hint FreeBSD might be capable of this too, even if
> > to some limited extent. And possibly even Windows/EXEC_BACKEND, although
> > that seems much trickier.
> >
> > FWIW I think it's perfectly fine to only support resizing on selected
> > platforms, especially considering Linux is the most widely used system
> > for running Postgres. We still need to be able to build/run on other
> > systems, of course. And maybe it'd be good to be able to disable this
> > even on Linux, if that eliminates some overhead and/or risks for people
> > who don't need the feature. Just a thought.
>
> +1. Plan is to make it work on Linux first and disable it on the other
> platforms. I think it's a good idea to be able to disable it at initdb
> time say. But I am not yet sure if that's going to avoid any overhead
> or risks. I haven't yet wrapped my head around the platform dependency
> completely.
>
> >
> > Anyway, my main point is that this information is important, but very
> > scattered over the thread. It's a bit foolish to expect everyone who
> > wants to do a review to read the whole thread (which will inevitably
> > grow longer over time), and assemble all these pieces again an again,
> > following all the changes in the design etc. Few people will get over
> > that hurdle, IMHO.
> >
> > So I think it'd be very helpful to write a README, explaining the
> > currnent design/approach, and summarizing all these aspects in a single
> > place. Including things like portability, interaction with the OS
> > accounting, OOM killer, this kind of stuff. Some of this stuff may be
> > already mentioned in code comments, but you it's hard to find those.
> >
> > Especially worth documenting are the states the processes need to go
> > through (using the barriers), and the transacitons between them (i.e.
> > what is allowed in each phase, what blocks can be visible, etc.).
> >
>
> That's a great idea. There is already a README in
> src/backend/storage/buffer/, I have added a section of buffer pool
> resizing. The section has pointers to relevant code comments. I will
> revise this to be as self sufficient as possible as the patches
> mature. The code managing the shared memory is scattered across
> backend/portability, storage/ipc etc. It's slightly troublesome to
> find one good place to add a README which explains everything.
>
> >
> > I'll go over some higher-level items first, and then over some comments
> > for individual patches.
> >
> >
> > 1) no user docs
> >
> > There are no user .sgml docs, and maybe it's time to write some,
> > explaining how to use this thing - how to configure it, how to trigger
> > the resizing, etc. It took me a while to realize I need to do ALTER
> > SYSTEM + pg_reload_conf() to kick this off.
>
> Latest patches have updated config.sgml and func-admin.sgml to
> document the functionality. Please review them.
>
> >
> > It should also document the user-visible limitations, e.g. what activity
> > is blocked during the resizing, etc.
> >
>
> We are still figuring this out. But we will document the limitations
> in func-admin.sgml. Right now there are no limitations and my
> intention is to keep the limitations as limited as possible.
>
> >
> > 2) pending GUC changes
> >
> > I'm somewhat skeptical about the GUC approach. I don't think it was
> > designed with this kind of use case in mind, and so I think it's quite
> > likely it won't be able to handle it well.
> >
> > For example, there's almost no validation of the values, so how do you
> > ensure the new value makes sense? Because if it doesn't, it can easily
> > crash the system (I've seen such crashes repeatedly, I'll get to that).
> > Sure, you may do ALTER SYSTEM to set shared_buffers to nonsense and it
> > won't start after restart/reboot, but crashing an instance is maybe a
> > little bit more annoying.
> >
> > Let's say we did the ALTER SYSTEM + pg_reload_conf(), and it gets stuck
> > waiting on something (can't evict a buffer or something). How do you
> > cancel it, when the change is already written to the .auto.conf file?
> > Can you simply do ALTER SYSTEM + pg_reload_conf() again?
> >
> > It also seems a bit strange that the "switch" gets to be be driven by a
> > randomly selected backend (unless I'm misunderstanding this bit). It
> > seems to be true for the buffer eviction during shrinking, at least.
> >
> > Perhaps this should be a separate utility command, or maybe even just
> > a new ALTER SYSTEM variant? Or even just a function, similar to what
> > the "online checksums" patch did, possibly combined with a bgworder
> > (but probably not needed, there are no db-specific tasks to do).
> >
>
> As documented in func-admin.sgml, and config.sgml, resizing the buffer
> pool requires ALTER SYSTEM + pg_reload_conf() followed by
> pg_resize_shared_buffers(). There is no pending flag, no arbitrary
> backend driving the process. The validation, failures are handled by
> pg_resize_shared_buffers(). The backend where
> pg_resize_shared_buffers() is called is responsible for driving the
> resizing process. I think, the new implementation takes care of all
> the concerns mentioned above.
>
> >
> > 3) max_available_memory
> >
> > Speaking of GUCs, I dislike how max_available_memory works. It seems a
> > bit backwards to me. I mean, we're specifying shared_buffers (and some
> > other parameters), and the system calculates the amount of shared memory
> > needed. But the limit determines the total limit?
> >
> > I think the GUC should specify the maximum shared_buffers we want to
> > allow, and then we'd work out the total to pre-allocate? Considering
> > we're only allowing to resize shared_buffers, that should be pretty
> > trivial. Yes, it might happen that the "total limit" happens to exceed
> > the available memory or something, but we already have the problem
> > with shared_buffers. Seems fine if we explain this in the docs, and
> > perhaps print the calculated memory limit on start.
> >
> > In any case, we should not allow setting a value that ends up
> > overflowing the internal reserved space. It's true we don't have a good
> > way to do checks for GUcs, but it's a bit silly to crash because of
> > hitting some non-obvious internal limit that we necessarily know about.
> >
> > Maybe this is a reason why GUC hooks are not a good way to set this.
>
> Latest patches introduce max_shared_buffers which specifies the
> maximum shared buffers allowed. max_available_memory, that some
> previous patchsets had, has been removed. This GUC is mentioned in the
> documentation changes.
>
> One relatively small thing to think about is what do we call
> shared_buffers GUC - it's not SIGHUP per say since it requires a
> function to be called after the reload but it's not POSTMASTER either
> since restart is not required. I think we need a new PGC_ for it but
> haven't yet thought of a good name. Do you have any suggestions?
>
> >
> >
> > 4) SHMEM_RESIZE_RATIO
> >
> > The SHMEM_RESIZE_RATIO thing seems a bit strange too. There's no way
> > these ratios can make sense. For example, BLCKSZ is 8192 but the buffer
> > descriptor is 64B. That's 128x difference, but the ratios says 0.6 and
> > 0.1, so 6x. Sure, we'll actually allocate only the memory we need, and
> > the rest is only "reserved".
> >
> > However, that just makes the max_available_memory a bit misleading,
> > because you can't ever use it. You can use the 60% for shared buffers
> > (which is not mentioned anywhere, and good luck not overflowing that,
> > as it's never checked), but those smaller regions are guaranteed to be
> > mostly unused. Unfortunate.
> >
> > And it's not just a matter of fixing those ratios, because then someone
> > rebuilds with 32kB blocks and you're in the same situation.
> >
> > Moreover, all of the above is for mappings sized based on NBuffers. But
> > if we allocate 10% for MAIN_SHMEM_SEGMENT, won't that be a problem the
> > moment someone increases of max_connection, max_locks_per_transaction
> > and possibly some other stuff?
> >
>
> +1. Fixed. max_shared_buffers only deals with the shared buffers.
>
> >
> > 5) no tests
> >
> > I mentioned no "user docs", but the patch has 0 tests too. Which seems
> > a bit strange for a patch of this age.
> >
> > A really serious part of the patch series seems to be the coordination
> > of processes when going through the phases, enforced by the barriers.
> > This seems like a perfect match for testing using injection points, and
> > I know we did something like this in the online checksums patch, which
> > needs to coordinate processes in a similar way.
> >
> > But even just a simple TAP test that does a bunch of (random?) resizes
> > while running a pgbench seem better than no tests. (That's what I did
> > manually, and it crashed right away.)
> >
> > There's a lot more stuff to test here, I think. Idle sessions with
> > buffers pinned by open cursors, multiple backends doing ALTER SYSTEM
> > + pg_reload_conf concurrently, other kinds of failures.
> >
>
> +1. The latest patches already have a few tests but we are adding more
> tests to cover different scenarios.
>
> >
> > 6) SIGBUS failures
> >
> > As mentioned, I did some simple tests with shrink/resize with a pgbench
> > in the background, and it almost immediately crashed for me :-( With a
> > SIGBUS, which I think is fairly rare on x86 (definitely much less common
> > than e.g. SIGSEGV).
> >
> > An example backtrace attached.
> >
>
> There is a test which runs pgbench concurrently with resizing and it
> fails much less frequently (1 in 20 or even lower). Quite likely the
> crash you have seen is also fixed. Please let me know if the test
> fails for you and if there is something that is in your test but not
> that test.
>
> >
> > 7) EXEC_BACKEND, FreeBSD
> >
> > We clearly need to keep this working on systems without the necessary
> > bits (so likely EXEC_BACKEND, FreeBSD etc.). But the builds currently
> > fail in both cases, it seems.
> >
> > I think it's fine to not support resizing on every platform, then we'd
> > never get it, but it still needs to build. It would be good to not have
> > two very different code versions, one for resizing and one without it,
> > though. I wonder if we can just have the "no-resize" use the same struct
> > (with the segments/mapping, ...) and all that, but skipping the space
> > reservation.
> >
>
> I have not thought fully through support on platforms other than
> Linux. However, my current plan is to not support GUC
> max_shared_buffers and the resizing function
> pg_resize_shared_buffers() on platforms which do not support the
> necessary bits. The code will still build and run on those platforms,
> but resizing shared buffers without restart won't be possible.
>
> >
> > 8) monitoring
> >
> > So, let's say I start a resize of shared buffers. How will I know what
> > it's currently doing, how much longer it might take, what it's waiting
> > for, etc.? I think it'd be good to have progress monitoring, through
> > the regular system view (e.g. pg_stat_shmem_resize_progress?).
> >
>
> When pg_resize_shared_buffers() finishes, user knows that the resizing
> is finished. It usually takes a few seconds for the resizing to
> finish. I am not sure whether we will need progress reporting. But
> there might be cases where the resizing has to wait for something or
> it may take longer for a given phase e.g. eviction. So we may require
> progress reporting. Have added TODO in the code for the same.
>
> >
> > 10) what to do about stuck resize?
> >
> > AFAICS the resize can get stuck for various reasons, e.g. because it
> > can't evict pinned buffers, possibly indefinitely. Not great, it's not
> > clear to me if there's a way out (canceling the resize) after a timeout,
> > or something like that? Not great to start an "online resize" only to
> > get stuck with all activity blocked for indefinite amount of time, and
> > get to restart anyway.
> >
> > Seems related to Thomas' message [2], but AFAICS the patch does not do
> > anything about this yet, right? What's the plan here?
> >
>
> pg_resize_shared_buffers() is affected by timeouts like
> statement_timeouts as well as query cancellation. However, current
> implementation lacks graceful handling of these events. Another TODO.
>
> >
> > 11) preparatory actions?
> >
> > Even if it doesn't get stuck, some of the actions can take a while, like
> > evicting dirty buffers before shrinking, etc. This is similar to what
> > happens on restart, when the shutdown checkpoint can take a while, while
> > the system is (partly) unavailable.
> >
> > The common mitigation is to do an explicit checkpoint right before the
> > restart, to make the shutdown checkpoint cheap. Could we do something
> > similar for the shrinking, e.g. flush buffers from the part to be
> > removed before actually starting the resize?
> >
>
> Eviction is carried out before changes to the memory or shared buffer
> metadata. If eviction fails, the resizing is rolled back. The system
> continues to work with old buffer pool size. A test particularly
> testing this aspect remains to be added.
>
> >
> > 12) does this affect e.g. fork() costs?
> >
> > I wonder if this affects the cost of fork() in some undesirable way?
> > Could it make fork() more measurably more expensive?
> >
>
> A new backend needs to attach to extra shared memory segments, AFAIK
> and detach those when exiting. But I don't think there's any other
> extra work that a backend needs to do when starting or exiting.
> Windows might be different. I haven't seen any noticeable difference
> in pgbench performance which uses new connection for every
> transaction. However, we haven't tried a benchmark with empty
> transactions so that fork() and exit() are exercised at a higher
> frequency. Another TODO.
>
> >
> > 13) resize "state" is all over the place
> >
> > For me, a big hurdle when reasoning about the resizing correctness is
> > that there's quite a lot of distinct pieces tracking what the current
> > "state" is. I mean, there's:
> >
> > - ShmemCtrl->NSharedBuffers
> > - NBuffers
> > - NBuffersOld
> > - NBuffersPending
> > - ... (I'm sure I missed something)
> >
> > There's no cohesive description how this fits together, it seems a bit
> > "ad hoc". Could be correct, but I find it hard to reason about.
> >
>
> Next set of patches will consolidate the state only in two places
> NBuffersPending and StrategyControl, which needs a new name. But even
> in the attached patches it's consolidated in NBuffersPending,
> ShmemCtrl and StrategyControl. There are instances of NBuffers which
> will be replaced by variables from ShmemCtrl or StrategyControl.
>
> >
> > 14) interesting messages from the thread
> >
> > While reading through the thread, I noticed a couple messages that I
> > think are still relevant:
> >
> > - I see Peter E posted some review in 2024/11 [3], but it seems his
> > comments were mostly ignored. I agree with most of them.
>
> Find detailed reply to that email at the end.
>
> >
> > - Robert mentioned a couple interesting failure scenarios in [4], not
> > sure if all of this was handled. He howerver assumes pointers would
> > not be stable (and that's something we should not allow, and the
> > current approach works OK in this regard, I think). He also outlines
> > how it'd happen in phases - this would be useful for the design README
> > I think. It also reminds me the "phases" in the checksums patch.
> >
>
> stable pointers in the latest patch. The phases are explained in the
> prologue of pg_resize_shared_buffers().
>
> > - Robert asked [5] if Linux might abruptly break this, but I find that
> > unlikely. We'd point out we rely on this, and they'd likely rethink.
> > This would be made safer if this was specified by POSIX - taking that
> > away once implemented seems way harder than for custom extensions.
> > It's likely they'd not take away the feature without an alternative
> > way to achieve the same effect, I think (yes, harder to maintain).
> > Tom suggests [7] this is not in POSIX.
> >
> > - Matthias mentioned [6] similar flags on other operating systems. Could
> > some of those be used to implement the same resizing?
>
> Part of portability TODO.
>
> >
> > - Andres had an interesting comment about how overcommit interacts with
> > MAP_NORESERVE. AFAIK it means we need the flag to not break overcommit
> > accounting. There's also some comments about from linux-mm people [9].
> >
>
> New implementation uses MAP_NORESERVE. See my earlier response about
> overcommit.
>
> > - There seem to be some issues with releasing memory backing a mapping
> > with hugetlb [10]. With the fd (and truncating the file), this seems
> > to release the memory, but it's linux-specific? But most of this stuff
> > is specific to linux, it seems. So is this a problem? With this it
> > should be working even for hugetlb ...
> >
>
> Right. mmap() + ftruncate(), instead of mmap() + mremap() allows us to
> avoid going through postmaster, which makes the implementation much
> simpler. And also support hugetlb. But it will be linux only.
>
> > - It seems FreeBSD has MFD_HUGETLB [11], so maybe we could use this and
> > make the hugetlb stuff work just like on Linux? Unclear. Also, I
> > thought the mfd stuff is linux-specific ... or am I confused?
> >
>
> Portability TODO.
>
> > - Andres objected to any approach without pointer stability, and I agree
> > with that. If we can figure out such solution, of course.
>
> Since we call mmap only once, the address of the mappping does not
> change; it is always stable. So pointer stability is guaranteed.
>
> >
> > - Thomas asked [13] why we need to stop all the backends, instead of
> > just waiting for them to acknowledge the new (smaller) NBuffers value
> > and then let them continue. I also don't quite see why this should
> > not work, and it'd limit the disruption when we have to wait for
> > eviction of buffers pinned by paused cursors, etc.
> >
>
> Approach in the latest patches does not stop all backends. Backends
> continue to work while resizing is in progress. They are synchronized
> at certain points using barriers. The details of synchronization for
> each subsystem and worker backend need to be worked out. We are
> working on that.
>
> >
> >
> > Now, some comments about the individual patches (some of this may be a
> > bit redundant with the earlier points):
>
> Since these patches have been heavily rewritten because we have
> redesigned and reimplemented the resizing process, some of the
> comments may not be relevant any more. I will try to answer the
> underlying concerns wherever applicable.
>
> >
> >
> > v5-0001-Process-config-reload-in-AIO-workers.patch
> >
> > 1) Hmmm, so which other workers may need such explicit handling? Do all
> > other processes participate in procsignal stuff, or does anything
> > need an explicit handling?
> >
>
> As Andres mentioned in [2], this is not needed. It's not part of the
> latest patches.
>
> >
> > v5-0003-Introduce-pss_barrierReceivedGeneration.patch
> >
> > 1) Do we actually need this? Isn't it enough to just have two barriers?
> > Or a barrier + condition variable, or something like that.
> >
> > 2) The comment talks about "coordinated way" when processing messages,
> > but it's not very clear to me. It should explain what is needed and
> > not possible with the current barrier code.
> >
> > 3) This very much reminds me what the online checksums patch needed to
> > do, and we managed to do it using plain barriers. So why does this
> > need this new thing? (No opinion on whether it's correct.)
> >
>
>
> This isn't required in the latest implementation. Removed from the
> latest patchset.
>
> >
> > v5-0004-Allow-to-use-multiple-shared-memory-mappings.patch
> >
> > 1) "int shmem_segment" - wouldn't it be better to have a separate enum
> > for this? I mean, we'll have a predefined list of segments, right?
> >
>
> +1. Andres suggested [2] to keep only two segments. So enum may be
> superfluous, but may be good if we need more segments in future. TODO
> for now.
>
> > 2) typedef struct AnonymousMapping would deserve some comment
>
> This structure no more exists. Instead we use MemoryMappingSizes to
> hold the required sizes of shared memory segments.
>
> >
> > 3) ANON_MAPPINGS - Probably should be MAX_ANON_MAPPINGS? But we'll know
> > how many we have, so why not to allocate exactly the right number?
> > Or even just an array of structs, like in similar cases?
> >
>
> +1. Renamed as NUM_MEMORY_MAPPINGS and is used to declare and traverse
> corresponding arrays.
>
> > 4) static int next_free_segment = 0;
> >
> > We exactly know what segments we'll create and in which order, no? So
> > why do we even bother with this next_free_segment thing? Can't we
> > simply declare an array of AnonymousMapping elements, with all the
> > elements, and then just walk it and calculate the sizes/pointers?
>
> next_free_segment is removed from the latest patches as it's not needed.
>
> >
> > 5) I'm a bit confused about the segment/mapping difference. The patch
> > seems to randomly mix those, or maybe I'm just confused. I mean,
> > we are creating just shmem segment, and the pieces are mappings,
> > right? So why do we index them by "shmem_segment"?
> >
> > Also, consider
> >
> > CreateAnonymousSegment(AnonymousMapping *mapping)
> >
> > so is that creating a segment or mapping? Or what's the difference?
> >
> > Or are we creating multiple segments, and I missed that? Or are there
> > different "segment" concepts, or what?
> >
> > 6) There should probably be some sort of API wrapping the mappings, so
> > that the various places don't need to mess with next_free_segments
> > directly, etc. Perhaps PGSharedMemoryCreate() shouldn't do this, and
> > should just pass size to CreateAnonymousSegment(), and that finding
> > empty slot in Mappings, etc.? Not sure that'll work, but it's a bit
> > error-prone if a struct is modified from multiple places like this.
>
> Fixed this confusion in the latest patches. There are multiple
> segments, each mapped to a different address space. Since we are using
> mmap() + ftruncate(), the memory allocated for each segment is
> controlled by a separate fds. If we use a single address space
> reservation and carve multiple segments out of it, we can not resize
> each segment separately. Just to clarify, we do not reserve a large
> part of address space and then carve it into smaller segments.
>
> >
> > 7) We should remember which segments got to use huge pages and which
> > did not. And we should make it optional for each segment. Although,
> > maybe I'm just confused about the "segment" definition - if we only
> > have one, that's where huge pages are applied.
> >
> > If we could have multiple segments for different segments (whatever
> > that means), not sure what we'll report for cases when some segments
> > get to use huge pages and others don't. Either because we don't want
> > to use that for some segments, or because we happen to run out of
> > the available huge pages.
>
> This is an interesting idea. In the current implementation either we
> use huge pages for all the segments or none of them. I think, per
> segment huge page usage will be an add-on feature in the next version.
> What do you think? Using separate segments for reserving separate
> address spaces will make it easy to use huge pages for some and not
> for others.
>
> >
> > 8) It seems PGSharedMemoryDetach got some significant changes, but the
> > comment was not modified at all. I'd guess that means the comment is
> > perhaps stale, or maybe there's something we should mention.
> >
>
> Done.
>
> > 9) I doubt the Assert on GetConfigOption needs to be repeated for all
> > segments (in CreateSharedMemoryAndSemaphores).
> >
>
> Done.
>
>
> > 10) Why do we have the Mapping and Segments indexed in different ways?
> > I mean, Mappings seem to be filled in FIFO (just grab the next free
> > slot), while Segments are indexed by segment ID.
> >
>
> Fixed this confusion in the latest patches. Both are indexed by segment ID.
>
> > 11) Actually, what's the difference between the contents of Mappings
> > and Segments? Isn't that the same thing, indexed in the same way?
> > Or could it be unified? Or are they conceptually different thing?
> >
>
> See explanation above.
>
> > 12) I believe we'll have a predefined list of segments, with fixed IDs,
> > so why not just have a MAX of those IDs as the capacity?
> >
>
> yes. Fixed in the latest patches.
>
> > 13) Would it be good to have some checks on shmem_segment values? That
> > it's valid with respect to defined segments, etc. An assert, maybe?
> > What about some asserts on the Mapping/Segment elements? To check
> > that the element is sensible, and that the arrays "match" (if we
> > need both).
> >
>
> That's a good idea. Added Asserts to that effect.
>
> > 14) Some of the lines got pretty long, e.g. in pg_get_shmem_allocations.
> > I suggest we define some macros to make this shorter, or something
> > like that.
> >
>
> Done.
>
> > 15) I'd maybe rename ShmemSegment to PGShmemSegment, for consistency
> > with PGShmemHeader?
>
> Actually we track information about the shared memory segments at
> multiple places. shmem.c has APIs similar to MemoryContext for shared
> memory. The information there is consolidated into ShmemSegment
> structure. pg_shmem.h has platform independent APIs to manage shared
> memory segments and then each implementation has implementation
> specific information. Some of that information is passed from parent
> process (postmaster) to child process (postgresql backends). I have
> created PGInhShmemSegment for the information that is passed from
> postmaster to backends and then AnonymousShmemSegment structure for
> tracking information about anonymous shared memory in sysv_shmem.c. I
> don't like PGInhShmemSegment name, but I haven't figured out a better
> name. I may rearrange these structures a bit more in the next patches.
>
> >
> > 16) Is MAIN_SHMEM_SEGMENT something we want to expose in a public header
> > file? Seems very much like an internal thing, people should access
> > it only through APIs ...
> >
>
> If we convert it to an enum, we can't avoid exposing it in a public
> header file. In future, we may allow users to allocate memory in
> specific segments. So I think it's fine to expose it.
>
> >
> > v5-0005-Address-space-reservation-for-shared-memory.patch
> >
> > 1) Shouldn't reserved_offset and huge_pages_on really be in the segment
> > info? Or maybe even in mapping info? (again, maybe I'm confused
> > about what these structs store)
>
> I can't find reserved_offset in the latest patches. huge_pages_on is
> now a global flag, either all segments use huge pages or none of them
> do. As mentioned earlier, per segment huge page usage may be an add-on
> separate feature.
>
> >
> > 2) CreateSharedMemoryAndSemaphores comment is rather light on what it
> > does, considering it now reserves space and then carves is into
> > segments.
> >
>
> The detailed comment is in PGSharedMemoryCreate(), which seems to be a
> better place for explaining memory reservation etc. I just change the
> comment here to mention plural shared memory segments and
> differentiate between shared memory segments and the shared memory
> structures. Please let me know if this looks good.
>
> > 3) So ReserveAnonymousMemory is what makes decisions about huge pages,
> > for the whole reserved space / all segments in it. That's a bit
> > unfortunate with respect to the desirability of some segments
> > benefiting from huge pages and others not. Maybe we should have two
> > "reserved" areas, one with huge pages, one without?
> >
>
> See my responses above about mapping vs segment and per segment huge page
> usage.
>
> > I guess we don't want too many segments, because that might make
> > fork() more expensive, etc. Just guessing, though. Also, how would
> > this work with threading?
>
> See my earlier response about fork() performance and also limiting
> number of segments to just 2. If we do see fork() is getting
> expensive, we can limit the number of segments to 1.
>
> >
> > 4) Any particular reason to define max_available_memory as
> > GUC_UNIT_BLOCKS and not GUC_UNIT_MB? Of course, if we change this
> > to have "max shared buffers limit" then it'd make sense to use
> > blocks, but "total limit" is not in blocks.
> >
>
> Right. See response about max_shared_buffers above.
>
> > 5) The general approach seems sound to me, but I'm not expert on this.
> > I wonder how portable this behavior is. I mean, will it work on other
> > Unix systems / Windows? Is it POSIX or Linux extension?
>
> See my earlier response about portability.
>
> >
> > 6) It might be a good idea to have Assert procedures to chech mappings
> > and segments (that it doesn't overflow reserved space, etc.). It
> > took me ages to realize I can change shared_buffers to >60% of the
> > limit, it'll happily oblige and then just crash with OOM when
> > calling mprotect().
> >
>
> This shouldn't happen with the latest patches. There are tests for the
> same. Please let me know if you still face the issue again in your
> tests.
>
> >
> > v5-0006-Introduce-multiple-shmem-segments-for-shared-buff.patch
> >
> > 1) I suspect the SHMEM_RESIZE_RATIO is the wrong direction, because it
> > entirely ignores relationships between the parts. See the earlier
> > comment about this.
> >
> > 2) In fact, what happens if the user tries to resize to a value that is
> > too large for one of the segments? How would the system know before
> > starting the resize (and failing)?
> >
>
> +1. No SHMEM_RESIZE_RATIO in the latest patches. Reservation is
> entirely based on max_shared_buffers.
>
> > 3) It seems wrong to modify the BufferManagerShmemSize like this. It's
> > probably better to have a "...SegmentSize" function for individual
> > segments, and let BufferManagerShmemSize() to still return a sum of
> > all segments.
> >
>
> I have reimplemented CalculateShmemSize() in the latest patches. Please
> review.
>
> > 4) I think MaxAvailableMemory is the wrong abstraction, because that's
> > not what people specify. See earlier comment.
> >
>
> +1. No MaxAvailableMemory in the latest patches.
>
> > 5) Let's say we change the shared memory size (ALTER SYSTEM), trigger
> > the config reload (pg_reload_conf). But then we find that we can't
> > actually shrink the buffers, for some unpredictable reason (e.g.
> > there's pinned buffers). How do we "undo" the change? We can't
> > really undo the ALTER SYSTEM, that's already written in the .conf
> > and we don't know the old value, IIRC. Is it reasonable to start
> > killing backends from the assign_hook or something? Seems weird.
> >
>
> The buffer pool resizing is done using the function
> pg_resize_shared_buffers(), which does not need rolling back the
> config value. As mentioned earlier, appropriate error handling is
> still a TBD.
>
> >
> > v5-0007-Allow-to-resize-shared-memory-without-restart.patch
> >
> > 1) Why would AdjustShmemSize be needed? Isn't that a sign of a bug
> > somewhere in the resizing?
> >
>
> It has been replaced by BufferManagerShmemResize() in the latest
> patches. BufferManagerShmemResize() only resizes the buffer manager
> related shared memory segments and data structures.
>
> > 2) Isn't the pg_memory_barrier() in CoordinateShmemResize a bit weird?
> > Why is it needed, exactly? If it's to flush stuff for processes
> > consuming EmitProcSignalBarrier, it's that too late? What if a
> > process consumes the barrier between the emit and memory barrier?
> >
> > 3) WaitOnShmemBarrier seem a bit under-documented.
> >
>
> Both of these functions are not required in the new implementation.
> Removed from the latest patches.
>
> > 4) Is this actually adding buffers to the freelist? I see buf_init only
> > links the new buffers by seeting freeNext, but where are the new
> > buffers added to the existing freelist?
> >
>
> Freelist does not exist anymore.
>
> > 5) The issue with a new backend seeing an old NBuffers value reminds me
> > of the "support enabling checksums online" thread, where we ran into
> > similar race conditions. See message [1], the part about race #2
> > (the other race might be relevant too, not sure). It's been a while,
> > but I think our conclusion ini that thread was that the "best" fix
> > would be to change the order of steps in InitPostgres(), i.e. setup
> > the ProcSignal stuff first, and only then "copy" the NBuffers value.
> > And handle the possibility that we receive a "duplicate" barriers.
> >
>
> I plan to remove NBuffers entirely and instead use shared memory
> variables so that we don't have to worry about backends inheriting
> stale values from Postmaster or even involving Postmaster in the
> resizing process. This is being worked upon and will be available in a
> future patchset.
>
> > 6) In fact, the online checksums thread seems like a possible source of
> > inspiration for some of the issues, because it needs to do similar
> > stuff (e.g. make sure all backends follow steps in a synchronized
> > way, etc.). And it didn't need new types of Barrier to do that.
> >
>
> Right. New implementation does not require any changes to the barrier
> implementation.
>
> > 7) Also, this seems like a perfect match for testing using injection
> > points. In fact, there's not a single test in the whole patch series.
> > Or a single line of .sgml docs, for that matter. It took me a while
> > to realize I'm supposed to change the size by ALTER SYSTEM + reload
> > the config.
> >
>
> There are some tests in the latest patches. More tests are being added.
>
> >
> > v5-0008-Support-shrinking-shared-buffers.patch
> >
> > 1) Why is ShmemCtrl->evictor_pid reset in AnonymousShmemResize? Isn't
> > there a place starting it and waiting for it to complete? Why
> > shouldn't it do EvictExtraBuffers itself?
> >
>
> evictor_pid is not used in the latest patches. Eviction is done in
> pg_resize_shared_buffers() itself by the backend which executes that
> function.
>
> > 2) Isn't the change to BufferManagerShmemInit wrong? How do we know the
> > last buffer is still at the end of the freelist? Seems unlikely.
> >
>
> No freelist anymore.
>
> > 3) Seems a bit strange to do it from a random backend. Shouldn't it
> > be the responsibility of a process like checkpointer/bgwriter, or
> > maybe a dedicated dynamic bgworker? Can we even rely on a backend
> > to be available?
> >
>
> The backend which executes pg_resize_shared_buffers() coordinates the
> resizing itself. I have not seen a need for it to use another
> background worker yet.
>
> > 4) Unsolved issues with buffers pinned for a long time. Could be an
> > issue if the buffer is pinned indefinitely (e.g. cursor in idle
> > connection), and the resizing blocks some activity (new connections
> > or stuff like that).
> >
>
> Resizing stops immediately after it encounters a pinned buffer. More
> sophisticated handling of such situations is a TBD and mostly v2.
>
> > 5) Funny that "AI suggests" something, but doesn't the block fail to
> > reset nextVictimBuffer of the clocksweep? It may point to a buffer
> > we're removing, and it'll be invalid, no?
> >
>
> TODO:
>
> > 6) It's not clear to me in what situations this triggers (in the call
> > to BufferManagerShmemInit)
> >
> > if (FirstBufferToInit < NBuffers) ...
> >
>
> That code does not exist in the latest patches.
>
> >
> > v5-0009-Reinitialize-StrategyControl-after-resizing-buffe.patch
> >
> > 1) IMHO this should be included in the earlier resize/shrink patches,
> > I don't see a reason to keep it separate (assuming this is the
> > correct way, and the "init" is not).
> >
>
> Right. Merged into the resizing patch in the latest patches.
>
> > 2) Doesn't StrategyPurgeFreeList already do some of this for the case
> > of shrinking memory?
>
> No freelist anymore.
>
> >
> > 3) Not great adding a bunch of static variables to bufmgr.c. Why do we
> > need to make "everything" static global? Isn't it enough to make
> > only the "valid" flag global? The rest can stay local, no?
> >
> > If everything needs to be global for some reason, could we at least
> > make it a struct, to group the fields, not just separate random
> > variables? And maybe at the top, not half-way throught the file?
> >
> > 4) Isn't the name BgBufferSyncAdjust misleading? It's not adjusting
> > anything, it's just invalidating the info about past runs.
>
> Being worked upon.
>
> >
> > 5) I don't quite understand why BufferSync needs to do the dance with
> > delay_shmem_resize. I mean, we certainly should not run BufferSync
> > from the code that resizes buffers, right? Certainly not after the
> > eviction, from the part that actually rebuilds shmem structs etc.
> > So perhaps something could trigger resize while we're running the
> > BufferSync()? Isn't that a bit strange? If this flag is needed, it
> > seems more like a band-aid for some issue in the architecture.
> >
> > 6) Also, why should it be fine to get into situation that some of the
> > buffers might not be valid, during shrinking? I mean, why should
> > this check (pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) !=
> NBuffers).
> > It seems better to ensure we never get into "sync" in a way that
> > might lead some of the buffers invalid. Seems way too lowlevel to
> > care about whether resize is happening.
> >
> > 7) I don't understand the new condition for "Execute the LRU scan".
> > Won't this stop LRU scan even in cases when we want it to happen?
> > Don't we want to scan the buffers in the remaining part (after
> > shrinking), for example? Also, we already checked this shmem flag at
> > the beginning of the function - sure, it could change (if some other
> > process modifies it), but does that make sense? Wouldn't it cause
> > problems if it can change at an arbitrary point while running the
> > BufferSync? IMHO just another sign it may not make sense to allow
> > this, i.e. buffer sync should not run during the "actual" resize.
> >
>
> Resizing runs for a few seconds. With default BufferSync frequency of
> 200ms, there will be a few cycles of BufferSync during resizing. If we
> don't let BufferSync run during resizing, we might have more dirty
> buffers to tackle post-resizing. I think it will be wise to let it run
> during resizing in the portion of the buffer pool which is not being
> removed. Working on that.
>
> >
> > v5-0010-Additional-validation-for-buffer-in-the-ring.patch
> >
> > 1) So the problem is we might create a ring before shrinking shared
> > buffers, and then GetBufferFromRing will see bogus buffers? OK, but
> > we should be more careful with these checks, otherwise we'll miss
> > real issues when we incorrectly get an invalid buffer. Can't the
> > backends do this only when they for sure know we did shrink the
> > shared buffers? Or maybe even handle that during the barrier?
> >
>
> AFAIK, the buffer rings are inside the Scan nodes, which are not
> accessible globally or to the barrier handling functions. So we can't
> purge the rings only once during or after resizing. If we already have
> a mechanism to access rings globally and hence in the barrier code, or
> we can develop such a mechanism, we can purge the rings only once.
> Please let me know if you have any ideas on this.
>
> > 2) IMHO a sign there's the "transitions" between different NBuffers
> > values may not be clear enough, and we're allowing stuff to happen
> > in the "blurry" area. I think that's likely to cause bugs (it did
> > cause issues for the online checksums patch, I think).
> >
>
> NBuffers is going to be replaced by shared memory variables. The
> current code relies on buffer pool size being static for the entire
> duration of the server. We are now changing that assumption. So, yes
> there will be some "blurry" areas and bugs to tackle. We are
> developing tests to uncover bugs and fix them.
>
> Here's reply to email by Peter E [3].
>
> On Tue, Nov 19, 2024 at 6:27 PM Peter Eisentraut <peter@eisentraut.org>
> wrote:
> >
> > On 18.10.24 21:21, Dmitry Dolgov wrote:
> > > v1-0001-Allow-to-use-multiple-shared-memory-mappings.patch
> > >
> > > Preparation, introduces the possibility to work with many shmem
> mappings. To
> > > make it less invasive, I've duplicated the shmem API to extend it with
> the
> > > shmem_slot argument, while redirecting the original API to it. There
> are
> > > probably better ways of doing that, I'm open for suggestions.
> >
> > After studying this a bit, I tend to think you should just change the
> > existing APIs in place. So for example,
> >
> > void *ShmemAlloc(Size size);
> >
> > becomes
> >
> > void *ShmemAlloc(int shmem_slot, Size size);
> >
> > There aren't that many callers, and all these duplicated interfaces
> > almost add more new code than they save.
> >
> > It might be worth making exceptions for interfaces that are likely to be
> > used by extensions. For example, I see pg_stat_statements using
> > ShmemInitStruct() and ShmemInitHash(). But that seems to be it. Are
> > there any other examples out there? Maybe there are many more that I
> > don't see right now. But at least for the initialization functions, it
> > doesn't seem worth it to preserve the existing interfaces exactly.
> >
> > In any case, I think the slot number should be the first argument. This
> > matches how MemoryContextAlloc() or also talloc() work.
> >
>
> Fixed. Please take a look at shmem.c changes.
>
> > (Now here is an idea: Could these just be memory contexts? Instead of
> > making six shared memory slots, could you make six memory contexts with
> > a special shared memory type. And ShmemAlloc becomes the allocation
> > function, etc.?)
>
> I don't see a need to do that right now. But will revisit this idea.
>
> >
> > I noticed the existing code made inconsistent use of PGShmemHeader * vs.
> > void *, which also bled into your patch. I made the attached little
> > patch to clean that up a bit.
> >
>
> The latest patches are rebased on top of your commit.
>
> > I suggest splitting the struct ShmemSegment into one struct for the
> > three memory addresses and a separate array just for the slock_t's. The
> > former struct can then stay private in storage/ipc/shmem.c, only the
> > locks need to be exported.
>
> > Also, maybe some of this should be declared in storage/shmem.h rather
> > than in storage/pg_shmem.h. We have the existing ShmemLock in there, so
> > it would be a bit confusing to have the per-segment locks elsewhere.
> >
>
> I have described the new structures above. Please review.
>
> >
> > Maybe rename ANON_MAPPINGS to something like NUM_ANON_MAPPINGS.
>
> Done.
>
> >
> > > v1-0003-Introduce-multiple-shmem-slots-for-shared-buffers.patch
> > >
> > > Splits shared_buffers into multiple slots, moving out structures that
> depend on
> > > NBuffers into separate mappings. There are two large gaps here:
> > >
> > > * Shmem size calculation for those mappings is not correct yet, it
> includes too
> > > many other things (no particular issues here, just haven't had
> time).
> > > * It makes hardcoded assumptions about what is the upper limit for
> resizing,
> > > which is currently low purely for experiments. Ideally there should
> be a new
> > > configuration option to specify the total available memory, which
> would be a
> > > base for subsequent calculations.
> >
> > Yes, I imagine a shared_buffers_hard_limit setting. We could maybe
> > default that to the total available memory, but it would also be good to
> > be able to specify it directly, for testing.
> >
>
> Latest patches introduce max_shared_buffers instead.
>
> >
> > > v1-0005-Use-anonymous-files-to-back-shared-memory-segment.patch
> > >
> > > Allows an anonyous file to back a shared mapping. This makes certain
> things
> > > easier, e.g. mappings visual representation, and gives an fd for
> possible
> > > future customizations.
> >
> > I think this could be a useful patch just by itself, without the rest of
> > the series, because of
> >
> > > * By default, Linux will not add file-backed shared mappings into a
> > > core dump, making it more convenient to work with them in PostgreSQL:
> > > no more huge dumps to process.
> >
> > This could be significant operational benefit.
> >
> > When you say "by default", is this adjustable? Does someone actually
> > want the whole shared memory in their core file? (If it's adjustable,
> > is it also adjustable for anonymous mappings?)
>
> 'man core' mentions that /proc/[PID]/coredump_filter file controls the
> memory segments written to the core file. The value in the file is a
> bitmask of memory mapping types. The bits correspond to the flags
> passed to mmap(). This file can be written to by the program through
> /proc/self/coredump_filter or externally by something like echo >
> /proc/PID/coredump_filter. A boot time setting becomes default for all
> the programs. A child process inherits the setting from parent through
> fork() and execve() does not change it. So it looks like DBAs should
> be able to control what gets dumped in postgresql core dumps by using
> any of those options. If we use file backed shared memory as the
> patches do, DBAs may have to change their default setting, which may
> not have considered filed backed shared memory. We will need to
> highlight this in the release notes. I have added a TODO for the same
> in the code for now. We will add this note in the commit message or
> update relevant document as the patches get committable.
>
> The man page says that bits 0, 1, 4 are set by default (anonymous
> private and shared mappings and ELF headers). But on my machine I see
> that the default is different. It's possible that different
> installations use different configurations.
>
> >
> > I'm wondering about this change:
> >
> > -#define PG_MMAP_FLAGS
> > (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE)
> > +#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE)
> >
> > It looks like this would affect all mmap() calls, not only the one
> > you're changing. But that's the only one that uses this macro! I don't
> > understand why we need this; I don't see anything in the commit log
> > about this ever being used for any portability. I think we should just
> > get rid of it and have mmap() use the right flags directly.
>
> This is committed as c100340729b66dc46d4f9d68a794957bf2c468d8.
>
> >
> > I see that FreeBSD has a memfd_create() function. Might be worth a try.
> > Obviously, this whole thing needs a configure test for memfd_create()
> > anyway.
>
> We are using memfd_create in the latest patches. I will look into
> creating a configure test for memfd_create(). I have added a TODO
> comment for the same.
>
> >
> > I see that memfd_create() has a MFD_HUGETLB flag. It's not very clear
> > how that interacts with the MAP_HUGETLB flag for mmap(). Do you need to
> > specify both of them if you want huge pages?
>
> We need both. I used toy program attached to [0] to verify that. Maybe
> we could use similar program for configure test.
>
> There are three important areas of puzzle that I will work on next:
> 1. Synchronization: The coordinator sends Proc signal barriers to all
> the concurrent backends during resizing. The code which deals with the
> buffer pool needs to react to these barriers and may need to adjust
> its course of action. Checkpointer, background writer, new buffer
> allocation, code scanning the buffer pool sequentially are a few
> examples that need to react to the barrier code.
>
> 2. Graceful handling of errors and failures during resizing.
>
> 3. Portability: A config test to check whether a platform has the APIs
> need to support this feature and enable the feature in the
> corresponding build. On other platforms build succeeds, regression
> runs but this feature is disabled.
>
> All the TODOs that mentioned above and those not covered by the above
> three points are added to the code. One of them being to reduce the
> number of segments to just two (main and shared buffer blocks). I plan
> to address them as the feature matures.
>
> [0]
> https://www.postgresql.org/message-id/CAExHW5sVxEwQsuzkgjjJQP9-XVe0H2njEVw1HxeYFdT7u7J%2BeQ%40mail.g...
> [1]
> https://www.postgresql.org/message-id/CAExHW5sOu8%2B9h6t7jsA5jVcQ--N-LCtjkPnCw%2BrpoN0ovT6PHg%40mail...
> [2]
> https://www.postgresql.org/message-id/qltuzcdxapofdtb5mrd4em3bzu2qiwhp3cdwdsosmn7rhrtn4u%40yaogvphfw...
> [3]
> https://www.postgresql.org/message-id/12add41a-7625-4639-a394-a5563e349322%40eisentraut.org
> [4]
> https://www.postgresql.org/message-id/CAExHW5vTWABxuM5fbQcFkGuTLwaxuZDEE2vtx2WuMUWk6JnF4g%40mail.gma...
> [5]
> https://www.postgresql.org/message-id/CAEze2WiMkmXUWg10y%2B_oGhJzXirZbYHB5bw0%3DVWte%2BYHwSBa%3DA%40...
>
> --
> Best Wishes,
> Ashutosh Bapat
>
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-06 09:25 ` Re: Changing shared_buffers without restart Bowen Shi <zxwsbg12138@gmail.com>
@ 2026-02-06 10:00 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2026-02-06 10:00 UTC (permalink / raw)
To: Bowen Shi <zxwsbg12138@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
HI Bowen,
Thanks for looking at the patches.
On Fri, Feb 6, 2026 at 2:56 PM Bowen Shi <zxwsbg12138@gmail.com> wrote:
>
> Hi Ashutosh,
>
> I tried applying the v20260128 patches but encountered conflicts. Could you let me know the base commit they were developed against?
>
The commit is recorded in the patch itself. base-commit:
e3094679b9835fed2ea5c7d7877e8ac8e7554d33.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-07 23:44 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2 siblings, 1 reply; 167+ messages in thread
From: Heikki Linnakangas @ 2026-02-07 23:44 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; +Cc: Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
I had a look at this (v20260128), focusing on
v20260128-0002-Memory-and-address-space-management-for-bu.patch It
introduces a new concept of "segment id" and exposes that to various places:
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
int segment_id)
So for each shmem struct, you now also specify 'segment_id'. (If you
call the old ShmemInitStruct() function, it defaults to MAIN_SHMEM_SEGMENT.)
I don't quite understand how you're supposed to use different segments
and when to use different "structs" in the same segment. The next patch
makes a segment resizeable:
+/*
+ * ShmemResizeStructInSegment -- Resize the given structure in shared
memory.
+ *
+ * This function resizes an existing shared memory structure while
preserving
+ * the existing memory location.
+ *
+ * Returns: pointer to the existing structure location, if the resize is
+ * successful, otherwise NULL.
+ */
+void *
+ShmemResizeStructInSegment(const char *name, Size size, bool *foundPtr,
+ int segment_id)
Ok, how does that actually work, if you allocate two structs in the
segment and start to resize them?
I think there's a tacit assumption here that if you want to be able to
resize a struct, it must be the only struct in the segment. If so,
what's the point of having a named struct in the segment in the first place?
I propose this API instead:
void
ShmemInitStructExt(const char *name, Size size, bool *foundPtr, bool
resizeable, Size max_size);
void *
ShmemResizeStruct(const char *name, Size size);
This completely hides the segment ids from the callers, it becomes
shmem.c's internal business. If you call ShmemInitStructExt with
resizeable==false, it can do the allocation from the main segment as
usual. But if you pass resizeable==true, then it creates a separate
segment for it, so that it can be resized.
What happens if you call ShmemInitStructExt() and the requested size
doesn't match the current size?
Could you write a standalone test module in src/test/modules to
demonstrate how to use the resizable shmem segments, please? That'd
allow focusing on that interface without worrying all the other
complexities of shared buffers.
- Heikki
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
@ 2026-02-09 15:15 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:14 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 22:37 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
0 siblings, 2 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2026-02-09 15:15 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On Sun, Feb 8, 2026 at 5:14 AM Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>
> I had a look at this (v20260128), focusing on
We committed refactoring of ShmemLock which is causing a merge
conflict. Attaching latest version of patchset with following changes
on top of 20260128.
1. The number of shared memory segments is reduced to two - main and
shared buffer blocks, as Andres suggested in [1]
2. Created separate structures to track properties of Anonymous shared
memory segments and UsedShmem segments.
3. Rebased on top of the ShmemLock infrastructure.
> v20260128-0002-Memory-and-address-space-management-for-bu.patch It
> introduces a new concept of "segment id" and exposes that to various places:
>
> +void *
> +ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr,
> int segment_id)
>
> So for each shmem struct, you now also specify 'segment_id'. (If you
> call the old ShmemInitStruct() function, it defaults to MAIN_SHMEM_SEGMENT.)
>
> I don't quite understand how you're supposed to use different segments
> and when to use different "structs" in the same segment. The next patch
> makes a segment resizeable:
>
> +/*
> + * ShmemResizeStructInSegment -- Resize the given structure in shared
> memory.
> + *
> + * This function resizes an existing shared memory structure while
> preserving
> + * the existing memory location.
> + *
> + * Returns: pointer to the existing structure location, if the resize is
> + * successful, otherwise NULL.
> + */
> +void *
> +ShmemResizeStructInSegment(const char *name, Size size, bool *foundPtr,
> + int segment_id)
>
> Ok, how does that actually work, if you allocate two structs in the
> segment and start to resize them?
e);
>
> I think there's a tacit assumption here that if you want to be able to
> resize a struct, it must be the only struct in the segment. If so,
> what's the point of having a named struct in the segment in the first place?
>
A segment may contain multiple structures but only the structure
allocated at the end of the segment is allowed to be resizable. If a
structure other than the end structure is tried to be resized,
following Assert will fail.
Assert((char *) segment->ShmemBase + ShmemAllocator->free_offset ==
(char *) result->location + result->allocated_size.
Please note the comment above this assertion is a bit misleading. Will
fix it in the next set of patches. Though in the attached patches we
allocate only one structure, buffer blocks, in BUFFERS_SHMEM_SEGMENT
(apart from PGShmemHeader and ShmemAllocator), it doesn't need to be
the only structure in the segment.
I may be missing something in your question, but the point of having a
named struct in the segment is to be able to discover it from
ShmemIndex.
> I propose this API instead:
>
> void
> ShmemInitStructExt(const char *name, Size size, bool *foundPtr, bool
> resizeable, Size max_size);
>
> void *
> ShmemResizeStruct(const char *name, Size size);
>
> This completely hides the segment ids from the callers, it becomes
> shmem.c's internal business. If you call ShmemInitStructExt with
> resizeable==false, it can do the allocation from the main segment as
> usual. But if you pass resizeable==true, then it creates a separate
> segment for it, so that it can be resized.
This is interesting ... more on this later.
>
> What happens if you call ShmemInitStructExt() and the requested size
> doesn't match the current size?
If the caller wants to fetch an existing resizable structure, it
shouldn't be required to know its current size because it may not know
its correct size when fetching it. The code in the patch is written in
a way to be compliant with current APIs as much. But if we are
introducing new APIs, I think we don't need to be that compliant.
> Could you write a standalone test module in src/test/modules to
> demonstrate how to use the resizable shmem segments, please? That'd
> allow focusing on that interface without worrying all the other
> complexities of shared buffers.
From your writeup it seems like you are leaning towards creating the
shared memory segments on-demand rather than having predefined
segments as done in the patch. I think your proposal is interesting
and might create a possibility for extensions to be able to create
resizable shared memory structures.
Let me first describe what happens in the current design and then
describe how we can implement on-demand (vs. dynamic) shared memory
segments that you seem to be leaning towards.
In the current patches, there are two predefined shared memory
segments: a. MAIN_SHMEM_SEGMENT b. BUFFER_SHMEM_SEGMENT. Each shared
memory segment in shmem.c is backed by a PGShmemHeader returned by
PGSharedMemoryCreate(). Each of these segments have one
PGUsedShmemInfo entry in UsedShmemInfo array and one AnonShmemData
entry in AnonShmemInfo array. When estimating the shared memory size
in CalculateShmemSize(), the sizes returned by all its minions are
counted against the main shared memory segment. Only
BufferManagerShmemSize() adds memory size of the buffer blocks against
BUFFERS_SHMEM_SEGMENT in the given MemoryMappingSizes array.
CreateSharedMemoryAndSemaphores() then creates shared memory for the
two segments and initializes ShmemSegments for the same. The
properties of the shared memory like mmap address or shmids etc. are
placed in PGUsedShmemInfo or AnonShmemData as appropriate. These
structures are used at the time of detaching/reattaching the shared
memory segments. Please note that the output produced by
CalculateShmemSize() is also used to report the total shared memory
used by InitializeShmemGUCs().
A side-note: I didn't find any existing README which describes how we
create and manage the shared memory segment. I think we should
probably write one and place it in storage/ipc as a separate patch and
then update the relevant parts in these patchset. Similarly I think we
could introduce PGUsedShmemInfo and AnonShmemData as a separate
commit.
After the shared memory segments are created shared data structures
are created. All the modules, including any extensions, create their
shared memory structures in the main memory segment. Only the buffer
manager creates BufferBlocks array in BUFFERS_SHMEM_SEGMENT, for which
it uses ShmemInitStructInSegment(). When resizing the buffer pool, the
shared memory segment is resized using PGSharedMemoryResize() and the
buffer blocks array is resized using ShmemResizeStructInSegment(). In
the current infrastructure ShmemInitStructExt() can not create a
shared memory segment right before allocating the data structure as
you suggest.
Since the shared memory segments are predefined, a test module, as you
suggest, does not have a segment that it can use. That led me to think
that you are imagining some kind of on-demand shared memory segments
that extensions or test modules can use. I think that will be a useful
feature by itself. Just as an example, imagine an extension to provide
shared plan cache which resizes the plan cache as needed. However, we
need to make sure that these on-demand shared memory segments work
well with CalculateShmemSize(), PGSharedMemoryDetach(),
AnonymousShmemDetach() etc. I can think of two ways to do this
1. In the current infrastructure, declare NUM_MEMORY_MAPPINGS to be
some larger number e.g. 10 to support 10 on-demand shared memory
segments. An internal module like BufferManagerShmemSize() directly
adds its sizes to the MemoryMappingSizes corresponding to the segment
of its choice (e.g. BUFFER_SHMEM_SEGMENT). Somehow we declare the
number of segments reserved for the internal modules. Each module then
adds its size requirements to the required number of
MemoryMappingSizes[] entries that is passed to shmem_request_hook. If
the number of entries used by all the modules higher than
NUM_MEMORY_MAPPING, the server can not start. These entries are later
transferred to MemoryMappingSizes[] that is passed to
CalculateShmemSize(). Each module is required to remember the
segment_ids it grabbed. CreateSharedMemoryAndSemaphores() then creates
the shared memory segments. The internal and external modules allocate
shared structures in the segments that they grabbed respectively. It
requires passing segment_id to ShmemInitStructExt() though.
2. There is no predefined limit on the number of segments. When
CalculateShmemSize() is called, BufferManagerShmemSize() outputs the
shared memory required in the main segment and total of shared memory
required in the other segments. Similarly shmem_request_hook of each
external module outputs the shared memory required in the main segment
and the total of shared memory required in the other segments. The
main segment is created in CreateSharedMemoryAndSemaphores() directly
using the requested size. The other output size is merely used for
reporting purposes in InitializeShmemGUCs(). When allocating shared
data structures, the main shared memory data structures are allocated
using ShmemInitStruct(), whereas the data structures in other memory
segments are allocated using ShmemInitStructExt() which takes
immediate allocation size and size of address space to be reserved as
arguments. It does not require segment_id though. For every new data
structure that ShmemInitStructExt() encounters, it a. creates a new
shared memory segment using PGSharedMemoryCreate(), b. allocates that
structure with the initial size. This function also needs to create
the PGShmemInfo and AnonShmemData entries corresponding to new
segments and make them available to PGSharedMemoryDetach(),
AnonymousShmemDetach() etc. For that we create a shared hash table in
the main shared memory segment (just like ShmemIndex) where we store
the metadata against each segment name (by coining it from the name of
the structure). We expect ShmemInitStructExt() to allocate structures
and segments only in the Postmaster and only at the beginning. When a
backend starts, it pulls the segment metadata from the shared hash
table which is ultimately used by PGSharedMemoryDetach(),
AnonymousShmemDetach(). At run time any backend which has access to
the shared memory should be able to call ShmemResizeStruct() given a
resizable structure and new size. Some higher level synchronization is
needed to ensure that the same structure is not resized simultaneously
by two backends.
The first approach is simple but has limited use given the fixed
number of segments. Second is more flexible but that's some work. I am
not sure whether it's worth doing all that if there are hardly any
extensions which could use resizable shared data structures.
Please let me know if you had something else in your mind when you
suggested a test module.
The first patch needs some clean up and work to make it committable.
But I was focusing on synchronization which still needs to be fleshed
out. However, if you think that a resizable shared memory segment is a
useful feature by itself that we can target for PG 19, I will focus on
making it committable.
[1] https://www.postgresql.org/message-id/qltuzcdxapofdtb5mrd4em3bzu2qiwhp3cdwdsosmn7rhrtn4u@yaogvphfwc4...
--
Best Wishes,
Ashutosh Bapat
Attachments:
[application/x-patch] v20260209-0001-Add-a-view-to-read-contents-of-shared-buff.patch (14.8K, ../../CAExHW5tSw8r06RLAArvf923cO4NGetitPhQ7AO0o7hsKx8jsNw@mail.gmail.com/2-v20260209-0001-Add-a-view-to-read-contents-of-shared-buff.patch)
download | inline diff:
From d602e084dbc76f441395665a23b0968474fce004 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 25 Aug 2025 19:23:50 +0530
Subject: [PATCH v20260209 1/7] Add a view to read contents of shared buffer
lookup table
The view exposes the contents of the shared buffer lookup table for
debugging, testing and investigation.
This helped me in debugging issues where the buffer descriptor array and
buffer lookup table were out of sync; either the buffer lookup table had
a mapping page->buffer which wasn't present in the buffer descriptor
array or a page in the buffer descriptor array didn't have corresponding
entry in the buffer lookup table. pg_buffercache doesn't help with those
kind of issues. Also doing that under the debugger in very painful.
I intend to keep this patch while the rest of the code matures. If it is
found useful as a debugging tool, we may consider make it committable
and commit it.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../expected/pg_buffercache.out | 39 ++++++++
.../pg_buffercache--1.5--1.6.sql | 24 +++++
contrib/pg_buffercache/pg_buffercache_pages.c | 18 ++++
contrib/pg_buffercache/sql/pg_buffercache.sql | 20 +++++
doc/src/sgml/system-views.sgml | 89 +++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 58 ++++++++++++
src/include/storage/buf_internals.h | 2 +
7 files changed, 250 insertions(+)
diff --git a/contrib/pg_buffercache/expected/pg_buffercache.out b/contrib/pg_buffercache/expected/pg_buffercache.out
index 886dea770f6..f0df2d2d3bc 100644
--- a/contrib/pg_buffercache/expected/pg_buffercache.out
+++ b/contrib/pg_buffercache/expected/pg_buffercache.out
@@ -33,6 +33,26 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
t
(1 row)
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -46,6 +66,10 @@ SELECT * FROM pg_buffercache_summary();
ERROR: permission denied for function pg_buffercache_summary
SELECT * FROM pg_buffercache_usage_counts();
ERROR: permission denied for function pg_buffercache_usage_counts
+SELECT * FROM pg_buffercache_lookup_table_entries();
+ERROR: permission denied for function pg_buffercache_lookup_table_entries
+SELECT * FROM pg_buffercache_lookup_table;
+ERROR: permission denied for view pg_buffercache_lookup_table
RESET role;
-- Check that pg_monitor is allowed to query view / function
SET ROLE pg_monitor;
@@ -73,6 +97,21 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts();
t
(1 row)
+RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
RESET role;
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
index 458f054a691..9bf58567878 100644
--- a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
+++ b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
@@ -44,3 +44,27 @@ CREATE FUNCTION pg_buffercache_evict_all(
OUT buffers_skipped int4)
AS 'MODULE_PATHNAME', 'pg_buffercache_evict_all'
LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Add the buffer lookup table function
+CREATE FUNCTION pg_buffercache_lookup_table_entries(
+ OUT tablespace oid,
+ OUT database oid,
+ OUT relfilenode oid,
+ OUT forknum int2,
+ OUT blocknum int8,
+ OUT bufferid int4)
+RETURNS SETOF record
+AS 'MODULE_PATHNAME', 'pg_buffercache_lookup_table_entries'
+LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Create a view for convenient access.
+CREATE VIEW pg_buffercache_lookup_table AS
+ SELECT * FROM pg_buffercache_lookup_table_entries();
+
+-- Don't want these to be available to public.
+REVOKE ALL ON FUNCTION pg_buffercache_lookup_table_entries() FROM PUBLIC;
+REVOKE ALL ON pg_buffercache_lookup_table FROM PUBLIC;
+
+-- Grant access to monitoring role.
+GRANT EXECUTE ON FUNCTION pg_buffercache_lookup_table_entries() TO pg_read_all_stats;
+GRANT SELECT ON pg_buffercache_lookup_table TO pg_read_all_stats;
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 89b86855243..f60f797a9b4 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -16,6 +16,7 @@
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "utils/rel.h"
+#include "utils/tuplestore.h"
#define NUM_BUFFERCACHE_PAGES_MIN_ELEM 8
@@ -107,6 +108,7 @@ PG_FUNCTION_INFO_V1(pg_buffercache_evict_all);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_relation);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_all);
+PG_FUNCTION_INFO_V1(pg_buffercache_lookup_table_entries);
/* Only need to touch memory once per backend process lifetime */
@@ -958,3 +960,19 @@ pg_buffercache_mark_dirty_all(PG_FUNCTION_ARGS)
PG_RETURN_DATUM(result);
}
+
+/*
+ * Return lookup table content as a set of records.
+ */
+Datum
+pg_buffercache_lookup_table_entries(PG_FUNCTION_ARGS)
+{
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* Fill the tuplestore */
+ BufTableGetContents(rsinfo->setResult, rsinfo->setDesc);
+
+ return (Datum) 0;
+}
diff --git a/contrib/pg_buffercache/sql/pg_buffercache.sql b/contrib/pg_buffercache/sql/pg_buffercache.sql
index 127d604905c..22e255c9721 100644
--- a/contrib/pg_buffercache/sql/pg_buffercache.sql
+++ b/contrib/pg_buffercache/sql/pg_buffercache.sql
@@ -18,6 +18,18 @@ from pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -26,6 +38,8 @@ SELECT * FROM pg_buffercache_os_pages;
SELECT * FROM pg_buffercache_pages() AS p (wrong int);
SELECT * FROM pg_buffercache_summary();
SELECT * FROM pg_buffercache_usage_counts();
+SELECT * FROM pg_buffercache_lookup_table_entries();
+SELECT * FROM pg_buffercache_lookup_table;
RESET role;
-- Check that pg_monitor is allowed to query view / function
@@ -36,6 +50,12 @@ SELECT buffers_used + buffers_unused > 0 FROM pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts();
RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+RESET role;
+
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 8b4abef8c68..c5683068470 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -71,6 +71,11 @@
<entry>backend memory contexts</entry>
</row>
+ <row>
+ <entry><link linkend="view-pg-buffer-lookup-table"><structname>pg_buffer_lookup_table</structname></link></entry>
+ <entry>shared buffer lookup table</entry>
+ </row>
+
<row>
<entry><link linkend="view-pg-config"><structname>pg_config</structname></link></entry>
<entry>compile-time configuration parameters</entry>
@@ -929,6 +934,90 @@ AND c1.path[c2.level] = c2.path[c2.level];
</para>
</sect1>
+ <sect1 id="view-pg-buffer-lookup-table">
+ <title><structname>pg_buffer_lookup_table</structname></title>
+ <indexterm>
+ <primary>pg_buffer_lookup_table</primary>
+ </indexterm>
+ <para>
+ The <structname>pg_buffer_lookup_table</structname> view exposes the current
+ contents of the shared buffer lookup table. Each row represents an entry in
+ the lookup table mapping a relation page to the ID of buffer in which it is
+ cached. The shared buffer lookup table is locked for a short duration while
+ reading so as to ensure consistency. This may affect performance if this view
+ is queried very frequently.
+ </para>
+ <table id="pg-buffer-lookup-table-view" xreflabel="pg_buffer_lookup_table">
+ <title><structname>pg_buffer_lookup_table</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>tablespace</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the tablespace containing the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>database</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database containing the relation (zero for shared relations)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relfilenode</structfield> <type>oid</type>
+ </para>
+ <para>
+ relfilenode identifying the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>forknum</structfield> <type>int2</type>
+ </para>
+ <para>
+ Fork number within the relation (see <xref linkend="storage-file-layout"/>)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>blocknum</structfield> <type>int8</type>
+ </para>
+ <para>
+ Block number within the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>bufferid</structfield> <type>int4</type>
+ </para>
+ <para>
+ ID of the buffer caching the page
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ <para>
+ Access to this view is restricted to members of the
+ <literal>pg_read_all_stats</literal> role by default.
+ </para>
+ </sect1>
+
<sect1 id="view-pg-config">
<title><structname>pg_config</structname></title>
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 23d85fd32e2..5089c7322f3 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -21,7 +21,12 @@
*/
#include "postgres.h"
+#include "fmgr.h"
+#include "funcapi.h"
#include "storage/buf_internals.h"
+#include "storage/lwlock.h"
+#include "utils/rel.h"
+#include "utils/builtins.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -159,3 +164,56 @@ BufTableDelete(BufferTag *tagPtr, uint32 hashcode)
if (!result) /* shouldn't happen */
elog(ERROR, "shared buffer hash table corrupted");
}
+
+/*
+ * BufTableGetContents
+ * Fill the given tuplestore with contents of the shared buffer lookup table
+ *
+ * This function is used by pg_buffercache extension to expose buffer lookup
+ * table contents via SQL. The caller is responsible for setting up the
+ * tuplestore and result set info.
+ */
+void
+BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc)
+{
+/* Expected number of attributes of the buffer lookup table entry. */
+#define BUFTABLE_CONTENTS_COLS 6
+
+ HASH_SEQ_STATUS hstat;
+ BufferLookupEnt *ent;
+ Datum values[BUFTABLE_CONTENTS_COLS];
+ bool nulls[BUFTABLE_CONTENTS_COLS];
+ int i;
+
+ memset(nulls, 0, sizeof(nulls));
+
+ Assert(tupdesc->natts == BUFTABLE_CONTENTS_COLS);
+
+ /*
+ * Lock all buffer mapping partitions to ensure a consistent view of the
+ * hash table during the scan. Must grab LWLocks in partition-number order
+ * to avoid LWLock deadlock.
+ */
+ for (i = 0; i < NUM_BUFFER_PARTITIONS; i++)
+ LWLockAcquire(BufMappingPartitionLockByIndex(i), LW_SHARED);
+
+ hash_seq_init(&hstat, SharedBufHash);
+ while ((ent = (BufferLookupEnt *) hash_seq_search(&hstat)) != NULL)
+ {
+ values[0] = ObjectIdGetDatum(ent->key.spcOid);
+ values[1] = ObjectIdGetDatum(ent->key.dbOid);
+ values[2] = ObjectIdGetDatum(ent->key.relNumber);
+ values[3] = ObjectIdGetDatum(ent->key.forkNum);
+ values[4] = Int64GetDatum(ent->key.blockNum);
+ values[5] = Int32GetDatum(ent->id);
+
+ tuplestore_putvalues(tupstore, tupdesc, values, nulls);
+ }
+
+ /*
+ * Release all buffer mapping partition locks in the reverse order so as
+ * to avoid LWLock deadlock.
+ */
+ for (i = NUM_BUFFER_PARTITIONS - 1; i >= 0; i--)
+ LWLockRelease(BufMappingPartitionLockByIndex(i));
+}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 27f12502d19..f4e9e703b8b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -29,6 +29,7 @@
#include "storage/spin.h"
#include "utils/relcache.h"
#include "utils/resowner.h"
+#include "utils/tuplestore.h"
/*
* Buffer state is a single 64-bit variable where following data is combined.
@@ -581,6 +582,7 @@ extern uint32 BufTableHashCode(BufferTag *tagPtr);
extern int BufTableLookup(BufferTag *tagPtr, uint32 hashcode);
extern int BufTableInsert(BufferTag *tagPtr, uint32 hashcode, int buf_id);
extern void BufTableDelete(BufferTag *tagPtr, uint32 hashcode);
+extern void BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc);
/* localbuf.c */
extern bool PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount);
base-commit: c67bef3f3252a3a38bf347f9f119944176a796ce
--
2.34.1
[application/x-patch] v20260209-0002-Memory-and-address-space-management-for-bu.patch (95.2K, ../../CAExHW5tSw8r06RLAArvf923cO4NGetitPhQ7AO0o7hsKx8jsNw@mail.gmail.com/3-v20260209-0002-Memory-and-address-space-management-for-bu.patch)
download | inline diff:
From 29f8a789f7e0931dc4e119d090a02e7b6a4c24cd Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 3 Feb 2026 10:58:03 +0530
Subject: [PATCH v20260209 2/7] Memory and address space management for buffer
resizing
This has three changes
1. Allow to use multiple shared memory mappings
============================================
Currently all the work with shared memory is done via a single anonymous
memory mapping, which limits ways how the shared memory could be organized.
Introduce possibility to allocate multiple shared memory mappings, where
a single mapping is associated with a specified shared memory segment.
Modifies pg_shmem_allocations to report shared memory segment as well.
Adds pg_shmem_segments to report shared memory segment information.
2. Address space reservation for shared memory
============================================
Currently the shared memory layout is designed to pack everything tight
together, leaving no space between mappings for resizing. Here is how it
looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
anonymous shared memory we talk about:
00400000-00490000 /path/bin/postgres
...
012d9000-0133e000 [heap]
7f443a800000-7f470a800000 /dev/zero (deleted)
7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
...
Make the layout more dynamic via splitting every shared memory segment
into two parts:
* An anonymous file, which actually contains shared memory content.
Such an anonymous file is created via memfd_create, it lives in
memory, behaves like a regular file and semantically equivalent to an
anonymous memory allocated via mmap with MAP_ANONYMOUS.
* A reservation mapping, which size is much larger than required shared
segment size. This mapping is created with flag MAP_NORESERVE (to not
count the reserved space against memory limits). The anonymous file is
mapped into this reservation mapping.
If we have to change the address maps while resizing the shared buffer
pool, it is needed to be done in Postmaster too, so that the new
backends will inherit the resized address space from the Postmaster.
However, Postmaster is not invovled in ProcSignalBarrier mechanism and
we don't want it to spend time in things other than its core
functionality. To achive that, maximum required address space maps are
setup upfront with read and write access when starting the server. When
resizing the buffer pool only the backing file object is resized from
the coordinator. This also makes the ProcSignalBarrier handling code
light for backends other than the coordinator.
The resulting layout looks like this:
00400000-00490000 /path/bin/postgres
...
3f526000-3f590000 rw-p [heap]
7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
To resize a shared memory segment in this layout it's possible to use
ftruncate on the memory mapped file.
This approach also do not impact the actual memory usage as reported by
the kernel.
TODO: Verify that Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup
was created with the memory limit 256 MB, then PostgreSQL was launched within
this cgroup with shared_buffers = 128 MB:
$ cd /sys/fs/cgroup
$ mkdir postgres
$ cd postres
$ echo 268435456 > memory.max
$ echo $MASTER_PID_SHELL > cgroup.procs
$ cat memory.current
17465344 (~16.6 MB)
$ echo $PATCH_PID_SHELL > cgroup.procs
$ cat memory.current
20770816 (~19.8 MB)
There are also few unrelated advantages of using memory mapped files:
* We've got a file descriptor, which could be used for regular file
operations (modification, truncation, you name it).
* The file could be given a name, which improves readability when it
comes to process maps.
* By default, Linux will not add file-backed shared mappings into a core dump,
making it more convenient to work with them in PostgreSQL: no more huge dumps
to process. - Some hackers have expressed concerns over it.
The downside is that memfd_create is Linux specific.
3. Refactor CalculateShmemSize()
=============================
This function calls many functions which return the amount of shared
memory required for different shared memory data structures. Up until
now, the returned total of these sizes was used to create a single
shared memory segment. With this change, CalculateShmemSize() needs to
estimate memory requirements for each of the segments. It now takes an
array of MemoryMappingSizes, containing as many elements as the number
of segments, as an argument. The sizes returned by all the function it
calls, except BufferManagerShmemSize(), are added and saved in the first
element (index 0) of the array. BufferManagerShmemSize() is modified to
save the amount of memory required for buffer manager related segments
in the corresponding array element. Additionally it also saves the
amount of reserved space. For now, the amount of reserved address space
is same as the amount of required memory but that is expected to change
with the next commit which implements buffer pool resize.
CalculateShmemSize() now returns the total of sizes corresponding to all
the sizes.
Author: Dmitrii Dolgov and Ashutosh Bapat
Reviewed-by: Tomas Vondra
---
doc/src/sgml/system-views.sgml | 9 +
src/backend/catalog/system_views.sql | 7 +
src/backend/port/posix_sema.c | 2 +-
src/backend/port/sysv_sema.c | 2 +-
src/backend/port/sysv_shmem.c | 552 ++++++++++++++++------
src/backend/port/win32_sema.c | 2 +-
src/backend/port/win32_shmem.c | 291 +++++++-----
src/backend/postmaster/launch_backend.c | 31 +-
src/backend/storage/buffer/buf_init.c | 47 +-
src/backend/storage/buffer/buf_table.c | 1 +
src/backend/storage/buffer/freelist.c | 7 +-
src/backend/storage/ipc/ipc.c | 4 +-
src/backend/storage/ipc/ipci.c | 100 +++-
src/backend/storage/ipc/shmem.c | 266 ++++++++---
src/backend/storage/lmgr/lwlock.c | 7 +-
src/backend/storage/lmgr/predicate.c | 3 +-
src/backend/utils/activity/pgstat_shmem.c | 3 +-
src/include/catalog/pg_proc.dat | 12 +-
src/include/storage/bufmgr.h | 3 +-
src/include/storage/ipc.h | 4 +-
src/include/storage/pg_shmem.h | 86 +++-
src/include/storage/shmem.h | 9 +-
src/test/regress/expected/rules.out | 9 +-
src/tools/pgindent/typedefs.list | 4 +
24 files changed, 1042 insertions(+), 419 deletions(-)
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index c5683068470..6fa47e3c63d 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -4305,6 +4305,15 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
</para></entry>
</row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>segment</structfield> <type>text</type>
+ </para>
+ <para>
+ The name of the shared memory segment concerning the allocation.
+ </para></entry>
+ </row>
+
<row>
<entry role="catalog_table_entry"><para role="column_definition">
<structfield>off</structfield> <type>int8</type>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 7553f31fef0..bc11589aeab 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -668,6 +668,13 @@ GRANT SELECT ON pg_shmem_allocations TO pg_read_all_stats;
REVOKE EXECUTE ON FUNCTION pg_get_shmem_allocations() FROM PUBLIC;
GRANT EXECUTE ON FUNCTION pg_get_shmem_allocations() TO pg_read_all_stats;
+CREATE VIEW pg_shmem_segments AS
+ SELECT * FROM pg_get_shmem_segments();
+
+REVOKE ALL ON pg_shmem_segments FROM PUBLIC;
+GRANT SELECT ON pg_shmem_segments TO pg_read_all_stats;
+REVOKE EXECUTE ON FUNCTION pg_get_shmem_segments() FROM PUBLIC;
+GRANT EXECUTE ON FUNCTION pg_get_shmem_segments() TO pg_read_all_stats;
CREATE VIEW pg_shmem_allocations_numa AS
SELECT * FROM pg_get_shmem_allocations_numa();
diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c
index e368e5ee7ed..5ad50c79dcd 100644
--- a/src/backend/port/posix_sema.c
+++ b/src/backend/port/posix_sema.c
@@ -216,7 +216,7 @@ PGReserveSemaphores(int maxSemas)
#else
sharedSemas = (PGSemaphore)
- ShmemAlloc(PGSemaphoreShmemSize(maxSemas));
+ ShmemAlloc(MAIN_SHMEM_SEGMENT, PGSemaphoreShmemSize(maxSemas));
#endif
numSems = 0;
diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c
index 86c4d359ef7..f0c7b064ffb 100644
--- a/src/backend/port/sysv_sema.c
+++ b/src/backend/port/sysv_sema.c
@@ -344,7 +344,7 @@ PGReserveSemaphores(int maxSemas)
DataDir)));
sharedSemas = (PGSemaphore)
- ShmemAlloc(PGSemaphoreShmemSize(maxSemas));
+ ShmemAlloc(MAIN_SHMEM_SEGMENT, PGSemaphoreShmemSize(maxSemas));
numSharedSemas = 0;
maxSharedSemas = maxSemas;
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 2e3886cf9fe..29ffcaa35f3 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -39,7 +39,17 @@
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
-
+/*
+ * TODO: The first two sentences in the first paragraph below make me feel like
+ * we should have only one SysV segment. Is that true? Needs investigation.
+ */
+/*
+ * TODO: third paragraph should mention that we use memfd_create to create
+ * shared memory segment, and possibly there's a way to share that segment
+ * between two processes using the file descriptor instead of going through SysV
+ * shared memory segment. So one day EXEC_BACKEND can also use anonymous shared
+ * memory.
+ */
/*
* As of PostgreSQL 9.3, we normally allocate only a very small amount of
* System V shared memory, and only for the purposes of providing an
@@ -91,12 +101,56 @@ typedef enum
SHMSTATE_UNATTACHED, /* pertinent to DataDir, no attached PIDs */
} IpcMemoryState;
+/*
+ * Anonymous mapping layout we use looks like this:
+ *
+ * 00400000-00c2a000 r-xp /bin/postgres
+ * ...
+ * 3f526000-3f590000 rw-p [heap]
+ * 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted)
+ * 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted)
+ * 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
+ * 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
+ * ...
+ *
+ * We need to place shared memory mappings in such a way, that there will be
+ * gaps between them in the address space. Those gaps have to be large enough
+ * to resize the mapping up to certain size, without counting towards the total
+ * memory consumption.
+ *
+ * To achieve this, for each shared memory segment we first create an anonymous
+ * file of specified size using memfd_create, which will accomodate actual
+ * shared memory mapping content. It is represented by the first /memfd:main
+ * with rw permissions. Then we create a mapping for this file using mmap, with
+ * size much larger than required and flags PROT_NONE (allows to make sure the
+ * reserved space will not be used) and MAP_NORESERVE (prevents the space from
+ * being counted against memory limits). The mapping serves as an address space
+ * reservation, into which shared memory segment can be extended and is
+ * represented by the second /memfd:main with no permissions.
+ */
+
+PGUsedShmemInfo UsedShmemInfo[NUM_MEMORY_MAPPINGS];
+
+ /*
+ * Structure to hold anonymous shared memory segment properties.
+ */
+typedef struct AnonShmemData
+{
+ int fd; /* fd for the backing anon file */
+ void *addr; /* Pointer to the start of the mapped memory */
+ Size size; /* Size of the mapped memory */
-unsigned long UsedShmemSegID = 0;
-void *UsedShmemSegAddr = NULL;
+} AnonShmemData;
-static Size AnonymousShmemSize;
-static void *AnonymousShmem = NULL;
+AnonShmemData AnonShmemInfo[NUM_MEMORY_MAPPINGS];
+
+/*
+ * Flag telling that we have decided to use huge pages.
+ *
+ * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false)
+ * instead, but it feels like an overkill.
+ */
+static bool huge_pages_on = false;
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
@@ -471,19 +525,20 @@ PGSharedMemoryAttach(IpcMemoryId shmId,
* hugepage sizes, we might want to think about more invasive strategies,
* such as increasing shared_buffers to absorb the extra space.
*
- * Returns the (real, assumed or config provided) page size into
- * *hugepagesize, and the hugepage-related mmap flags to use into
- * *mmap_flags if requested by the caller. If huge pages are not supported,
- * *hugepagesize and *mmap_flags are set to 0.
+ * Returns the (real, assumed or config provided) page size into *hugepagesize,
+ * the hugepage-related mmap and memfd flags to use into *mmap_flags and
+ * *memfd_flags if requested by the caller. If huge pages are not supported,
+ * *hugepagesize, *mmap_flags and *memfd_flags are set to 0.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
#ifdef MAP_HUGETLB
Size default_hugepagesize = 0;
Size hugepagesize_local = 0;
int mmap_flags_local = 0;
+ int memfd_flags_local = 0;
/*
* System-dependent code to find out the default huge page size.
@@ -542,6 +597,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
mmap_flags_local = MAP_HUGETLB;
+ memfd_flags_local = MFD_HUGETLB;
/*
* On recent enough Linux, also include the explicit page size, if
@@ -556,11 +612,22 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
}
#endif
+#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT)
+ if (hugepagesize_local != default_hugepagesize)
+ {
+ int shift = pg_ceil_log2_64(hugepagesize_local);
+
+ memfd_flags_local |= (shift & MFD_HUGE_MASK) << MFD_HUGE_SHIFT;
+ }
+#endif
+
/* assign the results found */
if (mmap_flags)
*mmap_flags = mmap_flags_local;
if (hugepagesize)
*hugepagesize = hugepagesize_local;
+ if (memfd_flags)
+ *memfd_flags = memfd_flags_local;
#else
@@ -568,6 +635,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags)
*hugepagesize = 0;
if (mmap_flags)
*mmap_flags = 0;
+ if (memfd_flags)
+ *memfd_flags = 0;
#endif /* MAP_HUGETLB */
}
@@ -589,84 +658,266 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
return true;
}
+/*
+ * Wrapper around posix_fallocate() to allocate memory for a given shared memory
+ * segment.
+ *
+ * Performs retry on EINTR, and raises error upon failure.
+ */
+static void
+shmem_fallocate(int fd, const char *mapping_name, Size size, int elevel)
+{
+#if defined(HAVE_POSIX_FALLOCATE) && defined(__linux__)
+ int ret;
+
+
+ /*
+ * If there is not enough memory, trying to access a hole in address space
+ * will cause SIGBUS. If supported, avoid that by allocating memory
+ * upfront.
+ *
+ * We still use a traditional EINTR retry loop to handle SIGCONT.
+ * posix_fallocate() doesn't restart automatically, and we don't want this
+ * to fail if you attach a debugger.
+ */
+ do
+ {
+ ret = posix_fallocate(fd, 0, size);
+ } while (ret == EINTR);
+
+ if (ret != 0)
+ {
+ ereport(elevel,
+ (errmsg("segment[%s]: could not allocate space for anonymous file: %s",
+ mapping_name, strerror(ret)),
+ (ret == ENOMEM) ?
+ errhint("This error usually means that PostgreSQL's request "
+ "for a shared memory segment exceeded available memory, "
+ "swap space, or huge pages. To reduce the request size "
+ "(currently %zu bytes), reduce PostgreSQL's shared "
+ "memory usage, perhaps by reducing \"shared_buffers\" or "
+ "\"max_connections\".",
+ size) : 0));
+ }
+#endif
+}
+
+/*
+ * Round up the required amount of memory and the amount of required reserved
+ * address space to the nearest huge page size.
+ */
+static inline void
+round_off_mapping_sizes_for_hugepages(MemoryMappingSizes *mapping, int hugepagesize)
+{
+ if (hugepagesize == 0)
+ return;
+
+ if (mapping->shmem_req_size % hugepagesize != 0)
+ mapping->shmem_req_size += add_size(mapping->shmem_req_size,
+ hugepagesize - (mapping->shmem_req_size % hugepagesize));
+
+ if (mapping->shmem_reserved % hugepagesize != 0)
+ mapping->shmem_reserved = add_size(mapping->shmem_reserved,
+ hugepagesize - (mapping->shmem_reserved % hugepagesize));
+}
+
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * This function will modify mapping size to the actual size of the allocation,
+ * if it ends up allocating a segment that is larger than requested. If needed,
+ * it also rounds up the mapping reserved size to be a multiple of huge page
+ * size.
+ *
+ * Note that we do not fallback from huge pages to regular pages in this
+ * function, this decision was already made in ReserveAnonymousMemory and we
+ * stick to it.
+ *
+ * TODO: Update the prologue to be consistent with the code.
*/
-static void *
-CreateAnonymousSegment(Size *size)
+static void
+CreateAnonymousSegment(int segment_id, MemoryMappingSizes *mapping)
{
- Size allocsize = *size;
void *ptr = MAP_FAILED;
- int mmap_errno = 0;
- int mmap_flags = MAP_SHARED | MAP_ANONYMOUS | MAP_HASSEMAPHORE;
+ int mmap_flags = MAP_SHARED | MAP_HASSEMAPHORE | MAP_NORESERVE;
+ AnonShmemData *anonshmem = &AnonShmemInfo[segment_id];
+ const char *segname = MappingName(segment_id);
+ int memfd_flags = 0;
#ifndef MAP_HUGETLB
- /* PGSharedMemoryCreate should have dealt with this case */
- Assert(huge_pages != HUGE_PAGES_ON);
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
#else
- if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ if (huge_pages_on)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
int huge_mmap_flags;
+ int huge_memfd_flags;
- GetHugePageSize(&hugepagesize, &huge_mmap_flags);
+ /* Make sure nothing is messed up */
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
- if (allocsize % hugepagesize != 0)
- allocsize = add_size(allocsize, hugepagesize - (allocsize % hugepagesize));
+ /* Round up the request size to a suitable large value */
+ GetHugePageSize(&hugepagesize, &huge_mmap_flags, &huge_memfd_flags);
+ round_off_mapping_sizes_for_hugepages(mapping, hugepagesize);
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | huge_mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ /* Verify that the new size is withing the reserved boundaries */
+ Assert(mapping->shmem_reserved >= mapping->shmem_req_size);
+
+ mmap_flags = mmap_flags | huge_mmap_flags;
+ memfd_flags = memfd_flags | huge_memfd_flags;
}
#endif
/*
- * Report whether huge pages are in use. This needs to be tracked before
- * the second mmap() call if attempting to use huge pages failed
- * previously.
+ * Prepare an anonymous file backing the segment. Its size will be
+ * specified later via ftruncate.
+ *
+ * The file behaves like a regular file, but lives in memory. Once all
+ * references to the file are dropped, it is automatically released.
+ * Anonymous memory is used for all backing pages of the file, thus it has
+ * the same semantics as anonymous memory allocations using mmap with the
+ * MAP_ANONYMOUS flag.
+ *
+ * TODO: Need a configuration test for memfd_create.
+ *
+ * TODO: Earlier releases did not use file backed shared memory segments.
+ * By setting bit 1 in /proc/<PID>/coredump_filter, those shared memory
+ * segments could be dumped to the core file. But dumping file backed
+ * shared memory segments requires bit 3 to be set. We need to document
+ * this change in the release notes.
*/
- SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
- PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ anonshmem->fd = memfd_create(segname, memfd_flags);
+ if (anonshmem->fd == -1)
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not create anonymous shared memory file: %m",
+ segname)));
- if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON)
- {
- /*
- * Use the original size, not the rounded-up value, when falling back
- * to non-huge pages.
- */
- allocsize = *size;
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags, -1, 0);
- mmap_errno = errno;
- }
+ elog(DEBUG1, "segment[%s]: mmap(%zu)", segname, mapping->shmem_req_size);
+ /*
+ * Reserve maximum required address space for future expansion of this
+ * memory segment. The whole address space will be setup for read/write
+ * access, so that memory allocated to this address space can be read or
+ * written to even if it is resized in the future using just ftruncate.
+ * MAP_NORESERVE alone should ensure that no memory is allocated. But when
+ * using huge pages, the memory is allocated at mmap time if PROT_WRITE |
+ * PROT_READ is used. Hence we create the mapping with PROT_NONE first and
+ * then use mprotect to set the required permissions.
+ */
+ ptr = mmap(NULL, mapping->shmem_reserved, PROT_NONE,
+ mmap_flags, anonshmem->fd, 0);
if (ptr == MAP_FAILED)
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not map anonymous shared memory: %m",
+ segname)));
+
+ if (mprotect(ptr, mapping->shmem_reserved, PROT_READ | PROT_WRITE) == -1)
+ ereport(FATAL,
+ (errmsg("segment[%s]: could not update anonymous shared memory permissions: %m",
+ segname)));
+
+
+ /*
+ * Resize the backing file to the required size. On platforms where it is
+ * supported, we also allocate the required memory upfront. On other
+ * platform the memory upto the size of file will be allocated on demand.
+ */
+ if (ftruncate(anonshmem->fd, mapping->shmem_req_size) == -1)
{
- errno = mmap_errno;
+ int save_errno = errno;
+
+ close(anonshmem->fd);
+ anonshmem->fd = -1;
+
+ errno = save_errno;
ereport(FATAL,
- (errmsg("could not map anonymous shared memory: %m"),
- (mmap_errno == ENOMEM) ?
+ (errmsg("segment[%s]: could not truncate anonymous file to size %zu: %m",
+ segname, mapping->shmem_req_size),
+ (save_errno == ENOMEM) ?
errhint("This error usually means that PostgreSQL's request "
"for a shared memory segment exceeded available memory, "
"swap space, or huge pages. To reduce the request size "
"(currently %zu bytes), reduce PostgreSQL's shared "
"memory usage, perhaps by reducing \"shared_buffers\" or "
"\"max_connections\".",
- allocsize) : 0));
+ mapping->shmem_req_size) : 0));
}
+ shmem_fallocate(anonshmem->fd, segname, mapping->shmem_req_size, FATAL);
- *size = allocsize;
- return ptr;
+ anonshmem->addr = ptr;
+ anonshmem->size = mapping->shmem_reserved;
+}
+
+/*
+ * PrepareHugePages
+ *
+ * Figure out if there are enough huge pages to allocate all shared memory
+ * segments, and report that information via huge_pages_status and
+ * huge_pages_on. It needs to be called before creating shared memory segments.
+ *
+ * It is necessary to maintain the same semantic (simple on/off) for
+ * huge_pages_status, even if there are multiple shared memory segments: all
+ * segments either use huge pages or not, there is no mix of segments with
+ * different page size. The latter might be actually beneficial, in particular
+ * because only some segments may require large amount of memory, but for now
+ * we go with a simple solution.
+ */
+void
+PrepareHugePages()
+{
+ void *ptr = MAP_FAILED;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+ int mmap_flags = (MAP_SHARED | MAP_HASSEMAPHORE);
+
+ CalculateShmemSize(mapping_sizes);
+
+ /* Complain if hugepages demanded but we can't possibly support them */
+#if !defined(MAP_HUGETLB)
+ if (huge_pages == HUGE_PAGES_ON)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("huge pages not supported on this platform")));
+#else
+ if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
+ {
+ Size hugepagesize,
+ total_size = 0;
+ int huge_mmap_flags;
+
+ GetHugePageSize(&hugepagesize, &huge_mmap_flags, NULL);
+
+ /*
+ * Figure out how much memory is needed for all segments, keeping in
+ * mind that for every segment this value will be rounding up by the
+ * huge page size. The resulting value will be used to probe memory
+ * and decide whether we will allocate huge pages or not.
+ */
+ for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ Size segment_size = mapping_sizes[segment].shmem_req_size;
+
+ if (segment_size % hugepagesize != 0)
+ segment_size += hugepagesize - (segment_size % hugepagesize);
+
+ total_size += segment_size;
+ }
+
+ /* Map total amount of memory to test its availability. */
+ elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB",
+ total_size);
+ ptr = mmap(NULL, total_size, PROT_NONE,
+ mmap_flags | MAP_ANONYMOUS | huge_mmap_flags, -1, 0);
+ }
+#endif
+
+ /*
+ * Report whether huge pages are in use. This needs to be tracked before
+ * creating shared memory segments.
+ */
+ SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+ huge_pages_on = ptr != MAP_FAILED;
}
/*
@@ -676,20 +927,29 @@ CreateAnonymousSegment(Size *size)
static void
AnonymousShmemDetach(int status, Datum arg)
{
- /* Release anonymous shared memory block, if any. */
- if (AnonymousShmem != NULL)
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ AnonShmemData *segment = &AnonShmemInfo[i];
+
+ /* Release anonymous shared memory block, if any. */
+ if (segment->addr != NULL)
+ {
+ Assert(segment->fd != -1);
+
+ if (munmap(segment->addr, segment->size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ segment->addr, segment->size);
+ segment->addr = NULL;
+ close(segment->fd);
+ segment->fd = -1;
+ }
}
}
/*
* PGSharedMemoryCreate
*
- * Create a shared memory segment of the given size and initialize its
+ * Create a shared memory segment for the given mapping and initialize its
* standard header. Also, register an on_shmem_exit callback to release
* the storage.
*
@@ -699,7 +959,7 @@ AnonymousShmemDetach(int status, Datum arg)
* postmaster or backend.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping,
PGShmemHeader **shim)
{
IpcMemoryKey NextShmemSegID;
@@ -707,6 +967,8 @@ PGSharedMemoryCreate(Size size,
PGShmemHeader *hdr;
struct stat statbuf;
Size sysvsize;
+ AnonShmemData *anonshmem = &AnonShmemInfo[segment_id];
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[segment_id];
/*
* We use the data directory's ID info (inode and device numbers) to
@@ -719,14 +981,6 @@ PGSharedMemoryCreate(Size size,
errmsg("could not stat data directory \"%s\": %m",
DataDir)));
- /* Complain if hugepages demanded but we can't possibly support them */
-#if !defined(MAP_HUGETLB)
- if (huge_pages == HUGE_PAGES_ON)
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("huge pages not supported on this platform")));
-#endif
-
/* For now, we don't support huge pages in SysV memory */
if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP)
ereport(ERROR,
@@ -734,12 +988,12 @@ PGSharedMemoryCreate(Size size,
errmsg("huge pages not supported with the current \"shared_memory_type\" setting")));
/* Room for a header? */
- Assert(size > MAXALIGN(sizeof(PGShmemHeader)));
+ Assert(mapping->shmem_req_size > MAXALIGN(sizeof(PGShmemHeader)));
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
- AnonymousShmemSize = size;
+ /* On success, mapping data will be modified. */
+ CreateAnonymousSegment(segment_id, mapping);
/* Register on-exit routine to unmap the anonymous segment */
on_shmem_exit(AnonymousShmemDetach, (Datum) 0);
@@ -749,7 +1003,7 @@ PGSharedMemoryCreate(Size size,
}
else
{
- sysvsize = size;
+ sysvsize = mapping->shmem_req_size;
/* huge pages are only available with mmap */
SetConfigOption("huge_pages_status", "off",
@@ -762,7 +1016,7 @@ PGSharedMemoryCreate(Size size,
* loop simultaneously. (CreateDataDirLockFile() does not entirely ensure
* that, but prefer fixing it over coping here.)
*/
- NextShmemSegID = statbuf.st_ino;
+ NextShmemSegID = statbuf.st_ino + usedShmem->UsedShmemSegID;
for (;;)
{
@@ -800,6 +1054,8 @@ PGSharedMemoryCreate(Size size,
errmsg("pre-existing shared memory block (key %lu, ID %lu) is still in use",
(unsigned long) NextShmemSegID,
(unsigned long) shmid),
+ errdetail("when trying to create shared memory block for segment \"%s\"",
+ MappingName(segment_id)),
errhint("Terminate any old server processes associated with data directory \"%s\".",
DataDir)));
break;
@@ -854,24 +1110,24 @@ PGSharedMemoryCreate(Size size,
/*
* Initialize space allocation status for segment.
*/
- hdr->totalsize = size;
+ hdr->totalsize = mapping->shmem_req_size;
+ hdr->reservedsize = mapping->shmem_reserved;
hdr->content_offset = MAXALIGN(sizeof(PGShmemHeader));
*shim = hdr;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegID = (unsigned long) NextShmemSegID;
+ usedShmem->UsedShmemSegAddr = memAddress;
+ usedShmem->UsedShmemSegID = (unsigned long) NextShmemSegID;
/*
- * If AnonymousShmem is NULL here, then we're not using anonymous shared
- * memory, and should return a pointer to the System V shared memory
- * block. Otherwise, the System V shared memory block is only a shim, and
- * we must return a pointer to the real block.
+ * If we're not using anonymous shared memory, return a pointer to the
+ * System V shared memory block. Otherwise, the System V shared memory
+ * block is only a shim, and we must return a pointer to the real block.
*/
- if (AnonymousShmem == NULL)
+ if (anonshmem->addr == NULL)
return hdr;
- memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader));
- return (PGShmemHeader *) AnonymousShmem;
+ memcpy(anonshmem->addr, hdr, sizeof(PGShmemHeader));
+ return anonshmem->addr;
}
#ifdef EXEC_BACKEND
@@ -884,9 +1140,9 @@ PGSharedMemoryCreate(Size size,
* EXEC_BACKEND case; otherwise postmaster children inherit the shared memory
* segment attachment via fork().
*
- * UsedShmemSegID and UsedShmemSegAddr are implicit parameters to this
- * routine. The caller must have already restored them to the postmaster's
- * values.
+ * Segments array is an implicit parameter to this
+ * routine. The caller must have already restored it to the postmaster's
+ * state.
*/
void
PGSharedMemoryReAttach(void)
@@ -894,32 +1150,42 @@ PGSharedMemoryReAttach(void)
IpcMemoryId shmid;
PGShmemHeader *hdr;
IpcMemoryState state;
- void *origUsedShmemSegAddr = UsedShmemSegAddr;
+ void *origUsedShmemSegAddr;
- Assert(UsedShmemSegAddr != NULL);
- Assert(IsUnderPostmaster);
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[i];
+
+ origUsedShmemSegAddr = usedShmem->UsedShmemSegAddr;
+
+ Assert(usedShmem->UsedShmemSegAddr != NULL);
+ Assert(IsUnderPostmaster);
#ifdef __CYGWIN__
- /* cygipc (currently) appears to not detach on exec. */
- PGSharedMemoryDetach();
- UsedShmemSegAddr = origUsedShmemSegAddr;
+ /* cygipc (currently) appears to not detach on exec. */
+ PGSharedMemoryDetach();
+ usedShmem->UsedShmemSegAddr = origUsedShmemSegAddr;
#endif
- elog(DEBUG3, "attaching to %p", UsedShmemSegAddr);
- shmid = shmget(UsedShmemSegID, sizeof(PGShmemHeader), 0);
- if (shmid < 0)
- state = SHMSTATE_FOREIGN;
- else
- state = PGSharedMemoryAttach(shmid, UsedShmemSegAddr, &hdr);
- if (state != SHMSTATE_ATTACHED)
- elog(FATAL, "could not reattach to shared memory (key=%d, addr=%p): %m",
- (int) UsedShmemSegID, UsedShmemSegAddr);
- if (hdr != origUsedShmemSegAddr)
- elog(FATAL, "reattaching to shared memory returned unexpected address (got %p, expected %p)",
- hdr, origUsedShmemSegAddr);
- dsm_set_control_handle(hdr->dsm_control);
-
- UsedShmemSegAddr = hdr; /* probably redundant */
+ elog(DEBUG3, "attaching to %p", usedShmem->UsedShmemSegAddr);
+ shmid = shmget(usedShmem->UsedShmemSegID, sizeof(PGShmemHeader), 0);
+ if (shmid < 0)
+ state = SHMSTATE_FOREIGN;
+ else
+ state = PGSharedMemoryAttach(shmid, usedShmem->UsedShmemSegAddr, &hdr);
+ if (state != SHMSTATE_ATTACHED)
+ elog(FATAL, "could not reattach to shared memory (key=%d, addr=%p): %m",
+ (int) usedShmem->UsedShmemSegID, usedShmem->UsedShmemSegAddr);
+ if (hdr != origUsedShmemSegAddr)
+ elog(FATAL, "reattaching to shared memory returned unexpected address (got %p, expected %p)",
+ hdr, origUsedShmemSegAddr);
+
+ /* Re-establish dsm_control mapping, if any */
+ if (hdr->dsm_control != 0)
+ dsm_set_control_handle(hdr->dsm_control);
+
+ usedShmem->UsedShmemSegAddr = hdr; /* probably redundant */
+ }
}
/*
@@ -933,14 +1199,13 @@ PGSharedMemoryReAttach(void)
* The child process startup logic might or might not call PGSharedMemoryDetach
* after this; make sure that it will be a no-op if called.
*
- * UsedShmemSegID and UsedShmemSegAddr are implicit parameters to this
- * routine. The caller must have already restored them to the postmaster's
- * values.
+ * Segments array is an implicit parameter to this
+ * routine. The caller must have already restored it to the postmaster's
+ * state.
*/
void
PGSharedMemoryNoReAttach(void)
{
- Assert(UsedShmemSegAddr != NULL);
Assert(IsUnderPostmaster);
#ifdef __CYGWIN__
@@ -948,10 +1213,16 @@ PGSharedMemoryNoReAttach(void)
PGSharedMemoryDetach();
#endif
- /* For cleanliness, reset UsedShmemSegAddr to show we're not attached. */
- UsedShmemSegAddr = NULL;
- /* And the same for UsedShmemSegID. */
- UsedShmemSegID = 0;
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[i];
+
+ Assert(usedShmem->UsedShmemSegAddr != NULL);
+ /* For cleanliness, reset UsedShmemSegAddr to show we're not attached. */
+ usedShmem->UsedShmemSegAddr = NULL;
+ /* And the same for UsedShmemSegID. */
+ usedShmem->UsedShmemSegID = 0;
+ }
}
#endif /* EXEC_BACKEND */
@@ -959,35 +1230,44 @@ PGSharedMemoryNoReAttach(void)
/*
* PGSharedMemoryDetach
*
- * Detach from the shared memory segment, if still attached. This is not
+ * Detach from the shared memory segments, if still attached. This is not
* intended to be called explicitly by the process that originally created the
- * segment (it will have on_shmem_exit callback(s) registered to do that).
+ * segments (it will have on_shmem_exit callback(s) registered to do that).
* Rather, this is for subprocesses that have inherited an attachment and want
* to get rid of it.
*
- * UsedShmemSegID and UsedShmemSegAddr are implicit parameters to this
- * routine, also AnonymousShmem and AnonymousShmemSize.
+ * PGUsedShmemInfo::UsedShmemSegID and PGUsedShmemInfo::UsedShmemSegAddr are
+ * implicit parameters to this routine obtained from entries in UsedShmemInfo
+ * array.
*/
void
PGSharedMemoryDetach(void)
{
- if (UsedShmemSegAddr != NULL)
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- if ((shmdt(UsedShmemSegAddr) < 0)
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[i];
+ AnonShmemData *anonshmem = &AnonShmemInfo[i];
+
+ if (usedShmem->UsedShmemSegAddr != NULL)
+ {
+ if ((shmdt(usedShmem->UsedShmemSegAddr) < 0)
#if defined(EXEC_BACKEND) && defined(__CYGWIN__)
- /* Work-around for cygipc exec bug */
- && shmdt(NULL) < 0
+ /* Work-around for cygipc exec bug */
+ && shmdt(NULL) < 0
#endif
- )
- elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr);
- UsedShmemSegAddr = NULL;
- }
+ )
+ elog(LOG, "shmdt(%p) failed: %m", usedShmem->UsedShmemSegAddr);
+ usedShmem->UsedShmemSegAddr = NULL;
+ }
- if (AnonymousShmem != NULL)
- {
- if (munmap(AnonymousShmem, AnonymousShmemSize) < 0)
- elog(LOG, "munmap(%p, %zu) failed: %m",
- AnonymousShmem, AnonymousShmemSize);
- AnonymousShmem = NULL;
+ if (anonshmem->addr != NULL)
+ {
+ if (munmap(anonshmem->addr, anonshmem->size) < 0)
+ elog(LOG, "munmap(%p, %zu) failed: %m",
+ anonshmem->addr, anonshmem->size);
+ anonshmem->addr = NULL;
+ close(anonshmem->fd);
+ anonshmem->fd = -1;
+ }
}
}
diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c
index ba97c9b2d64..4683736415b 100644
--- a/src/backend/port/win32_sema.c
+++ b/src/backend/port/win32_sema.c
@@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas)
* process exits.
*/
void
-PGReserveSemaphores(int maxSemas)
+PGReserveSemaphores(int maxSemas, int shmem_segment)
{
mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE));
if (mySemSet == NULL)
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 794e4fcb2ad..034460eab96 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -39,15 +39,14 @@
* address space and is negligible relative to the 64-bit address space.
*/
#define PROTECTIVE_REGION_SIZE (10 * WIN32_STACK_RLIMIT)
-void *ShmemProtectiveRegion = NULL;
-
-HANDLE UsedShmemSegID = INVALID_HANDLE_VALUE;
-void *UsedShmemSegAddr = NULL;
-static Size UsedShmemSegSize = 0;
static bool EnableLockPagesPrivilege(int elevel);
static void pgwin32_SharedMemoryDelete(int status, Datum shmId);
+PGUsedShmemInfo UsedShmemInfo[NUM_MEMORY_MAPPINGS];
+
+static Size UsedShmemSegSizes[NUM_MEMORY_MAPPINGS] = {0};
+
/*
* Generate shared memory segment name. Expand the data directory, to generate
* an identifier unique for this data directory. Then replace all backslashes
@@ -202,9 +201,11 @@ EnableLockPagesPrivilege(int elevel)
*
* Create a shared memory segment of the given size and initialize its
* standard header.
+ *
+ * TODO: Check that the segment_id is a valid one before indexing corresponding arrays.
*/
-PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+void
+PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping_sizes,
PGShmemHeader **shim)
{
void *memAddress;
@@ -216,13 +217,14 @@ PGSharedMemoryCreate(Size size,
DWORD size_high;
DWORD size_low;
SIZE_T largePageSize = 0;
- Size orig_size = size;
+ Size size = mapping_sizes->shmem_req_size;
DWORD flProtect = PAGE_READWRITE;
DWORD desiredAccess;
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[segment_id];
- ShmemProtectiveRegion = VirtualAlloc(NULL, PROTECTIVE_REGION_SIZE,
- MEM_RESERVE, PAGE_NOACCESS);
- if (ShmemProtectiveRegion == NULL)
+ usedShmem->ShmemProtectiveRegion = VirtualAlloc(NULL, PROTECTIVE_REGION_SIZE,
+ MEM_RESERVE, PAGE_NOACCESS);
+ if (usedShmem->ShmemProtectiveRegion == NULL)
elog(FATAL, "could not reserve memory region: error code %lu",
GetLastError());
@@ -231,8 +233,12 @@ PGSharedMemoryCreate(Size size,
szShareMem = GetSharedMemName();
- UsedShmemSegAddr = NULL;
+ usedShmem->UsedShmemSegAddr = NULL;
+ /*
+ * TODO: We don't need to perform this as many times as the number of
+ * segments. Instead do something similar to sysv_shmem.c
+ */
if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
{
/* Does the processor support large pages? */
@@ -304,7 +310,7 @@ retry:
* Use the original size, not the rounded-up value, when
* falling back to non-huge pages.
*/
- size = orig_size;
+ size = mapping_sizes->shmem_req_size;
flProtect = PAGE_READWRITE;
goto retry;
}
@@ -337,6 +343,8 @@ retry:
if (!hmap)
ereport(FATAL,
(errmsg("pre-existing shared memory block is still in use"),
+ errdetail("when trying to create shared memory block for segment \"%s\"",
+ PGShmemSegmentName(segment)),
errhint("Check if there are any old server processes still running, and terminate them.")));
free(szShareMem);
@@ -393,9 +401,9 @@ retry:
hdr->dsm_control = 0;
/* Save info for possible future use */
- UsedShmemSegAddr = memAddress;
- UsedShmemSegSize = size;
- UsedShmemSegID = hmap2;
+ usedShmem->UsedShmemSegAddr = memAddress;
+ UsedShmemSegSizes[segment_id] = size;
+ usedShmem->UsedShmemSegID = (unsigned long) hmap2;
/* Register on-exit routine to delete the new segment */
on_shmem_exit(pgwin32_SharedMemoryDelete, PointerGetDatum(hmap2));
@@ -405,8 +413,6 @@ retry:
/* Report whether huge pages are in use */
SetConfigOption("huge_pages_status", (flProtect & SEC_LARGE_PAGES) ?
"on" : "off", PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
-
- return hdr;
}
/*
@@ -416,42 +422,52 @@ retry:
* an already existing shared memory segment, using the handle inherited from
* the postmaster.
*
- * ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr are implicit
- * parameters to this routine. The caller must have already restored them to
- * the postmaster's values.
+ * Segments is an implicit parameters to this routine. The caller must have
+ * already restored ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr
+ * in each Segment to the postmaster's values.
*/
void
PGSharedMemoryReAttach(void)
{
PGShmemHeader *hdr;
- void *origUsedShmemSegAddr = UsedShmemSegAddr;
+ void *origUsedShmemSegAddr;
- Assert(ShmemProtectiveRegion != NULL);
- Assert(UsedShmemSegAddr != NULL);
Assert(IsUnderPostmaster);
- /*
- * Release memory region reservations made by the postmaster
- */
- if (VirtualFree(ShmemProtectiveRegion, 0, MEM_RELEASE) == 0)
- elog(FATAL, "failed to release reserved memory region (addr=%p): error code %lu",
- ShmemProtectiveRegion, GetLastError());
- if (VirtualFree(UsedShmemSegAddr, 0, MEM_RELEASE) == 0)
- elog(FATAL, "failed to release reserved memory region (addr=%p): error code %lu",
- UsedShmemSegAddr, GetLastError());
-
- hdr = (PGShmemHeader *) MapViewOfFileEx(UsedShmemSegID, FILE_MAP_READ | FILE_MAP_WRITE, 0, 0, 0, UsedShmemSegAddr);
- if (!hdr)
- elog(FATAL, "could not reattach to shared memory (key=%p, addr=%p): error code %lu",
- UsedShmemSegID, UsedShmemSegAddr, GetLastError());
- if (hdr != origUsedShmemSegAddr)
- elog(FATAL, "reattaching to shared memory returned unexpected address (got %p, expected %p)",
- hdr, origUsedShmemSegAddr);
- if (hdr->magic != PGShmemMagic)
- elog(FATAL, "reattaching to shared memory returned non-PostgreSQL memory");
- dsm_set_control_handle(hdr->dsm_control);
-
- UsedShmemSegAddr = hdr; /* probably redundant */
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[i];
+
+ Assert(usedShmem->ShmemProtectiveRegion != NULL);
+ Assert(usedShmem->UsedShmemSegAddr != NULL);
+
+ origUsedShmemSegAddr = usedShmem->UsedShmemSegAddr;
+
+ /*
+ * Release memory region reservations made by the postmaster
+ */
+ if (VirtualFree(usedShmem->ShmemProtectiveRegion, 0, MEM_RELEASE) == 0)
+ elog(FATAL, "failed to release reserved memory region (addr=%p): error code %lu",
+ usedShmem->ShmemProtectiveRegion, GetLastError());
+ if (VirtualFree(usedShmem->UsedShmemSegAddr, 0, MEM_RELEASE) == 0)
+ elog(FATAL, "failed to release reserved memory region (addr=%p): error code %lu",
+ usedShmem->UsedShmemSegAddr, GetLastError());
+
+ hdr = (PGShmemHeader *) MapViewOfFileEx(usedShmem->UsedShmemSegID, FILE_MAP_READ | FILE_MAP_WRITE, 0, 0, 0, usedShmem->UsedShmemSegAddr);
+ if (!hdr)
+ elog(FATAL, "could not reattach to shared memory (key=%p, addr=%p): error code %lu",
+ usedShmem->UsedShmemSegID, usedShmem->UsedShmemSegAddr, GetLastError());
+ if (hdr != origUsedShmemSegAddr)
+ elog(FATAL, "reattaching to shared memory returned unexpected address (got %p, expected %p)",
+ hdr, origUsedShmemSegAddr);
+ if (hdr->magic != PGShmemMagic)
+ elog(FATAL, "reattaching to shared memory returned non-PostgreSQL memory");
+ /* Re-establish dsm_control mapping, if any */
+ if (hdr->dsm_control != 0)
+ dsm_set_control_handle(hdr->dsm_control);
+
+ usedShmem->UsedShmemSegAddr = hdr; /* probably redundant */
+ }
}
/*
@@ -464,22 +480,28 @@ PGSharedMemoryReAttach(void)
* The child process startup logic might or might not call PGSharedMemoryDetach
* after this; make sure that it will be a no-op if called.
*
- * ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr are implicit
- * parameters to this routine. The caller must have already restored them to
- * the postmaster's values.
+ * Segments is an implicit parameters to this routine. The caller must have
+ * already restored ShmemProtectiveRegion and UsedShmemSegAddr
+ * in each Segment to the postmaster's values.
*/
void
PGSharedMemoryNoReAttach(void)
{
- Assert(ShmemProtectiveRegion != NULL);
- Assert(UsedShmemSegAddr != NULL);
Assert(IsUnderPostmaster);
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[i];
- /*
- * Under Windows we will not have mapped the segment, so we don't need to
- * un-map it. Just reset UsedShmemSegAddr to show we're not attached.
- */
- UsedShmemSegAddr = NULL;
+ Assert(usedShmem->ShmemProtectiveRegion != NULL);
+ Assert(usedShmem->UsedShmemSegAddr != NULL);
+
+ /*
+ * Under Windows we will not have mapped the segment, so we don't need
+ * to un-map it. Just reset UsedShmemSegAddr to show we're not
+ * attached.
+ */
+ usedShmem->UsedShmemSegAddr = NULL;
+ }
/*
* We *must* close the inherited shmem segment handle, else Windows will
@@ -492,49 +514,55 @@ PGSharedMemoryNoReAttach(void)
/*
* PGSharedMemoryDetach
*
- * Detach from the shared memory segment, if still attached. This is not
+ * Detach from the shared memory segments, if still attached. This is not
* intended to be called explicitly by the process that originally created the
- * segment (it will have an on_shmem_exit callback registered to do that).
- * Rather, this is for subprocesses that have inherited an attachment and want
- * to get rid of it.
+ * segments (it will have an on_shmem_exit callback registered to do that).
+ * Rather, this is for subprocesses that have inherited an attachment and want to
+ * get rid of it.
*
- * ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr are implicit
- * parameters to this routine.
+ * UsedShmemInfo is an implicit parameters to this routine. The caller must have
+ * already restored ShmemProtectiveRegion, UsedShmemSegID and UsedShmemSegAddr in
+ * each Segment to the postmaster's values.
*/
void
PGSharedMemoryDetach(void)
{
- /*
- * Releasing the protective region liberates an unimportant quantity of
- * address space, but be tidy.
- */
- if (ShmemProtectiveRegion != NULL)
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- if (VirtualFree(ShmemProtectiveRegion, 0, MEM_RELEASE) == 0)
- elog(LOG, "failed to release reserved memory region (addr=%p): error code %lu",
- ShmemProtectiveRegion, GetLastError());
+ PGUsedShmemInfo *segment = &UsedShmemInfo[i];
- ShmemProtectiveRegion = NULL;
- }
+ /*
+ * Releasing the protective region liberates an unimportant quantity
+ * of address space, but be tidy.
+ */
+ if (segment->ShmemProtectiveRegion != NULL)
+ {
+ if (VirtualFree(segment->ShmemProtectiveRegion, 0, MEM_RELEASE) == 0)
+ elog(LOG, "failed to release reserved memory region (addr=%p): error code %lu",
+ segment->ShmemProtectiveRegion, GetLastError());
- /* Unmap the view, if it's mapped */
- if (UsedShmemSegAddr != NULL)
- {
- if (!UnmapViewOfFile(UsedShmemSegAddr))
- elog(LOG, "could not unmap view of shared memory: error code %lu",
- GetLastError());
+ segment->ShmemProtectiveRegion = NULL;
+ }
- UsedShmemSegAddr = NULL;
- }
+ /* Unmap the view, if it's mapped */
+ if (segment->UsedShmemSegAddr != NULL)
+ {
+ if (!UnmapViewOfFile(segment->UsedShmemSegAddr))
+ elog(LOG, "could not unmap view of shared memory: error code %lu",
+ GetLastError());
- /* And close the shmem handle, if we have one */
- if (UsedShmemSegID != INVALID_HANDLE_VALUE)
- {
- if (!CloseHandle(UsedShmemSegID))
- elog(LOG, "could not close handle to shared memory: error code %lu",
- GetLastError());
+ segment->UsedShmemSegAddr = NULL;
+ }
- UsedShmemSegID = INVALID_HANDLE_VALUE;
+ /* And close the shmem handle, if we have one */
+ if (segment->UsedShmemSegID != INVALID_HANDLE_VALUE)
+ {
+ if (!CloseHandle(segment->UsedShmemSegID))
+ elog(LOG, "could not close handle to shared memory: error code %lu",
+ GetLastError());
+
+ segment->UsedShmemSegID = INVALID_HANDLE_VALUE;
+ }
}
}
@@ -574,50 +602,55 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
{
void *address;
- Assert(ShmemProtectiveRegion != NULL);
- Assert(UsedShmemSegAddr != NULL);
- Assert(UsedShmemSegSize != 0);
-
- /* ShmemProtectiveRegion */
- address = VirtualAllocEx(hChild, ShmemProtectiveRegion,
- PROTECTIVE_REGION_SIZE,
- MEM_RESERVE, PAGE_NOACCESS);
- if (address == NULL)
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
{
- /* Don't use FATAL since we're running in the postmaster */
- elog(LOG, "could not reserve shared memory region (addr=%p) for child %p: error code %lu",
- ShmemProtectiveRegion, hChild, GetLastError());
- return false;
- }
- if (address != ShmemProtectiveRegion)
- {
- /*
- * Should never happen - in theory if allocation granularity causes
- * strange effects it could, so check just in case.
- *
- * Don't use FATAL since we're running in the postmaster.
- */
- elog(LOG, "reserved shared memory region got incorrect address %p, expected %p",
- address, ShmemProtectiveRegion);
- return false;
- }
+ PGUsedShmemInfo *segment = &UsedShmemInfo[i];
- /* UsedShmemSegAddr */
- address = VirtualAllocEx(hChild, UsedShmemSegAddr, UsedShmemSegSize,
- MEM_RESERVE, PAGE_READWRITE);
- if (address == NULL)
- {
- elog(LOG, "could not reserve shared memory region (addr=%p) for child %p: error code %lu",
- UsedShmemSegAddr, hChild, GetLastError());
- return false;
- }
- if (address != UsedShmemSegAddr)
- {
- elog(LOG, "reserved shared memory region got incorrect address %p, expected %p",
- address, UsedShmemSegAddr);
- return false;
- }
+ Assert(segment->ShmemProtectiveRegion != NULL);
+ Assert(segment->UsedShmemSegAddr != NULL);
+ Assert(UsedShmemSegSizes[i] != 0);
+
+ /* ShmemProtectiveRegion */
+ address = VirtualAllocEx(hChild, segment->ShmemProtectiveRegion,
+ PROTECTIVE_REGION_SIZE,
+ MEM_RESERVE, PAGE_NOACCESS);
+ if (address == NULL)
+ {
+ /* Don't use FATAL since we're running in the postmaster */
+ elog(LOG, "could not reserve shared memory region (addr=%p) for child %p: error code %lu",
+ segment->ShmemProtectiveRegion, hChild, GetLastError());
+ return false;
+ }
+ if (address != segment->ShmemProtectiveRegion)
+ {
+ /*
+ * Should never happen - in theory if allocation granularity
+ * causes strange effects it could, so check just in case.
+ *
+ * Don't use FATAL since we're running in the postmaster.
+ */
+ elog(LOG, "reserved shared memory region got incorrect address %p, expected %p",
+ address, segment->ShmemProtectiveRegion);
+ return false;
+ }
+
+ /* UsedShmemSegAddr */
+ address = VirtualAllocEx(hChild, segment->UsedShmemSegAddr, UsedShmemSegSizes[i],
+ MEM_RESERVE, PAGE_READWRITE);
+ if (address == NULL)
+ {
+ elog(LOG, "could not reserve shared memory region (addr=%p) for child %p: error code %lu",
+ segment->UsedShmemSegAddr, hChild, GetLastError());
+ return false;
+ }
+ if (address != segment->UsedShmemSegAddr)
+ {
+ elog(LOG, "reserved shared memory region got incorrect address %p, expected %p",
+ address, segment->UsedShmemSegAddr);
+ return false;
+ }
+ }
return true;
}
@@ -627,7 +660,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild)
* use GetLargePageMinimum() instead.
*/
void
-GetHugePageSize(Size *hugepagesize, int *mmap_flags)
+GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags)
{
if (hugepagesize)
*hugepagesize = 0;
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index 926fd6f2700..b58ae118af1 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -89,13 +89,7 @@ typedef int InheritableSocket;
typedef struct
{
char DataDir[MAXPGPATH];
-#ifndef WIN32
- unsigned long UsedShmemSegID;
-#else
- void *ShmemProtectiveRegion;
- HANDLE UsedShmemSegID;
-#endif
- void *UsedShmemSegAddr;
+ PGUsedShmemInfo UsedShmemInfo[NUM_MEMORY_MAPPINGS];
#ifdef USE_INJECTION_POINTS
struct InjectionPointsCtl *ActiveInjectionPoints;
#endif
@@ -677,8 +671,13 @@ SubPostmasterMain(int argc, char *argv[])
process_shared_preload_libraries();
/* Restore basic shared memory pointers */
- if (UsedShmemSegAddr != NULL)
- InitShmemAllocator(UsedShmemSegAddr);
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[i];
+
+ if (usedShmem->UsedShmemSegAddr != NULL)
+ InitShmemAllocator(i, usedShmem->UsedShmemSegAddr);
+ }
/*
* Run the appropriate Main function
@@ -719,12 +718,7 @@ save_backend_variables(BackendParameters *param,
strlcpy(param->DataDir, DataDir, MAXPGPATH);
param->MyPMChildSlot = child_slot;
-
-#ifdef WIN32
- param->ShmemProtectiveRegion = ShmemProtectiveRegion;
-#endif
- param->UsedShmemSegID = UsedShmemSegID;
- param->UsedShmemSegAddr = UsedShmemSegAddr;
+ memcpy(param->UsedShmemInfo, UsedShmemInfo, sizeof(UsedShmemInfo));
#ifdef USE_INJECTION_POINTS
param->ActiveInjectionPoints = ActiveInjectionPoints;
@@ -979,12 +973,7 @@ restore_backend_variables(BackendParameters *param)
SetDataDir(param->DataDir);
MyPMChildSlot = param->MyPMChildSlot;
-
-#ifdef WIN32
- ShmemProtectiveRegion = param->ShmemProtectiveRegion;
-#endif
- UsedShmemSegID = param->UsedShmemSegID;
- UsedShmemSegAddr = param->UsedShmemSegAddr;
+ memcpy(UsedShmemInfo, param->UsedShmemInfo, sizeof(UsedShmemInfo));
#ifdef USE_INJECTION_POINTS
ActiveInjectionPoints = param->ActiveInjectionPoints;
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index c0c223b2e32..42112109af9 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proclist.h"
BufferDescPadded *BufferDescriptors;
@@ -56,6 +57,10 @@ CkptSortItem *CkptBufferIds;
* Pins must be released before end of transaction. For efficiency the
* shared refcount isn't increased if an individual backend pins a buffer
* multiple times. Check the PrivateRefCount infrastructure in bufmgr.c.
+ *
+ * All the data structures except the buffer blocks are allocated in the main
+ * shared memory segment. The buffer blocks are allocated in a separate segment
+ * to allow dynamic resizing of the buffer pool.
*/
@@ -75,22 +80,22 @@ BufferManagerShmemInit(void)
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
- ShmemInitStruct("Buffer Descriptors",
- NBuffers * sizeof(BufferDescPadded),
- &foundDescs);
+ ShmemInitStructInSegment("Buffer Descriptors",
+ NBuffers * sizeof(BufferDescPadded),
+ &foundDescs, MAIN_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
- ShmemInitStruct("Buffer Blocks",
- NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
- &foundBufs));
+ ShmemInitStructInSegment("Buffer Blocks",
+ NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
- ShmemInitStruct("Buffer IO Condition Variables",
- NBuffers * sizeof(ConditionVariableMinimallyPadded),
- &foundIOCV);
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &foundIOCV, MAIN_SHMEM_SEGMENT);
/*
* The array used to sort to-be-checkpointed buffer ids is located in
@@ -100,8 +105,9 @@ BufferManagerShmemInit(void)
* painful.
*/
CkptBufferIds = (CkptSortItem *)
- ShmemInitStruct("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt);
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ MAIN_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
@@ -147,21 +153,28 @@ BufferManagerShmemInit(void)
*
* compute the size of shared memory for the buffer pool including
* data pages, buffer descriptors, hash tables, etc.
+ *
+ * The function adds the amount of required memory for buffer blocks to
+ * BUFFERS_SHMEM_SEGMENT segment. Amount of memory required for other structures
+ * is returned.
*/
Size
-BufferManagerShmemSize(void)
+BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
{
- Size size = 0;
+ size_t size;
+
+ /* size of data pages, plus alignment padding */
+ size = add_size(0, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_reserved = size;
+ size = 0;
/* size of buffer descriptors */
size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
/* to allow aligning buffer descriptors */
size = add_size(size, PG_CACHE_LINE_SIZE);
- /* size of data pages, plus alignment padding */
- size = add_size(size, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
-
/* size of stuff controlled by freelist.c */
size = add_size(size, StrategyShmemSize());
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 5089c7322f3..a33786a460b 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -25,6 +25,7 @@
#include "funcapi.h"
#include "storage/buf_internals.h"
#include "storage/lwlock.h"
+#include "storage/pg_shmem.h"
#include "utils/rel.h"
#include "utils/builtins.h"
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index b7687836188..403890055be 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -19,6 +19,7 @@
#include "port/atomics.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var))))
@@ -418,9 +419,9 @@ StrategyInitialize(bool init)
* Get or create the shared strategy control block
*/
StrategyControl = (BufferStrategyControl *)
- ShmemInitStruct("Buffer Strategy Status",
- sizeof(BufferStrategyControl),
- &found);
+ ShmemInitStructInSegment("Buffer Strategy Status",
+ sizeof(BufferStrategyControl),
+ &found, MAIN_SHMEM_SEGMENT);
if (!found)
{
diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c
index cb944edd8df..4af7782795e 100644
--- a/src/backend/storage/ipc/ipc.c
+++ b/src/backend/storage/ipc/ipc.c
@@ -62,6 +62,8 @@ static void proc_exit_prepare(int code);
* but provide some additional features we need --- in particular,
* we want to register callbacks to invoke when we are disconnecting
* from a broken shared-memory context but not exiting the postmaster.
+ * Maximum number of such exit callbacks depends on the number of shared
+ * segments.
*
* Callback functions can take zero, one, or two args: the first passed
* arg is the integer exitcode, the second is the Datum supplied when
@@ -69,7 +71,7 @@ static void proc_exit_prepare(int code);
* ----------------------------------------------------------------
*/
-#define MAX_ON_EXITS 20
+#define MAX_ON_EXITS 40
struct ONEXIT
{
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 1f7e933d500..6d2c4520c43 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -50,6 +50,7 @@
#include "storage/procarray.h"
#include "storage/procsignal.h"
#include "storage/sinvaladt.h"
+#include "utils/builtins.h"
#include "utils/guc.h"
#include "utils/injection_point.h"
@@ -81,10 +82,17 @@ RequestAddinShmemSpace(Size size)
/*
* CalculateShmemSize
- * Calculates the amount of shared memory needed.
+ * Calculates the amount of shared memory needed.
+ *
+ * The amount of shared memory required per segment is saved in mapping_sizes,
+ * which is expected to be an array of size NUM_MEMORY_MAPPINGS. The total
+ * amount of memory needed across all the segments is returned. For the memory
+ * mappings which reserve address space for future expansion, the required
+ * amount of reserved space is saved in mapping_sizes of those segments.
+ * This memory is not included in the returned value.
*/
Size
-CalculateShmemSize(void)
+CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
{
Size size;
@@ -102,7 +110,13 @@ CalculateShmemSize(void)
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
size = add_size(size, DSMRegistryShmemSize());
- size = add_size(size, BufferManagerShmemSize());
+
+ /*
+ * Buffer manager adds estimates for memory requirements for every shared
+ * memory segment that it uses in the corresponding AnonymousMappings.
+ * Consider size required from only the main shared memory segment here.
+ */
+ size = add_size(size, BufferManagerShmemSize(mapping_sizes));
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
size = add_size(size, ProcGlobalShmemSize());
@@ -145,8 +159,22 @@ CalculateShmemSize(void)
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
+ /*
+ * All the shared memory allocations considered so far happen in the main
+ * shared memory segment.
+ */
+ mapping_sizes[MAIN_SHMEM_SEGMENT].shmem_req_size = size;
+ mapping_sizes[MAIN_SHMEM_SEGMENT].shmem_reserved = size;
+
+ size = 0;
/* might as well round it off to a multiple of a typical page size */
- size = add_size(size, 8192 - (size % 8192));
+ for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ mapping_sizes[segment].shmem_req_size = add_size(mapping_sizes[segment].shmem_req_size, 8192 - (mapping_sizes[segment].shmem_req_size % 8192));
+ mapping_sizes[segment].shmem_reserved = add_size(mapping_sizes[segment].shmem_reserved, 8192 - (mapping_sizes[segment].shmem_reserved % 8192));
+ /* Compute the total size of all segments */
+ size = size + mapping_sizes[segment].shmem_req_size;
+ }
return size;
}
@@ -185,25 +213,21 @@ AttachSharedMemoryStructs(void)
/*
* CreateSharedMemoryAndSemaphores
- * Creates and initializes shared memory and semaphores.
+ * Creates shared memory segments and initializes shared memory structures
+ * and semaphores.
*/
void
CreateSharedMemoryAndSemaphores(void)
{
- PGShmemHeader *shim;
- PGShmemHeader *seghdr;
- Size size;
+ PGShmemHeader *main_seg_shim = NULL;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
Assert(!IsUnderPostmaster);
- /* Compute the size of the shared-memory block */
- size = CalculateShmemSize();
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+ CalculateShmemSize(mapping_sizes);
- /*
- * Create the shmem segment
- */
- seghdr = PGSharedMemoryCreate(size, &shim);
+ /* Decide if we use huge pages or regular size pages */
+ PrepareHugePages();
/*
* Make sure that huge pages are never reported as "unknown" while the
@@ -212,16 +236,42 @@ CreateSharedMemoryAndSemaphores(void)
Assert(strcmp("unknown",
GetConfigOption("huge_pages_status", false, false)) != 0);
- /*
- * Set up shared memory allocation mechanism
- */
- InitShmemAllocator(seghdr);
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ MemoryMappingSizes *mapping = &mapping_sizes[i];
+ PGUsedShmemInfo *usedShmem = &UsedShmemInfo[i];
+ PGShmemHeader *shim;
+ PGShmemHeader *seghdr;
+
+ /*
+ * Set seed shmem identifier which will be changed to the final one
+ * when creating the shared memory segment.
+ */
+ usedShmem->UsedShmemSegID = i;
+
+ /* Compute the size of the shared-memory block */
+ elog(DEBUG3, "invoking IpcMemoryCreate(segment %s, size=%zu, reserved address space=%zu)",
+ MappingName(i), mapping->shmem_req_size, mapping->shmem_reserved);
+
+ /*
+ * Create the shmem segment.
+ */
+ seghdr = PGSharedMemoryCreate(i, mapping, &shim);
+
+ /*
+ * Set up shared memory allocation mechanism
+ */
+ InitShmemAllocator(i, seghdr);
+
+ if (i == MAIN_SHMEM_SEGMENT)
+ main_seg_shim = shim;
+ }
/* Initialize subsystems */
CreateOrAttachShmemStructs();
/* Initialize dynamic shared memory facilities. */
- dsm_postmaster_startup(shim);
+ dsm_postmaster_startup(main_seg_shim);
/*
* Now give loadable modules a chance to set up their shmem allocations
@@ -334,7 +384,9 @@ CreateOrAttachShmemStructs(void)
* InitializeShmemGUCs
*
* This function initializes runtime-computed GUCs related to the amount of
- * shared memory required for the current configuration.
+ * shared memory required for the current configuration. It assumes that the
+ * memory required by the shared memory segments is already calculated and is
+ * available in AnonymousMappings.
*/
void
InitializeShmemGUCs(void)
@@ -343,11 +395,13 @@ InitializeShmemGUCs(void)
Size size_b;
Size size_mb;
Size hp_size;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+
/*
* Calculate the shared memory size and round up to the nearest megabyte.
*/
- size_b = CalculateShmemSize();
+ size_b = CalculateShmemSize(mapping_sizes);
size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
SetConfigOption("shared_memory_size", buf,
@@ -356,7 +410,7 @@ InitializeShmemGUCs(void)
/*
* Calculate the number of huge pages required.
*/
- GetHugePageSize(&hp_size, NULL);
+ GetHugePageSize(&hp_size, NULL, NULL);
if (hp_size != 0)
{
Size hp_required;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9f362ce8641..5cb59d97871 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -63,6 +63,7 @@
* unnecessary.
*/
+
#include "postgres.h"
#include "common/int.h"
@@ -93,17 +94,28 @@ typedef struct ShmemAllocatorData
slock_t shmem_lock;
} ShmemAllocatorData;
-static void *ShmemAllocRaw(Size size, Size *allocated_size);
+/* Structure managing one shared memory segment. */
+typedef struct ShmemSegment
+{
+ PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
+ ShmemAllocatorData *ShmemAllocator;
+ void *ShmemBase; /* start address of shared memory */
+ const char *ShmemSegmentName; /* name of the segment for logging */
+} ShmemSegment;
+
+ShmemSegment Segments[NUM_MEMORY_MAPPINGS];
-/* shared memory global variables */
+static void *ShmemAllocRaw(ShmemSegment *segment, Size size, Size *allocated_size);
-static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
-static void *ShmemBase; /* start address of shared memory */
-static void *ShmemEnd; /* end+1 address of shared memory */
+/* Expose ShmemLock from the main segment for allocating LWLock tranches. */
+slock_t *ShmemLock;
-static ShmemAllocatorData *ShmemAllocator;
-slock_t *ShmemLock; /* points to ShmemAllocator->shmem_lock */
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+/*
+ * Primary index hashtable for shmem, for simplicity we use a single for all
+ * shared memory segments. There can be performance consequences of that, and
+ * an alternative option would be to have one index per shared memory segments.
+ */
+static HTAB *ShmemIndex = NULL;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -111,7 +123,7 @@ static bool firstNumaTouch = true;
Datum pg_numa_available(PG_FUNCTION_ARGS);
/*
- * InitShmemAllocator() --- set up basic pointers to shared memory.
+ * InitShmemAllocator() --- set up basic pointers to shared memory in the given segment.
*
* Called at postmaster or stand-alone backend startup, to initialize the
* allocator's data structure in the shared memory segment. In EXEC_BACKEND,
@@ -119,9 +131,13 @@ Datum pg_numa_available(PG_FUNCTION_ARGS);
* memory areas.
*/
void
-InitShmemAllocator(PGShmemHeader *seghdr)
+InitShmemAllocator(int segment_id, PGShmemHeader *seghdr)
{
+ ShmemSegment *segment;
+
Assert(seghdr != NULL);
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+ segment = &Segments[segment_id];
/*
* We assume the pointer and offset are MAXALIGN. Not a hard requirement,
@@ -130,23 +146,24 @@ InitShmemAllocator(PGShmemHeader *seghdr)
Assert(seghdr == (void *) MAXALIGN(seghdr));
Assert(seghdr->content_offset == MAXALIGN(seghdr->content_offset));
- ShmemSegHdr = seghdr;
- ShmemBase = seghdr;
- ShmemEnd = (char *) ShmemBase + seghdr->totalsize;
+ segment->ShmemSegHdr = seghdr;
+ segment->ShmemBase = seghdr;
+ segment->ShmemSegmentName = MappingName(segment_id);
#ifndef EXEC_BACKEND
Assert(!IsUnderPostmaster);
#endif
if (IsUnderPostmaster)
{
- PGShmemHeader *shmhdr = ShmemSegHdr;
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+
+ segment->ShmemAllocator = (ShmemAllocatorData *) ((char *) shmhdr + shmhdr->content_offset);
- ShmemAllocator = (ShmemAllocatorData *) ((char *) shmhdr + shmhdr->content_offset);
- ShmemLock = &ShmemAllocator->shmem_lock;
}
else
{
Size offset;
+ ShmemAllocatorData *ShmemAllocator;
/*
* Allocations after this point should go through ShmemAlloc, which
@@ -163,47 +180,68 @@ InitShmemAllocator(PGShmemHeader *seghdr)
ShmemAllocator = (ShmemAllocatorData *) ((char *) seghdr + seghdr->content_offset);
SpinLockInit(&ShmemAllocator->shmem_lock);
- ShmemLock = &ShmemAllocator->shmem_lock;
ShmemAllocator->free_offset = offset;
/* ShmemIndex can't be set up yet (need LWLocks first) */
ShmemAllocator->index = NULL;
+
+ segment->ShmemAllocator = ShmemAllocator;
ShmemIndex = (HTAB *) NULL;
}
+
+ /* Expose ShmemLock from the main segment for allocating LWLock tranches. */
+ if (segment_id == MAIN_SHMEM_SEGMENT)
+ ShmemLock = &segment->ShmemAllocator->shmem_lock;
}
/*
- * ShmemAlloc -- allocate max-aligned chunk from shared memory
+ * ShmemAlloc --
+ * allocate max-aligned chunk from given shared memory segment
*
* Throws error if request cannot be satisfied.
*
- * Assumes ShmemLock and ShmemSegHdr are initialized.
+ * Assumes ShmemLock and ShmemSegHdr in the given segment are initialized.
*/
-void *
-ShmemAlloc(Size size)
+
+static void *
+ShmemAllocInternal(ShmemSegment *segment, Size size)
{
void *newSpace;
Size allocated_size;
- newSpace = ShmemAllocRaw(size, &allocated_size);
+ newSpace = ShmemAllocRaw(segment, size, &allocated_size);
if (!newSpace)
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of shared memory (%zu bytes requested)",
- size)));
+ errmsg("out of shared memory in segment %s (%zu bytes requested)",
+ segment->ShmemSegmentName, size)));
return newSpace;
}
+void *
+ShmemAlloc(int segment_id, Size size)
+{
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+
+ return ShmemAllocInternal(&Segments[segment_id], size);
+}
+
/*
* ShmemAllocNoError -- allocate max-aligned chunk from shared memory
*
* As ShmemAlloc, but returns NULL if out of space, rather than erroring.
+ *
+ * This is used as a memory allocation callback for hash tables created using
+ * dynahash.c APIs. It's a bit of work to make the callback specify the segment
+ * where to allocate the memory. For now, there is not need to create shared
+ * memory hash tables in shared memory segments other than main memory segment.
+ * Hence we do not support segment_id parameter here.
*/
void *
ShmemAllocNoError(Size size)
{
Size allocated_size;
- return ShmemAllocRaw(size, &allocated_size);
+ return ShmemAllocRaw(&Segments[MAIN_SHMEM_SEGMENT], size, &allocated_size);
}
/*
@@ -213,11 +251,13 @@ ShmemAllocNoError(Size size)
* be equal to the number requested plus any padding we choose to add.
*/
static void *
-ShmemAllocRaw(Size size, Size *allocated_size)
+ShmemAllocRaw(ShmemSegment *segment, Size size, Size *allocated_size)
{
Size newStart;
Size newFree;
void *newSpace;
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+ ShmemAllocatorData *ShmemAllocator = segment->ShmemAllocator;
/*
* Ensure all space is adequately aligned. We used to only MAXALIGN this
@@ -233,22 +273,21 @@ ShmemAllocRaw(Size size, Size *allocated_size)
size = CACHELINEALIGN(size);
*allocated_size = size;
- Assert(ShmemSegHdr != NULL);
+ Assert(shmhdr != NULL);
- SpinLockAcquire(ShmemLock);
+ SpinLockAcquire(&ShmemAllocator->shmem_lock);
newStart = ShmemAllocator->free_offset;
-
newFree = newStart + size;
- if (newFree <= ShmemSegHdr->totalsize)
+ if (newFree <= shmhdr->totalsize)
{
- newSpace = (char *) ShmemBase + newStart;
+ newSpace = (char *) segment->ShmemBase + newStart;
ShmemAllocator->free_offset = newFree;
}
else
newSpace = NULL;
- SpinLockRelease(ShmemLock);
+ SpinLockRelease(&ShmemAllocator->shmem_lock);
/* note this assert is okay with newSpace == NULL */
Assert(newSpace == (void *) CACHELINEALIGN(newSpace));
@@ -257,14 +296,23 @@ ShmemAllocRaw(Size size, Size *allocated_size)
}
/*
- * ShmemAddrIsValid -- test if an address refers to shared memory
+ * ShmemAddrIsValid
+ * test if an address refers to the given shared memory segment.
*
* Returns true if the pointer points within the shared memory segment.
*/
bool
-ShmemAddrIsValid(const void *addr)
+ShmemAddrIsValid(int segment_id, const void *addr)
{
- return (addr >= ShmemBase) && (addr < ShmemEnd);
+ ShmemSegment *segment;
+ void *shmemEnd;
+
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+
+ segment = &Segments[segment_id];
+ shmemEnd = (char *) segment->ShmemBase + segment->ShmemSegHdr->totalsize;
+
+ return (addr >= segment->ShmemBase) && (addr < shmemEnd);
}
/*
@@ -318,6 +366,9 @@ InitShmemIndex(void)
* Note: before Postgres 9.0, this function returned NULL for some failure
* cases. Now, it always throws error instead, so callers need not check
* for NULL.
+ *
+ * See prologue of ShmemAllocNoError for explanation about lack of segment_id
+ * parameter.
*/
HTAB *
ShmemInitHash(const char *name, /* table string name for shmem index */
@@ -341,9 +392,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
/* look it up in the shmem index */
- location = ShmemInitStruct(name,
- hash_get_shared_size(infoP, hash_flags),
- &found);
+ location = ShmemInitStructInSegment(name,
+ hash_get_shared_size(infoP, hash_flags),
+ &found, MAIN_SHMEM_SEGMENT);
/*
* if it already exists, attach to it rather than allocate and initialize
@@ -376,15 +427,32 @@ ShmemInitHash(const char *name, /* table string name for shmem index */
*/
void *
ShmemInitStruct(const char *name, Size size, bool *foundPtr)
+{
+ return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT);
+}
+
+void *
+ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, int segment_id)
{
ShmemIndexEnt *result;
void *structPtr;
+ ShmemSegment *segment;
+
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+
+ segment = &Segments[segment_id];
LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
if (!ShmemIndex)
{
- /* Must be trying to create/attach to ShmemIndex itself */
+ ShmemAllocatorData *ShmemAllocator = segment->ShmemAllocator;
+
+ /*
+ * Must be trying to create/attach to ShmemIndex itself in the main
+ * shared memory segment.
+ */
+ Assert(segment_id == MAIN_SHMEM_SEGMENT);
Assert(strcmp(name, "ShmemIndex") == 0);
if (IsUnderPostmaster)
@@ -405,7 +473,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
* process can be accessing shared memory yet.
*/
Assert(ShmemAllocator->index == NULL);
- structPtr = ShmemAlloc(size);
+ structPtr = ShmemAllocInternal(segment, size);
ShmemAllocator->index = structPtr;
*foundPtr = false;
}
@@ -422,8 +490,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
LWLockRelease(ShmemIndexLock);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("could not create ShmemIndex entry for data structure \"%s\"",
- name)));
+ errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d",
+ name, segment_id)));
}
if (*foundPtr)
@@ -448,7 +516,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Size allocated_size;
/* It isn't in the table yet. allocate and initialize it */
- structPtr = ShmemAllocRaw(size, &allocated_size);
+ structPtr = ShmemAllocRaw(segment, size, &allocated_size);
if (structPtr == NULL)
{
/* out of memory; remove the failed ShmemIndex entry */
@@ -463,18 +531,18 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
result->size = size;
result->allocated_size = allocated_size;
result->location = structPtr;
+ result->segment_id = segment_id;
}
LWLockRelease(ShmemIndexLock);
- Assert(ShmemAddrIsValid(structPtr));
+ Assert(ShmemAddrIsValid(segment_id, structPtr));
Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
return structPtr;
}
-
/*
* Add two Size values, checking for overflow
*/
@@ -509,13 +577,14 @@ mul_size(Size s1, Size s2)
Datum
pg_get_shmem_allocations(PG_FUNCTION_ARGS)
{
-#define PG_GET_SHMEM_SIZES_COLS 4
+#define PG_GET_SHMEM_SIZES_COLS 5
ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
HASH_SEQ_STATUS hstat;
ShmemIndexEnt *ent;
- Size named_allocated = 0;
+ Size named_allocated[NUM_MEMORY_MAPPINGS] = {0};
Datum values[PG_GET_SHMEM_SIZES_COLS];
bool nulls[PG_GET_SHMEM_SIZES_COLS];
+ int i;
InitMaterializedSRF(fcinfo, 0);
@@ -527,30 +596,49 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
memset(nulls, 0, sizeof(nulls));
while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL)
{
+ ShmemSegment *segment = &Segments[ent->segment_id];
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+
values[0] = CStringGetTextDatum(ent->key);
- values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
- values[2] = Int64GetDatum(ent->size);
- values[3] = Int64GetDatum(ent->allocated_size);
- named_allocated += ent->allocated_size;
+ values[1] = CStringGetTextDatum(segment->ShmemSegmentName);
+ values[2] = Int64GetDatum((char *) ent->location - (char *) shmhdr);
+ values[3] = Int64GetDatum(ent->size);
+ values[4] = Int64GetDatum(ent->allocated_size);
+ named_allocated[ent->segment_id] += ent->allocated_size;
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
values, nulls);
}
/* output shared memory allocated but not counted via the shmem index */
- values[0] = CStringGetTextDatum("<anonymous>");
- nulls[1] = true;
- values[2] = Int64GetDatum(ShmemAllocator->free_offset - named_allocated);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ ShmemSegment *segment = &Segments[i];
+ ShmemAllocatorData *ShmemAllocator = segment->ShmemAllocator;
+
+ values[0] = CStringGetTextDatum("<anonymous>");
+ values[1] = CStringGetTextDatum(segment->ShmemSegmentName);
+ nulls[2] = true;
+ values[3] = Int64GetDatum(ShmemAllocator->free_offset - named_allocated[i]);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
/* output as-of-yet unused shared memory */
- nulls[0] = true;
- values[1] = Int64GetDatum(ShmemAllocator->free_offset);
- nulls[1] = false;
- values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemAllocator->free_offset);
- values[3] = values[2];
- tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ memset(nulls, 0, sizeof(nulls));
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+ ShmemAllocatorData *ShmemAllocator = segment->ShmemAllocator;
+
+ nulls[0] = true;
+ values[1] = CStringGetTextDatum(segment->ShmemSegmentName);
+ values[2] = Int64GetDatum(ShmemAllocator->free_offset);
+ values[3] = Int64GetDatum(shmhdr->totalsize - ShmemAllocator->free_offset);
+ values[4] = values[3];
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
LWLockRelease(ShmemIndexLock);
@@ -575,7 +663,7 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size os_page_size;
void **page_ptrs;
int *pages_status;
- uint64 shm_total_page_count,
+ uint64 shm_total_page_count = 0,
shm_ent_page_count,
max_nodes;
Size *nodes;
@@ -610,7 +698,13 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
* this is not very likely, and moreover we have more entries, each of
* them using only fraction of the total pages.
*/
- shm_total_page_count = (ShmemSegHdr->totalsize / os_page_size) + 1;
+ for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
+ {
+ PGShmemHeader *shmhdr = Segments[segment].ShmemSegHdr;
+
+ shm_total_page_count += (shmhdr->totalsize / os_page_size) + 1;
+ }
+
page_ptrs = palloc0_array(void *, shm_total_page_count);
pages_status = palloc_array(int, shm_total_page_count);
@@ -751,7 +845,7 @@ pg_get_shmem_pagesize(void)
Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
+ GetHugePageSize(&os_page_size, NULL, NULL);
return os_page_size;
}
@@ -761,3 +855,45 @@ pg_numa_available(PG_FUNCTION_ARGS)
{
PG_RETURN_BOOL(pg_numa_init() != -1);
}
+
+/* SQL SRF showing shared memory segments */
+Datum
+pg_get_shmem_segments(PG_FUNCTION_ARGS)
+{
+#define PG_GET_SHMEM_SEGS_COLS 5
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ Datum values[PG_GET_SHMEM_SEGS_COLS];
+ bool nulls[PG_GET_SHMEM_SEGS_COLS];
+ int i;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* output all allocated entries */
+ for (i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ ShmemSegment *segment = &Segments[i];
+ PGShmemHeader *shmhdr = segment->ShmemSegHdr;
+ ShmemAllocatorData *ShmemAllocator = segment->ShmemAllocator;
+ int j;
+
+ if (shmhdr == NULL)
+ {
+ for (j = 0; j < PG_GET_SHMEM_SEGS_COLS; j++)
+ nulls[j] = true;
+ }
+ else
+ {
+ memset(nulls, 0, sizeof(nulls));
+ values[0] = Int32GetDatum(i);
+ values[1] = CStringGetTextDatum(segment->ShmemSegmentName);
+ values[2] = Int64GetDatum(shmhdr->totalsize);
+ values[3] = Int64GetDatum(ShmemAllocator->free_offset);
+ values[4] = Int64GetDatum(shmhdr->reservedsize);
+ }
+
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
+
+ return (Datum) 0;
+}
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 517c55375b4..160b6927c23 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -80,6 +80,8 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "port/pg_bitutils.h"
+#include "postmaster/postmaster.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/procnumber.h"
@@ -446,7 +448,7 @@ CreateLWLocks(void)
char *ptr;
/* Allocate space */
- ptr = (char *) ShmemAlloc(spaceLocks);
+ ptr = (char *) ShmemAlloc(MAIN_SHMEM_SEGMENT, spaceLocks);
/* Initialize the dynamic-allocation counter for tranches */
LWLockCounter = (int *) ptr;
@@ -612,6 +614,9 @@ LWLockNewTrancheId(const char *name)
/*
* We use the ShmemLock spinlock to protect LWLockCounter and
* LWLockTrancheNames.
+ *
+ * XXX: Looks like this is the only use of Segments outside of shmem.c,
+ * it's maybe worth it to reshape this part to hide Segments structure.
*/
SpinLockAcquire(ShmemLock);
diff --git a/src/backend/storage/lmgr/predicate.c b/src/backend/storage/lmgr/predicate.c
index fe75ead3501..9aab75f54d6 100644
--- a/src/backend/storage/lmgr/predicate.c
+++ b/src/backend/storage/lmgr/predicate.c
@@ -207,6 +207,7 @@
#include "miscadmin.h"
#include "pgstat.h"
#include "port/pg_lfind.h"
+#include "storage/pg_shmem.h"
#include "storage/predicate.h"
#include "storage/predicate_internals.h"
#include "storage/proc.h"
@@ -595,7 +596,7 @@ CreatePredXact(void)
static void
ReleasePredXact(SERIALIZABLEXACT *sxact)
{
- Assert(ShmemAddrIsValid(sxact));
+ Assert(ShmemAddrIsValid(MAIN_SHMEM_SEGMENT, sxact));
dlist_delete(&sxact->xactLink);
dlist_push_tail(&PredXact->availableList, &sxact->xactLink);
diff --git a/src/backend/utils/activity/pgstat_shmem.c b/src/backend/utils/activity/pgstat_shmem.c
index 33fbdca9609..c6d9157f417 100644
--- a/src/backend/utils/activity/pgstat_shmem.c
+++ b/src/backend/utils/activity/pgstat_shmem.c
@@ -13,6 +13,7 @@
#include "postgres.h"
#include "pgstat.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "utils/memutils.h"
#include "utils/pgstat_internal.h"
@@ -233,7 +234,7 @@ StatsShmemInit(void)
int idx = kind - PGSTAT_KIND_CUSTOM_MIN;
Assert(kind_info->shared_size != 0);
- ctl->custom_data[idx] = ShmemAlloc(kind_info->shared_size);
+ ctl->custom_data[idx] = ShmemAlloc(MAIN_SHMEM_SEGMENT, kind_info->shared_size);
ptr = ctl->custom_data[idx];
}
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 83f6501df38..4b27f2a245e 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8592,8 +8592,8 @@
{ oid => '5052', descr => 'allocations from the main shared memory segment',
proname => 'pg_get_shmem_allocations', prorows => '50', proretset => 't',
provolatile => 'v', prorettype => 'record', proargtypes => '',
- proallargtypes => '{text,int8,int8,int8}', proargmodes => '{o,o,o,o}',
- proargnames => '{name,off,size,allocated_size}',
+ proallargtypes => '{text,text,int8,int8,int8}', proargmodes => '{o,o,o,o,o}',
+ proargnames => '{name,segment,off,size,allocated_size}',
prosrc => 'pg_get_shmem_allocations' },
{ oid => '4099', descr => 'Is NUMA support available?',
@@ -8616,6 +8616,14 @@
proargmodes => '{o,o,o}', proargnames => '{name,type,size}',
prosrc => 'pg_get_dsm_registry_allocations' },
+# shared memory segments
+{ oid => '5101', descr => 'shared memory segments',
+ proname => 'pg_get_shmem_segments', prorows => '6', proretset => 't',
+ provolatile => 'v', prorettype => 'record', proargtypes => '',
+ proallargtypes => '{int4,text,int8,int8,int8}', proargmodes => '{o,o,o,o,o}',
+ proargnames => '{id,name,size,freeoffset,reserved_size}',
+ prosrc => 'pg_get_shmem_segments' },
+
# memory context of local backend
{ oid => '2282',
descr => 'information about all memory contexts of local backend',
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index a40adf6b2a8..93348a34378 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -19,6 +19,7 @@
#include "storage/block.h"
#include "storage/buf.h"
#include "storage/bufpage.h"
+#include "storage/pg_shmem.h"
#include "storage/relfilelocator.h"
#include "utils/relcache.h"
#include "utils/snapmgr.h"
@@ -367,7 +368,7 @@ extern void MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
-extern Size BufferManagerShmemSize(void);
+extern Size BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index da32787ab51..f1d0802d048 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -18,6 +18,8 @@
#ifndef IPC_H
#define IPC_H
+#include "storage/pg_shmem.h"
+
typedef void (*pg_on_exit_callback) (int code, Datum arg);
typedef void (*shmem_startup_hook_type) (void);
@@ -77,7 +79,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
-extern Size CalculateShmemSize(void);
+extern Size CalculateShmemSize(MemoryMappingSizes *mapping_sizes);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 10c7b065861..eafab1dae52 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -26,12 +26,20 @@
#include "storage/dsm_impl.h"
+
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
{
int32 magic; /* magic # to identify Postgres segments */
#define PGShmemMagic 679834894
pid_t creatorPID; /* PID of creating process (set but unread) */
+
+ /*
+ * TODO: We might have to rename these fields to allocSize (for amount of
+ * memory allocated currently in this segment), maxSize (for maximum size
+ * the segment can grow to.)
+ */
Size totalsize; /* total size of segment */
+ Size reservedsize; /* Size of the reserved mapping */
Size content_offset; /* offset to the data, i.e. size of this
* header */
dsm_handle dsm_control; /* ID of dynamic shared memory control seg */
@@ -41,6 +49,55 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */
#endif
} PGShmemHeader;
+/*
+ * Information about the shared memory segment that is required to be passed
+ * from the Postmaster to each backend.
+ */
+typedef struct PGUsedShmemInfo
+{
+ void *UsedShmemSegAddr; /* SysV shared memory for the header */
+#ifndef WIN32
+ unsigned long UsedShmemSegID; /* IPC key */
+#else
+ void *ShmemProtectiveRegion; /* Protective region for Windows
+ * shared memory */
+ HANDLE UsedShmemSegID;
+#endif
+} PGUsedShmemInfo;
+
+/*
+ * To be able to dynamically resize the shared buffer pool, we allocate shared
+ * memory in two segments. Main segment which contains everything except the
+ * buffer blocks and BUFFERS_SHMEM_SEGMENT which contains the buffer blocks.
+ * Main segment is fixed sized whereas BUFFERS_SHMEM_SEGMENT can be resized
+ * during runtime.
+ *
+ * TODO: convert this to enum?
+ */
+
+#define MAIN_SHMEM_SEGMENT 0
+
+/* Buffer blocks */
+#define BUFFERS_SHMEM_SEGMENT 1
+
+/* Number of available segments for anonymous memory mappings */
+#define NUM_MEMORY_MAPPINGS 2
+
+/*
+ * Structure to hold required sizes of each shared memory segment as calculated
+ * by CalculateShmemSize().
+ *
+ * TODO: Does ShmemMappingSizes sound better?
+ */
+typedef struct MemoryMappingSizes
+{
+ Size shmem_req_size; /* Required size of the segment */
+ Size shmem_reserved; /* Required size of the reserved address
+ * space. */
+} MemoryMappingSizes;
+
+extern PGDLLIMPORT PGUsedShmemInfo UsedShmemInfo[NUM_MEMORY_MAPPINGS];
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
@@ -64,14 +121,6 @@ typedef enum
SHMEM_TYPE_MMAP,
} PGShmemType;
-#ifndef WIN32
-extern PGDLLIMPORT unsigned long UsedShmemSegID;
-#else
-extern PGDLLIMPORT HANDLE UsedShmemSegID;
-extern PGDLLIMPORT void *ShmemProtectiveRegion;
-#endif
-extern PGDLLIMPORT void *UsedShmemSegAddr;
-
#if !defined(WIN32) && !defined(EXEC_BACKEND)
#define DEFAULT_SHARED_MEMORY_TYPE SHMEM_TYPE_MMAP
#elif !defined(WIN32)
@@ -85,10 +134,27 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(Size size,
+extern PGShmemHeader *PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping_sizes,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+extern void PrepareHugePages(void);
+
+static inline const char *
+MappingName(int segment_id)
+{
+ switch (segment_id)
+ {
+ case MAIN_SHMEM_SEGMENT:
+ return "main";
+ case BUFFERS_SHMEM_SEGMENT:
+ return "buffers";
+ default:
+ return "unknown";
+ }
+}
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 89d45287c17..5d58b3b39e6 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -29,14 +29,16 @@
extern PGDLLIMPORT slock_t *ShmemLock;
typedef struct PGShmemHeader PGShmemHeader; /* avoid including
* storage/pg_shmem.h here */
-extern void InitShmemAllocator(PGShmemHeader *seghdr);
-extern void *ShmemAlloc(Size size);
+extern void InitShmemAllocator(int segment_id, PGShmemHeader *seghdr);
+extern void *ShmemAlloc(int segment_id, Size size);
extern void *ShmemAllocNoError(Size size);
-extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemAddrIsValid(int segment_id, const void *addr);
extern void InitShmemIndex(void);
extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
HASHCTL *infoP, int hash_flags);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+extern void *ShmemInitStructInSegment(const char *name, Size size,
+ bool *foundPtr, int segment_id);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
@@ -58,6 +60,7 @@ typedef struct
void *location; /* location in shared mem */
Size size; /* # bytes requested for the structure */
Size allocated_size; /* # bytes actually allocated */
+ int segment_id; /* segment in which the structure is allocated */
} ShmemIndexEnt;
#endif /* SHMEM_H */
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index f4ee2bd7459..1e1bd1eb8b4 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1768,14 +1768,21 @@ pg_shadow| SELECT pg_authid.rolname AS usename,
LEFT JOIN pg_db_role_setting s ON (((pg_authid.oid = s.setrole) AND (s.setdatabase = (0)::oid))))
WHERE pg_authid.rolcanlogin;
pg_shmem_allocations| SELECT name,
+ segment,
off,
size,
allocated_size
- FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size);
+ FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, segment, off, size, allocated_size);
pg_shmem_allocations_numa| SELECT name,
numa_node,
size
FROM pg_get_shmem_allocations_numa() pg_get_shmem_allocations_numa(name, numa_node, size);
+pg_shmem_segments| SELECT id,
+ name,
+ size,
+ freeoffset,
+ reserved_size
+ FROM pg_get_shmem_segments() pg_get_shmem_segments(id, name, size, freeoffset, reserved_size);
pg_stat_activity| SELECT s.datid,
d.datname,
s.pid,
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 77518489412..e4dbe1b787b 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -120,6 +120,7 @@ AmcheckOptions
AnalyzeAttrComputeStatsFunc
AnalyzeAttrFetchFunc
AnalyzeForeignTable_function
+AnonShmemData
AnlExprData
AnlIndexData
AnyArrayType
@@ -1685,6 +1686,7 @@ MVNDistinct
MVNDistinctItem
ManyTestResource
ManyTestResourceKind
+MemoryMappingSizes
Material
MaterialPath
MaterialState
@@ -1887,6 +1889,7 @@ PGFInfoFunction
PGFileType
PGFunction
PGIOAlignedBlock
+PGUsedShmemInfo
PGLZ_HistEntry
PGLZ_Strategy
PGLoadBalanceType
@@ -2807,6 +2810,7 @@ ShippableCacheEntry
ShmemAllocatorData
ShippableCacheKey
ShmemIndexEnt
+ShmemSegment
ShutdownForeignScan_function
ShutdownInformation
ShutdownMode
--
2.34.1
[application/x-patch] v20260209-0003-Allow-to-resize-shared-memory-without-rest.patch (147.0K, ../../CAExHW5tSw8r06RLAArvf923cO4NGetitPhQ7AO0o7hsKx8jsNw@mail.gmail.com/4-v20260209-0003-Allow-to-resize-shared-memory-without-rest.patch)
download | inline diff:
From 8cf16c7d4db804442a4b420c393443002bfaa06d Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Tue, 17 Jun 2025 14:16:55 +0200
Subject: [PATCH v20260209 3/7] Allow to resize shared memory without restart
shared_buffers is now PGC_SIGHUP instead of PGC_POSTMASTER. The value
of this GUC is saved in NBuffersPending instead of NBuffers. When the
server starts, the shared memory size is estimated and the memory is
allocated using NBuffersPending.
When a server is running, the new value of GUC (set using ALTER SYSTEM
... SET shared_buffers = ...; followed by SELECT pg_reload_conf()) does
not come into effect immediately. Instead a function
pg_resize_shared_buffers() is used to resize the buffer pool. The
function uses the current value of GUC in the backends where it is
executed. The function also coordinates the buffer resizing
synchronization across backends.
SHOW shared_buffers now shows the current size of the shared buffer pool
but it also shows pending size of shared buffers, if any.
A new GUC max_shared_buffers is introduced to control the maximum value
of shared_buffers that can be set. By default it is 0. When explicitly
set it needs to be higher than 'shared_buffers'. When max_shared_buffers
is set to 0, it assumes the same value as GUC shared_buffers. This GUC
determines the size of address space reserved for future buffer pool
sizes and the size of buffer look up table.
TBD: Describe the protocol used by pg_resize_shared_buffers() to
synchronize buffer resizing operation with other backends.
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking.
If a buffer being evicted is pinned, we abort the resizing operation.
There are other alternative which are not implemented in the current
patches 1. to wait for the pinned buffer to get unpinned, 2. the backend
is killed or it itself cancels the query or 3. rollback the operation.
Note that option 1 and 2 would require the pinning related local and
shared records to be accessed. But we need infrastructure to do either
of this right now.
So far the buffer pool metdata (NBuffers and the shared memory segment
address space) is saved in process local heap memory since it's static
for the life of a server. It is passed to a new backend through
Postmaster. But with buffer pool being resized while the server running,
we need Postmaster to update its buffer pool metadata as the resizing
progresses and pass it to the new backend. This has few complications:
1. Postmaster does not receive ProcSignalBarrier. So we need to signal
it separately.
2. Postmaster's local state is inherited by the new backend when
fork()ed. But we need more complex implementation to pass it to an
exec()ed backend.
3. A new backend may receive the updated state from Postmaster and also
the signal barrier which prompts the same update. Thus the proc
signal barrier code needs to be idempotent; adding further complexity
to it.
4. This task takes away Postmaster resources from it's core
functionality.
This can be avoided by following two changes:
1. The shared memory is resized only in a single backend without
requiring any changes to the memory address space.
2. Maintaining the buffer pool metadata in the shared memory instead of
process local memory. This change may affect performance so verify
that performance is not degraded.
TODO: In case the backend executing pg_resize_shared_buffers() exits
before the operation finishes, we need to make sure that the changes
made to the shared memory while resizing are cleaned up properly.
Removing the evicted buffers from buffer ring
=============================================
If the buffer pool has been shrunk, the buffers in the buffer ring may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition and masks any bug that
introduces an invalid buffer into the ring. The alternative fix is more
complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure thus accessible when processing buffer
resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work.
Author: Ashutosh Bapat
Author: Dmitrii Dolgov
Author of some tests: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
Reviewed-by: Tomas Vondra
More detailed note follow: Need to see which of those fit in the commit
message and which should be removed.
Reinitializing strategry control area
=====================================
The commit introduces a separate function StrategyReInitialize() instead
of reusing StrategyInitialize() since some of the things that the second
one does are not required in the first one. Here's list of what
StrategyReInitialize() does and how does it differ from
StrategyInitialize().
1. StrategyControl pointer needn't be fetched again since it should not
change. But added an Assert to make sure the pointer is valid.
2. &StrategyControl->buffer_strategy_lock need not be initialized again.
3. nextVictimBuffer, completePasses and numBufferAllocs are viewed in
the context of NBuffers. Now that NBuffers itself has changed, those
three do not make sense. Reset them as if the server has restarted
again.
Ability to delay resizing operation
===================================
This commit introduces a flag delay_shmem_resize, which postgresql
backends and workers can use to signal the coordinator to delay resizing
operation. Background writer sets this flag when its scanning buffers.
Background writer operation (needs a rethink)
===========================
Background writer is blocked when the actual resizing is in progress. It
stops a scan in progress when it sees that the resizing has begun or is
about to begin. Once the buffer resizing is finished, before resuming
the regular operation, bgwriter resets the information saved so far.
This information is viewed in the context of NBuffers and hence does not
make sense after resizing which chanegs NBuffers.
Buffer lookup table
===================
Right now there is no way to free shared memory. Even if we shrink the
buffer lookup table when shrinking the buffer pool the unused hash table
entries can not be freed. When we expand the buffer pool, more entries
can be allocated but we can not resize the hash table directory without
rehashing all the entries. Just allocating more entries will lead to
more contention. Hence we setup the buffer lookup table considering the
maximum possible size of the buffer pool which is MaxAvailableMemory
only once at the beginning. Shared buffer lookup table and
StrategyControl are not resized even if the buffer pool is resized hence
they are allocated in the main shared memory segment
BgWriter refactoring
====================
The way BgBufferSync is written today, it packs four functionalities:
setting up the buffer sync state, performing the buffer sync, resetting
the buffer sync state when bgwriter_lru_maxpages <= 0 and setting it up
again after bgwriter_lru_maxpages > 0. That makes the code hard to read.
It will be good to divide this function into 3/4 different functions
each performing one functionality. Then pack all the state (the local
variables from that function converted to static global) into a
structure, which is passed to these functions. Once that happens
BgBufferSyncReset() will call one of the functions to reset the state
when buffer pool is resized.
---
contrib/pg_buffercache/pg_buffercache_pages.c | 18 +-
doc/src/sgml/config.sgml | 45 +-
doc/src/sgml/func/func-admin.sgml | 57 +++
src/backend/access/transam/slru.c | 2 +-
src/backend/access/transam/xlog.c | 2 +-
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/optimizer/path/costsize.c | 8 +-
src/backend/port/sysv_shmem.c | 164 +++++++
src/backend/port/win32_shmem.c | 6 +
src/backend/postmaster/checkpointer.c | 23 +-
src/backend/postmaster/postmaster.c | 10 +-
src/backend/storage/aio/aio_init.c | 2 +-
src/backend/storage/buffer/Makefile | 3 +-
src/backend/storage/buffer/README | 74 +++
src/backend/storage/buffer/buf_init.c | 220 +++++++--
src/backend/storage/buffer/buf_resize.c | 455 ++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 10 +-
src/backend/storage/buffer/bufmgr.c | 178 ++++++-
src/backend/storage/buffer/freelist.c | 101 +++-
src/backend/storage/buffer/meson.build | 1 +
src/backend/storage/ipc/ipci.c | 13 +-
src/backend/storage/ipc/procsignal.c | 56 +++
src/backend/storage/ipc/shmem.c | 103 +++-
src/backend/tcop/postgres.c | 11 +
.../utils/activity/wait_event_names.txt | 2 +
src/backend/utils/init/globals.c | 5 +-
src/backend/utils/init/postinit.c | 49 ++
src/backend/utils/misc/guc.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 15 +-
src/include/catalog/pg_proc.dat | 6 +
src/include/miscadmin.h | 8 +
src/include/storage/buf.h | 2 +-
src/include/storage/buf_internals.h | 5 +-
src/include/storage/bufmgr.h | 20 +-
src/include/storage/ipc.h | 1 +
src/include/storage/lwlocklist.h | 1 +
src/include/storage/pg_shmem.h | 62 ++-
src/include/storage/procsignal.h | 9 +
src/include/storage/shmem.h | 3 +
src/include/utils/guc.h | 2 +
src/test/Makefile | 3 +-
src/test/README | 3 +
src/test/buffermgr/Makefile | 30 ++
src/test/buffermgr/README | 26 +
src/test/buffermgr/buffermgr_test.conf | 11 +
src/test/buffermgr/expected/buffer_resize.out | 312 ++++++++++++
src/test/buffermgr/meson.build | 23 +
src/test/buffermgr/sql/buffer_resize.sql | 97 ++++
src/test/buffermgr/t/001_resize_buffer.pl | 143 ++++++
.../buffermgr/t/003_parallel_resize_buffer.pl | 71 +++
.../t/004_client_join_buffer_resize.pl | 243 ++++++++++
src/test/meson.build | 1 +
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 76 +++
src/tools/pgindent/typedefs.list | 1 +
54 files changed, 2685 insertions(+), 111 deletions(-)
create mode 100644 src/backend/storage/buffer/buf_resize.c
create mode 100644 src/test/buffermgr/Makefile
create mode 100644 src/test/buffermgr/README
create mode 100644 src/test/buffermgr/buffermgr_test.conf
create mode 100644 src/test/buffermgr/expected/buffer_resize.out
create mode 100644 src/test/buffermgr/meson.build
create mode 100644 src/test/buffermgr/sql/buffer_resize.sql
create mode 100644 src/test/buffermgr/t/001_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/003_parallel_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/004_client_join_buffer_resize.pl
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index f60f797a9b4..8a17319ff2a 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -125,6 +125,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
TupleDesc tupledesc;
TupleDesc expected_tupledesc;
HeapTuple tuple;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
if (SRF_IS_FIRSTCALL())
{
@@ -181,10 +182,10 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
/* Allocate NBuffers worth of BufferCachePagesRec records. */
fctx->record = (BufferCachePagesRec *)
MemoryContextAllocHuge(CurrentMemoryContext,
- sizeof(BufferCachePagesRec) * NBuffers);
+ sizeof(BufferCachePagesRec) * currentNBuffers);
/* Set max calls and remember the user function context. */
- funcctx->max_calls = NBuffers;
+ funcctx->max_calls = currentNBuffers;
funcctx->user_fctx = fctx;
/* Return to original context when allocating transient memory */
@@ -198,13 +199,24 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
* snapshot across all buffers, but we do grab the buffer header
* locks, so the information of each buffer is self-consistent.
*/
- for (i = 0; i < NBuffers; i++)
+ for (i = 0; i < currentNBuffers; i++)
{
BufferDesc *bufHdr;
uint64 buf_state;
CHECK_FOR_INTERRUPTS();
+ /*
+ * TODO: We should just scan the entire buffer descriptor array
+ * instead of relying on curent buffer pool size. But that can
+ * happen if only we setup the descriptor array large enough at
+ * the server startup time.
+ */
+ if (currentNBuffers != pg_atomic_read_u32(&ShmemCtrl->currentNBuffers))
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("number of shared buffers changed during scan of buffer cache")));
+
bufHdr = GetBufferDescriptor(i);
/* Lock each buffer header before inspecting. */
buf_state = LockBufHdr(bufHdr);
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index f1af1505cf3..71654f94ce8 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -1732,7 +1732,6 @@ include_dir 'conf.d'
that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
(Non-default values of <symbol>BLCKSZ</symbol> change the minimum
value.)
- This parameter can only be set at server start.
</para>
<para>
@@ -1755,6 +1754,50 @@ include_dir 'conf.d'
appropriate, so as to leave adequate space for the operating system.
</para>
+ <para>
+ The shared memory consumed by the buffer pool is allocated and
+ initialized according to the value of the GUC at the time of starting
+ the server. A desired new value of GUC can be loaded while the server is
+ running using <systemitem>SIGHUP</systemitem>. But the buffer pool will
+ not be resized immediately. Use
+ <function>pg_resize_shared_buffers()</function> to dynamically resize
+ the shared buffer pool (see <xref linkend="functions-admin"/> for details).
+ <command>SHOW shared_buffers</command> shows the current number of
+ shared buffers and pending number, if any. Please note that when the GUC
+ is changed, the other GUCS which use this GUCs value to set their
+ defaults will not be changed. They may still require a server restart to
+ consider new value.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-max-shared-buffers" xreflabel="max_shared_buffers">
+ <term><varname>max_shared_buffers</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>max_shared_buffers</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Sets the upper limit for the <varname>shared_buffers</varname> value.
+ The default value is <literal>0</literal>,
+ which means no explicit limit is set and <varname>max_shared_buffers</varname>
+ will be automatically set to the value of <varname>shared_buffers</varname>
+ at server startup.
+ If this value is specified without units, it is taken as blocks,
+ that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
+ This parameter can only be set at server start.
+ </para>
+
+ <para>
+ This parameter determines the amount of memory address space to reserve
+ in each backend for expanding the buffer pool in future. While the
+ memory for buffer pool is allocated on demand as it is resized, the
+ memory required to hold the buffer manager metadata is allocated
+ statically at the server start accounting for the largest buffer pool
+ size allowed by this parameter.
+ <!-- TODO: Provide a numeric example of how much extra memory say max_shared_buffers = 1GB consume. -->
+ </para>
</listitem>
</varlistentry>
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 3ac81905d1f..22bae4e0bac 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -99,6 +99,63 @@
<returnvalue>off</returnvalue>
</para></entry>
</row>
+
+ <row>
+ <entry role="func_table_entry"><para role="func_signature">
+ <indexterm>
+ <primary>pg_resize_shared_buffers</primary>
+ </indexterm>
+ <function>pg_resize_shared_buffers</function> ()
+ <returnvalue>boolean</returnvalue>
+ </para>
+ <para>
+ Dynamically resizes the shared buffer pool to match the current
+ value of the <varname>shared_buffers</varname> parameter. This
+ function implements a coordinated resize process that ensures all
+ backend processes acknowledge the change before completing the
+ operation. The resize happens in multiple phases to maintain
+ data consistency and system stability. Returns <literal>true</literal>
+ if the resize was successful, or raises an error if the operation
+ fails. This function can only be called by superusers.
+ </para>
+ <para>
+ To resize shared buffers, first update the <varname>shared_buffers</varname>
+ setting and reload the configuration, then verify the new value is loaded
+ before calling this function. For example:
+<programlisting>
+postgres=# ALTER SYSTEM SET shared_buffers = '256MB';
+ALTER SYSTEM
+postgres=# SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers;
+ shared_buffers
+-------------------------
+ 128MB (pending: 256MB)
+(1 row)
+
+postgres=# SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+</programlisting>
+ The <command>SHOW shared_buffers</command> step is important to verify
+ that the configuration reload was successful and the new value is
+ available to the current session before attempting the resize. The
+ output shows both the current and pending values when a change is waiting
+ to be applied.
+ </para></entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/access/transam/slru.c b/src/backend/access/transam/slru.c
index 549c7e3e64b..2f30be34731 100644
--- a/src/backend/access/transam/slru.c
+++ b/src/backend/access/transam/slru.c
@@ -232,7 +232,7 @@ SimpleLruAutotuneBuffers(int divisor, int max)
{
return Min(max - (max % SLRU_BANK_SIZE),
Max(SLRU_BANK_SIZE,
- NBuffers / divisor - (NBuffers / divisor) % SLRU_BANK_SIZE));
+ NBuffersPending / divisor - (NBuffersPending / divisor) % SLRU_BANK_SIZE));
}
/*
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 13ec6225b85..b2605dbf87a 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4694,7 +4694,7 @@ XLOGChooseNumBuffers(void)
{
int xbuffers;
- xbuffers = NBuffers / 32;
+ xbuffers = NBuffersPending / 32;
if (xbuffers > (wal_segment_size / XLOG_BLCKSZ))
xbuffers = (wal_segment_size / XLOG_BLCKSZ);
if (xbuffers < 8)
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index 7d32cd0e159..638940363b4 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -337,6 +337,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ InitializeMaxNBuffers();
+
CreateSharedMemoryAndSemaphores();
/*
diff --git a/src/backend/optimizer/path/costsize.c b/src/backend/optimizer/path/costsize.c
index c30d6e84672..23d87bdb3c9 100644
--- a/src/backend/optimizer/path/costsize.c
+++ b/src/backend/optimizer/path/costsize.c
@@ -19,10 +19,10 @@
* is normally considerably less than random_page_cost. (However, if the
* database is fully cached in RAM, it is reasonable to set them equal.)
*
- * We also use a rough estimate "effective_cache_size" of the number of
- * disk pages in Postgres + OS-level disk cache. (We can't simply use
- * NBuffers for this purpose because that would ignore the effects of
- * the kernel's disk cache.)
+ * We also use a rough estimate "effective_cache_size" of the number of disk
+ * pages in Postgres + OS-level disk cache. (We can't simply use size of the
+ * buffer pool for this purpose because that would ignore the effects of the
+ * kernel's disk cache.)
*
* Obviously, taking constants for these values is an oversimplification,
* but it's tough enough to get any useful estimates even at this level of
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 29ffcaa35f3..c766f5c44fd 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -30,14 +30,19 @@
#include "miscadmin.h"
#include "port/pg_bitutils.h"
#include "portability/mem.h"
+#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/fd.h"
#include "storage/ipc.h"
+#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
#include "storage/shmem.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
#include "utils/pidfile.h"
+#include "utils/wait_event.h"
/*
* TODO: The first two sentences in the first paragraph below make me feel like
@@ -101,6 +106,8 @@ typedef enum
SHMSTATE_UNATTACHED, /* pertinent to DataDir, no attached PIDs */
} IpcMemoryState;
+volatile bool delay_shmem_resize = false;
+
/*
* Anonymous mapping layout we use looks like this:
*
@@ -127,6 +134,9 @@ typedef enum
* being counted against memory limits). The mapping serves as an address space
* reservation, into which shared memory segment can be extended and is
* represented by the second /memfd:main with no permissions.
+ *
+ * The reserved space for buffer manager related segments is calculated based on
+ * MaxNBuffers.
*/
PGUsedShmemInfo UsedShmemInfo[NUM_MEMORY_MAPPINGS];
@@ -152,6 +162,42 @@ AnonShmemData AnonShmemInfo[NUM_MEMORY_MAPPINGS];
*/
static bool huge_pages_on = false;
+/*
+ * Currently broadcasted value of NBuffers in shared memory.
+ *
+ * Most of the time this value is going to be equal to NBuffers. But if
+ * postmaster is resizing shared memory and a new backend was created
+ * at the same time, there is a possibility for the new backend to inherit the
+ * old NBuffers value, but miss the resize signal if ProcSignal infrastructure
+ * was not initialized yet. Consider this situation:
+ *
+ * Postmaster ------> New Backend
+ * | |
+ * | Launch
+ * | |
+ * | Inherit NBuffers
+ * | |
+ * Resize NBuffers |
+ * | |
+ * Emit Barrier |
+ * | Init ProcSignal
+ * | |
+ * Finish resize |
+ * | |
+ * New NBuffers Old NBuffers
+ *
+ * In this case the backend is not yet ready to receive a signal from
+ * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal
+ * is initialized even later, after the resizing was finished.
+ *
+ * To address resulting inconsistency, postmaster broadcasts the current
+ * NBuffers value via shared memory. Every new backend has to verify this value
+ * before it will access the buffer pool: if it differs from its own value,
+ * this indicates a shared memory resize has happened and the backend has to
+ * first synchronize with rest of the pack.
+ */
+ShmemControl *ShmemCtrl = NULL;
+
static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size);
static void IpcMemoryDetach(int status, Datum shmaddr);
static void IpcMemoryDelete(int status, Datum shmId);
@@ -946,6 +992,70 @@ AnonymousShmemDetach(int status, Datum arg)
}
}
+/*
+ * Resize all shared memory segments based on the new shared_buffers value (saved
+ * in ShmemCtrl area). The actual segment resizing is done via ftruncate, which
+ * will fail if there is not sufficient space to expand the anon file.
+ *
+ * TODO: Rename this to BufferShmemResize() or something. Only buffer manager's
+ * memory should be resized in this function.
+ *
+ * TODO: This function changes the amount of shared memory used. So it should
+ * also update the show only GUCs shared_memory_size and
+ * shared_memory_size_in_huge_pages in all backends. SetConfigOption() may be
+ * used for that. But it's not clear whether is_reload parameter is safe to use
+ * while resizing is going on; also at what stage it should be done.
+ */
+static bool
+AnonymousShmemResize(int segment_id, MemoryMappingSizes *mapping, bool expanding)
+{
+ Size hugepagesize;
+ AnonShmemData *anonshmem = &AnonShmemInfo[segment_id];
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ elog(DEBUG1, "Resize shmem from %d to %d", NBuffers, NBuffersPending);
+
+ if (anonshmem->fd == -1)
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("segment[%s]: only anonymous (mmaped) file backed segments can be resized",
+ MappingName(segment_id))));
+
+#ifndef MAP_HUGETLB
+ /* PrepareHugePages should have dealt with this case */
+ Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on);
+#else
+ if (huge_pages_on)
+ {
+ Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY);
+ GetHugePageSize(&hugepagesize, NULL, NULL);
+ round_off_mapping_sizes_for_hugepages(mapping, hugepagesize);
+ }
+#endif
+ Assert(anonshmem->addr);
+
+ /*
+ * Size of the reserved address space should not change, since it depends
+ * upon MaxNBuffers, which can be changed only on restart.
+ */
+ Assert(anonshmem->size == mapping->shmem_reserved);
+
+ /*
+ * Resize the backing file to resize the allocated memory, and allocate
+ * more memory on supported platforms if required.
+ */
+ if (ftruncate(anonshmem->fd, mapping->shmem_req_size) == -1)
+ ereport(ERROR,
+ (errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not truncate anonymous file segment for \"%s\": %m",
+ MappingName(segment_id))));
+ if (expanding)
+ shmem_fallocate(anonshmem->fd, MappingName(segment_id), mapping->shmem_req_size, ERROR);
+
+ return true;
+}
+
/*
* PGSharedMemoryCreate
*
@@ -1130,6 +1240,41 @@ PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping,
return anonshmem->addr;
}
+bool
+PGSharedMemoryResize(int segment_id, MemoryMappingSizes *mapping)
+{
+ AnonShmemData *anonshmem = &AnonShmemInfo[segment_id];
+ PGShmemHeader *hdr;
+
+ /* For now, we allow only mmapped memory to be resized. */
+ if (shared_memory_type != SHMEM_TYPE_MMAP || anonshmem->fd == -1)
+ elog(ERROR, "only anonymous (mmaped) file backed memory can be resized");
+
+ /* Anonymous memory has header as the first chunk. */
+ hdr = (PGShmemHeader *) anonshmem->addr;
+ /* Main shared memory segment is always static. */
+ Assert(segment_id != MAIN_SHMEM_SEGMENT);
+
+ /*
+ * We should have reserved enough address space for resizing. PANIC if
+ * that's not the case.
+ */
+ if (hdr->reservedsize < mapping->shmem_req_size)
+ ereport(PANIC,
+ (errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough shared memory is reserved")));
+
+ /* Nothing to do if size is unchanged */
+ if (hdr->totalsize == mapping->shmem_req_size)
+ return true;
+
+ AnonymousShmemResize(segment_id, mapping, mapping->shmem_req_size > hdr->totalsize);
+
+ /* Update the available size. */
+ hdr->totalsize = mapping->shmem_req_size;
+ return true;
+}
+
#ifdef EXEC_BACKEND
/*
@@ -1271,3 +1416,22 @@ PGSharedMemoryDetach(void)
}
}
}
+
+void
+ShmemControlInit(void)
+{
+ bool foundShmemCtrl;
+
+ ShmemCtrl = (ShmemControl *)
+ ShmemInitStruct("Shmem Control", sizeof(ShmemControl),
+ &foundShmemCtrl);
+
+ if (!foundShmemCtrl)
+ {
+ pg_atomic_init_u32(&ShmemCtrl->targetNBuffers, 0);
+ pg_atomic_init_u32(&ShmemCtrl->currentNBuffers, 0);
+ pg_atomic_init_flag(&ShmemCtrl->resize_in_progress);
+
+ ShmemCtrl->coordinator = 0;
+ }
+}
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 034460eab96..3601d13d1cf 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -415,6 +415,12 @@ retry:
"on" : "off", PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
}
+bool
+PGSharedMemoryResize(int segment_id, MemoryMappingSizes *mapping)
+{
+ elog(ERROR, "shared memory resizing is not supported on Windows");
+}
+
/*
* PGSharedMemoryReAttach
*
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index e03c19123bc..53282da5e25 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -659,9 +659,12 @@ CheckpointerMain(const void *startup_data, size_t startup_data_len)
static void
ProcessCheckpointerInterrupts(void)
{
- if (ProcSignalBarrierPending)
- ProcessProcSignalBarrier();
-
+ /*
+ * Reloading config can trigger further signals, complicating interrupts
+ * processing -- so let it run first.
+ *
+ * XXX: Is there any need in memory barrier after ProcessConfigFile?
+ */
if (ConfigReloadPending)
{
ConfigReloadPending = false;
@@ -681,6 +684,9 @@ ProcessCheckpointerInterrupts(void)
UpdateSharedMemoryConfig();
}
+ if (ProcSignalBarrierPending)
+ ProcessProcSignalBarrier();
+
/* Perform logging of memory contexts of this process */
if (LogMemoryContextPending)
ProcessLogMemoryContextInterrupt();
@@ -958,12 +964,13 @@ CheckpointerShmemSize(void)
Size size;
/*
- * The size of the requests[] array is arbitrarily set equal to NBuffers.
- * But there is a cap of MAX_CHECKPOINT_REQUESTS to prevent accumulating
- * too many checkpoint requests in the ring buffer.
+ * The size of the requests[] array is arbitrarily set equal to the
+ * initial size of buffer pool. But there is a cap of
+ * MAX_CHECKPOINT_REQUESTS to prevent accumulating too many checkpoint
+ * requests in the ring buffer.
*/
size = offsetof(CheckpointerShmemStruct, requests);
- size = add_size(size, mul_size(Min(NBuffers,
+ size = add_size(size, mul_size(Min(NBuffersPending,
MAX_CHECKPOINT_REQUESTS),
sizeof(CheckpointerRequest)));
@@ -994,7 +1001,7 @@ CheckpointerShmemInit(void)
*/
MemSet(CheckpointerShmem, 0, size);
SpinLockInit(&CheckpointerShmem->ckpt_lck);
- CheckpointerShmem->max_requests = Min(NBuffers, MAX_CHECKPOINT_REQUESTS);
+ CheckpointerShmem->max_requests = Min(NBuffersPending, MAX_CHECKPOINT_REQUESTS);
CheckpointerShmem->head = CheckpointerShmem->tail = 0;
ConditionVariableInit(&CheckpointerShmem->start_cv);
ConditionVariableInit(&CheckpointerShmem->done_cv);
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index d6133bfebc6..4c95376648d 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -110,11 +110,15 @@
#include "replication/slotsync.h"
#include "replication/walsender.h"
#include "storage/aio_subsys.h"
+#include "storage/bufmgr.h"
#include "storage/fd.h"
#include "storage/io_worker.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
#include "tcop/backend_startup.h"
#include "tcop/tcopprot.h"
#include "utils/datetime.h"
@@ -125,7 +129,6 @@
#ifdef EXEC_BACKEND
#include "common/file_utils.h"
-#include "storage/pg_shmem.h"
#endif
@@ -958,6 +961,11 @@ PostmasterMain(int argc, char *argv[])
*/
InitializeFastPathLocks();
+ /*
+ * Calculate MaxNBuffers for buffer pool resizing.
+ */
+ InitializeMaxNBuffers();
+
/*
* Give preloaded libraries a chance to request additional shared memory.
*/
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index d3c68d8b04c..b8dac7eb9d5 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -101,7 +101,7 @@ AioChooseMaxConcurrency(void)
/* Similar logic to LimitAdditionalPins() */
max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
- max_proportional_pins = NBuffers / max_backends;
+ max_proportional_pins = NBuffersPending / max_backends;
max_proportional_pins = Max(max_proportional_pins, 1);
diff --git a/src/backend/storage/buffer/Makefile b/src/backend/storage/buffer/Makefile
index fd7c40dcb08..3bc9aee85de 100644
--- a/src/backend/storage/buffer/Makefile
+++ b/src/backend/storage/buffer/Makefile
@@ -17,6 +17,7 @@ OBJS = \
buf_table.o \
bufmgr.o \
freelist.o \
- localbuf.o
+ localbuf.o \
+ buf_resize.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..5dd753718a0 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -261,3 +261,77 @@ As of 8.4, background writer starts during recovery mode when there is
some form of potentially extended recovery to perform. It performs an
identical service to normal processing, except that checkpoints it
writes are technically restartpoints.
+
+Resizing shared buffers at runtime
+----------------------------------
+
+Before <TODO: Add version>, the size of the shared buffer pool (i.e. the number
+of shared buffers) was given by the global variable NBuffers and was fixed at
+server start time using GUC shared_buffers. In order to change the size of the
+shared buffer pool, one needed to change the GUC and restart the server.
+Starting <TODO: add version> PostgreSQL supports resizing the buffer pool
+without a server restart. The new GUC variable max_shared_buffers defines the
+maximum size of the shared buffer pool. Existing GUC shared_buffers controls the
+size of the buffer pool at run time. See configure.sgml for more details about
+these GUCs.
+
+The buffer manager maintains following data structures in shared memory.
+1. Buffer Descriptors: An array of BufferDesc structures, one per buffer.
+2. Buffer Blocks: An array of buffer blocks, forming the buffer pool.
+3. Buffer Lookup Table: A hash table mapping a page to the buffer containing
+ that page.
+4. IO conditional variables: An array of conditional variables, one per buffer.
+5. Checkpoint buffer ids: An array of buffer ids used during checkpointing.
+
+Except for the hash table, all the above data structures are required to be
+allocated as contiguous memory chunks. The code also relies on their start
+addresses being stable throughout the server lifetime. Resizing these structures
+means stretching or shrinking their tails. This is done by a. reserving address
+spaces for each of these structures during server start up, b. allocating or
+deallocating physical memory pages as needed during resizing, and c.
+initializing the array elements in the newly allocated memory regions.
+
+TODO: Should the memory management strategy be documented here or somewhere near
+storage/pg_shmem.c or sysv_shmem.c?
+
+Following protocol is used to make each of the segments elastic in a system
+which supports anonymous shared memory segments (like Linux). For each of the
+segments:
+1. Allocate a file descriptor using memfd_create() system call.
+2. Reserve address space for the segment using mmap(), passing it the fd
+ obtained in the first step with PROT_NONE and MAP_NORESERVE flags.
+3. Change the protection of the whole address space to PROT_READ | PROT_WRITE
+ using mprotect().
+4. Allocate or deallocate physical memory pages using ftruncate() as needed.
+
+The first three steps are executed when starting the server. The last step is
+executed when starting the server and during resizing.
+
+We should ideally call mmap() with PROT_READ | PROT_WRITE and avoid calling
+mprotect() separately. However, on Linux, when using huge pages, mmap() with
+PROT_READ | PROT_WRITE allocates memory for the whole region even if
+MAP_NORESERVE is specified. Above sequence avoids this problem.
+
+We could use mmap and mremap to resize the segments. However these calls will
+need to be executed by each backend process as well as the postmaster when
+resizing. This requires additional Synchronization between the postmaster and
+the resizing coordinator and requires careful implementation to avoid missing a
+newly joining backend during resizing. Using memfd_create and ftruncate avoids
+this complexity as ftruncate can be called from only one process, usually the
+coordinator and the changes are visible to all other processes automatically.
+
+Global variables (like NBuffers before this change) are inherited by each
+backend from the postmaster. If we continue to rely on NBuffers to provide the
+size of the buffer pool at a given point in time, we need to involve the
+postmater in resizing. To avoid this complexity, we store the current size of
+the buffer pool in the shared structure ShmemCtrl (name subject to change) and
+change only that variable during resizing. We use ProcSignalBarrier to know when
+all backends have observed the new size after resizing.
+
+See CreateAnonymousSegment() and AnonymousShmemResize() for detailed
+implementation.
+
+To resize shared buffers at runtime a user performs the steps mentioned in the
+description of shared_buffers GUC variable in configure.sgml. Actual resizing
+protocol is documented in the prologue of function pg_resize_shared_buffers() in
+src/backend/storage/buffer/buf_resize.c
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 42112109af9..d57c8bfb209 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,8 +17,8 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
-#include "storage/pg_shmem.h"
#include "storage/proclist.h"
+#include "utils/guc.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -63,6 +63,23 @@ CkptSortItem *CkptBufferIds;
* to allow dynamic resizing of the buffer pool.
*/
+/*
+ * Initialize a single buffer.
+ */
+static void
+InitializeBuffer(int buf_id)
+{
+ BufferDesc *buf = GetBufferDescriptor(buf_id);
+
+ ClearBufferTag(&buf->tag);
+ pg_atomic_init_u64(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ buf->buf_id = buf_id;
+ pgaio_wref_clear(&buf->io_wref);
+ proclist_init(&buf->lock_waiters);
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+}
+
/*
* Initialize shared buffer pool
@@ -77,24 +94,25 @@ BufferManagerShmemInit(void)
foundDescs,
foundIOCV,
foundBufCkpt;
+ int i;
/* Align descriptors to a cacheline boundary. */
BufferDescriptors = (BufferDescPadded *)
ShmemInitStructInSegment("Buffer Descriptors",
- NBuffers * sizeof(BufferDescPadded),
+ MaxNBuffers * sizeof(BufferDescPadded),
&foundDescs, MAIN_SHMEM_SEGMENT);
/* Align buffer pool on IO page size boundary. */
BufferBlocks = (char *)
TYPEALIGN(PG_IO_ALIGN_SIZE,
ShmemInitStructInSegment("Buffer Blocks",
- NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ NBuffersPending * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
&foundBufs, BUFFERS_SHMEM_SEGMENT));
/* Align condition variables to cacheline boundary. */
BufferIOCVArray = (ConditionVariableMinimallyPadded *)
ShmemInitStructInSegment("Buffer IO Condition Variables",
- NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ MaxNBuffers * sizeof(ConditionVariableMinimallyPadded),
&foundIOCV, MAIN_SHMEM_SEGMENT);
/*
@@ -106,46 +124,41 @@ BufferManagerShmemInit(void)
*/
CkptBufferIds = (CkptSortItem *)
ShmemInitStructInSegment("Checkpoint BufferIds",
- NBuffers * sizeof(CkptSortItem), &foundBufCkpt,
+ MaxNBuffers * sizeof(CkptSortItem), &foundBufCkpt,
MAIN_SHMEM_SEGMENT);
if (foundDescs || foundBufs || foundIOCV || foundBufCkpt)
{
/* should find all of these, or none of them */
Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt);
- /* note: this path is only taken in EXEC_BACKEND case */
+
+ /*
+ * note: this path is only taken in EXEC_BACKEND case when
+ * initializing shared memory.
+ */
}
else
{
- int i;
-
/*
* Initialize all the buffer headers.
*/
- for (i = 0; i < NBuffers; i++)
- {
- BufferDesc *buf = GetBufferDescriptor(i);
-
- ClearBufferTag(&buf->tag);
-
- pg_atomic_init_u64(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
-
- buf->buf_id = i;
-
- pgaio_wref_clear(&buf->io_wref);
-
- proclist_init(&buf->lock_waiters);
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
- }
+ for (i = 0; i < NBuffersPending; i++)
+ InitializeBuffer(i);
}
- /* Init other shared buffer-management stuff */
+ /*
+ * Init other shared buffer-management stuff.
+ */
StrategyInitialize(!foundDescs);
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
&backend_flush_after);
+
+ /* Declare the size of current buffer pool. */
+ NBuffers = NBuffersPending;
+ pg_atomic_init_u32(&ShmemCtrl->currentNBuffers, NBuffersPending);
+ pg_atomic_init_u32(&ShmemCtrl->targetNBuffers, NBuffersPending);
}
/*
@@ -154,6 +167,10 @@ BufferManagerShmemInit(void)
* compute the size of shared memory for the buffer pool including
* data pages, buffer descriptors, hash tables, etc.
*
+ * All the data structures except the buffer blocks are maximally allocated to
+ * accomodate a buffer pool of size MaxNBuffers. Only buffer blocks are
+ * allocated based on the initial size of the buffer pool (NBuffersPending).
+ *
* The function adds the amount of required memory for buffer blocks to
* BUFFERS_SHMEM_SEGMENT segment. Amount of memory required for other structures
* is returned.
@@ -165,13 +182,15 @@ BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
/* size of data pages, plus alignment padding */
size = add_size(0, PG_IO_ALIGN_SIZE);
- size = add_size(size, mul_size(NBuffers, BLCKSZ));
+ size = add_size(size, mul_size(NBuffersPending, BLCKSZ));
mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_req_size = size;
+ size = add_size(0, PG_IO_ALIGN_SIZE);
+ size = add_size(size, mul_size(MaxNBuffers, BLCKSZ));
mapping_sizes[BUFFERS_SHMEM_SEGMENT].shmem_reserved = size;
size = 0;
/* size of buffer descriptors */
- size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded)));
+ size = add_size(size, mul_size(MaxNBuffers, sizeof(BufferDescPadded)));
/* to allow aligning buffer descriptors */
size = add_size(size, PG_CACHE_LINE_SIZE);
@@ -179,13 +198,156 @@ BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes)
size = add_size(size, StrategyShmemSize());
/* size of I/O condition variables */
- size = add_size(size, mul_size(NBuffers,
+ size = add_size(size, mul_size(MaxNBuffers,
sizeof(ConditionVariableMinimallyPadded)));
/* to allow aligning the above */
size = add_size(size, PG_CACHE_LINE_SIZE);
/* size of checkpoint sort array in bufmgr.c */
- size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem)));
+ size = add_size(size, mul_size(MaxNBuffers, sizeof(CkptSortItem)));
return size;
}
+
+/*
+ * Resize and reinitialize shared buffer manager structures when resizing the buffer pool.
+ *
+ * This function is called in the backend which coordinates buffer resizing
+ * operation.
+ */
+void
+BufferManagerShmemResize(int currentNBuffers, int targetNBuffers)
+{
+ bool found;
+ int i;
+ void *tmpPtr;
+
+ /*
+ * The only resizable buffer manager structure is the buffer blocks. All
+ * other structures are maximally allocated when the server starts.
+ */
+ tmpPtr = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemResizeStructInSegment("Buffer Blocks",
+ targetNBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &found, BUFFERS_SHMEM_SEGMENT));
+ if (BufferBlocks != tmpPtr || !found)
+ elog(FATAL, "resizing buffer blocks failed: expected pointer %p, got %p, found=%d",
+ BufferBlocks, tmpPtr, found);
+
+ /*
+ * TODO: Initialize the headers for new buffers. If we are shrinking the
+ * buffers, invalidate the extra buffers. If expanding initialize the new
+ * buffers.
+ */
+ for (i = currentNBuffers; i < targetNBuffers; i++)
+ InitializeBuffer(i);
+}
+
+/*
+ * BufferManagerShmemValidate
+ * Validate that buffer manager shared memory structures have correct
+ * pointers and sizes after a resize operation.
+ *
+ * This function is called by backends during ProcessBarrierShmemResizeStruct
+ * to ensure their view of the buffer structures is consistent after memory
+ * remapping.
+ */
+void
+BufferManagerShmemValidate(int targetNBuffers)
+{
+ bool found;
+ void *tmpPtr;
+
+ /* Validate Buffer Descriptors */
+ tmpPtr = (BufferDescPadded *)
+ ShmemInitStructInSegment("Buffer Descriptors",
+ MaxNBuffers * sizeof(BufferDescPadded),
+ &found, MAIN_SHMEM_SEGMENT);
+ if (!found || BufferDescriptors != tmpPtr)
+ elog(FATAL, "validating buffer descriptors failed: expected pointer %p, got %p, found=%d",
+ BufferDescriptors, tmpPtr, found);
+
+ /* Validate Buffer IO Condition Variables */
+ tmpPtr = (ConditionVariableMinimallyPadded *)
+ ShmemInitStructInSegment("Buffer IO Condition Variables",
+ MaxNBuffers * sizeof(ConditionVariableMinimallyPadded),
+ &found, MAIN_SHMEM_SEGMENT);
+ if (!found || BufferIOCVArray != tmpPtr)
+ elog(FATAL, "validating buffer IO condition variables failed: expected pointer %p, got %p, found=%d",
+ BufferIOCVArray, tmpPtr, found);
+
+ /* Validate Checkpoint BufferIds */
+ tmpPtr = (CkptSortItem *)
+ ShmemInitStructInSegment("Checkpoint BufferIds",
+ MaxNBuffers * sizeof(CkptSortItem), &found,
+ MAIN_SHMEM_SEGMENT);
+ if (!found || CkptBufferIds != tmpPtr)
+ elog(FATAL, "validating checkpoint buffer IDs failed: expected pointer %p, got %p, found=%d",
+ CkptBufferIds, tmpPtr, found);
+
+ /* Validate Buffer Blocks */
+ tmpPtr = (char *)
+ TYPEALIGN(PG_IO_ALIGN_SIZE,
+ ShmemInitStructInSegment("Buffer Blocks",
+ targetNBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE,
+ &found, BUFFERS_SHMEM_SEGMENT));
+ if (!found || BufferBlocks != tmpPtr)
+ elog(FATAL, "validating buffer blocks failed: expected pointer %p, got %p, found=%d",
+ BufferBlocks, tmpPtr, found);
+}
+
+/*
+ * check_shared_buffers
+ * GUC check_hook for shared_buffers
+ *
+ * When reloading the configuration, shared_buffers should not be set to a value
+ * higher than max_shared_buffers fixed at the boot time.
+ */
+bool
+check_shared_buffers(int *newval, void **extra, GucSource source)
+{
+ if (finalMaxNBuffers && *newval > MaxNBuffers)
+ {
+ GUC_check_errdetail("\"shared_buffers\" must be less than \"max_shared_buffers\".");
+ return false;
+ }
+ return true;
+}
+
+/*
+ * show_shared_buffers
+ * GUC show_hook for shared_buffers
+ *
+ * Shows both current and pending buffer counts with proper unit formatting.
+ */
+const char *
+show_shared_buffers(void)
+{
+ static char buffer[128];
+ int64 current_value,
+ pending_value;
+ const char *current_unit,
+ *pending_unit;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+
+ if (currentNBuffers == NBuffersPending)
+ {
+ /* No buffer pool resizing pending. */
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s", current_value, current_unit);
+ }
+ else
+ {
+ /*
+ * Shared buffer pool is pending to be resized, show both current and
+ * pending sizes.
+ */
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ convert_int_from_base_unit(NBuffersPending, GUC_UNIT_BLOCKS, &pending_value, &pending_unit);
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s (pending: " INT64_FORMAT "%s)",
+ current_value, current_unit, pending_value, pending_unit);
+ }
+
+ return buffer;
+}
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
new file mode 100644
index 00000000000..7b6d46c05ba
--- /dev/null
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -0,0 +1,455 @@
+/*-------------------------------------------------------------------------
+ *
+ * buf_resize.c
+ * shared buffer pool resizing functionality
+ *
+ * This module contains the implementation of shared buffer pool resizing,
+ * including the main resize coordination function and barrier processing
+ * functions that synchronize all backends during resize operations.
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/storage/buffer/buf_resize.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "miscadmin.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
+#include "utils/injection_point.h"
+
+
+/*
+ * Prepare ShmemCtrl for resizing the shared buffer pool.
+ */
+static void
+MarkBufferResizingStart(int targetNBuffers, int currentNBuffers)
+{
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ Assert(pg_atomic_read_u32(&ShmemCtrl->currentNBuffers) == currentNBuffers);
+
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, targetNBuffers);
+ ShmemCtrl->coordinator = MyProcPid;
+}
+
+/*
+ * Reset ShmemCtrl after resizing the shared buffer pool is done.
+ */
+static void
+MarkBufferResizingEnd(int newNBuffers)
+{
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ Assert(pg_atomic_read_u32(&ShmemCtrl->currentNBuffers) == newNBuffers);
+
+ /*
+ * TODO: should we leave targetNBuffers as is? We are setting it to
+ * NBuffers in BufferManagerShmemInit().
+ */
+ pg_atomic_write_u32(&ShmemCtrl->targetNBuffers, 0);
+ ShmemCtrl->coordinator = -1;
+}
+
+/*
+ * Communicate given buffer pool resize barrier to all other backends and the Postmaster.
+ *
+ * ProcSignalBarrier is not sent to the Postmaster but we need the Postmaster to
+ * update its knowledge about the buffer pool so that it can be inherited by the
+ * child processes.
+ */
+static void
+SharedBufferResizeBarrier(ProcSignalBarrierType barrier, const char *barrier_name)
+{
+ WaitForProcSignalBarrier(EmitProcSignalBarrier(barrier));
+ elog(LOG, "all backends acknowledged %s barrier", barrier_name);
+
+#ifdef USE_INJECTION_POINTS
+ /* Injection point specific to this barrier type */
+ switch (barrier)
+ {
+ case PROCSIGNAL_BARRIER_SHBUF_SHRINK:
+ INJECTION_POINT("pgrsb-shrink-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM:
+ INJECTION_POINT("pgrsb-resize-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_EXPAND:
+ INJECTION_POINT("pgrsb-expand-barrier-sent", NULL);
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED:
+ /* TODO: Add an injection point here. */
+ break;
+ case PROCSIGNAL_BARRIER_SMGRRELEASE:
+ case PROCSIGNAL_BARRIER_UPDATE_XLOG_LOGICAL_INFO:
+
+ /*
+ * Not relevant in this function but it's here so that the
+ * compiler can detect any missing shared buffer resizing barrier
+ * enum here.
+ */
+ break;
+ }
+#endif /* USE_INJECTION_POINTS */
+}
+
+/*
+ * C implementation of SQL interface to update the shared buffers according to
+ * the current values of shared_buffers GUCs.
+ *
+ * The current boundaries of the buffer pool are given by two ranges.
+ *
+ * - [1, StrategyControl::activeNBuffers] is the range of buffers from which new
+ * allocations can happen at any time.
+ *
+ * - [1, ShmemCtrl::currentNBuffers] is the range of valid buffers at any given
+ * time.
+ *
+ * Let's assume that before resizing, the number of buffers in the buffer pool is
+ * NBuffersOld. After resizing it is NBuffersNew. Before resizing
+ * StrategyControl::activeNBuffers == ShmemCtrl::currentNBuffers == NBuffersOld.
+ * After the resizing finishes StrategyControl::activeNBuffers ==
+ * ShmemCtrl::currentNBuffers == NBuffersNew. Thus when no resizing happens these
+ * two ranges are same.
+ *
+ * Following steps are performed by the coordinator during resizing.
+ *
+ * 1. Marks resizing in progress to avoid multiple concurrent invocations of this
+ * function.
+ *
+ * 2. When shrinking the shared buffer pool, the coordinator sends SHBUF_SHRINK
+ * ProcSignalBarrier. In response to this barrier background writer is expected
+ * to set StrategyControl::activeNBuffers = NBuffersNew to restrict the new
+ * buffer allocations only to the new buffer pool size and also reset its
+ * internal state. Once every backend has acknowledged the barrier, the
+ * coordinator can be sure that new allocations will not happen in the buffer
+ * pool area being shrunk. Then it evicts the buffers in that area. Note that
+ * ShmemCtrl::currentNBuffers is still NBuffersOld, since backend may still
+ * access buffers allocated before the resizing started. Buffer eviction may fail
+ * if a buffer being evicted is pinned and the resizing operatino is aborted.
+ * Once the eviction is finished, the extra memory can be freed in the next step.
+ *
+ * 2. This step is executed in both cases, when expanding the buffer pool or
+ * shrinking the buffer pool. The anonymous file backing each of the shared
+ * memory segment containg the buffer pool shared data structures is resized to
+ * the amount of memory required for the new buffer pool size. When expanding the
+ * expanded portion of memory is initialized appropriately.
+ * ShmemCtrl::currentNBuffers is set to NBuffersNew to indicate new range of
+ * valid shared buffers. Every backend is sent SHBUF_RESIZE_MAP_AND_MEM barrier.
+ * All the backends validate that their pointers to the shared buffers structure
+ * are valid and have the right size. Once every backend has acknowledged the
+ * barrier, this step finishes.
+ *
+ * 3. When expanding the buffer pool, the coordinator sends SHBUF_EXPAND barrier
+ * to signal end of expansion. When expadning the background writer, in response
+ * to StrategyControl::activeNBuffers = NBufferNew so that new allocations can
+ * use expanded range of buffer pool.
+ *
+ * TODO: Handle the case when the backend executing this function dies or the
+ * query is cancelled or it hits an error while resizing.
+ */
+Datum
+pg_resize_shared_buffers(PG_FUNCTION_ARGS)
+{
+ bool result = true;
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+ int targetNBuffers = NBuffersPending;
+ MemoryMappingSizes mapping_sizes[NUM_MEMORY_MAPPINGS];
+
+ if (currentNBuffers == targetNBuffers)
+ {
+ elog(LOG, "shared buffers are already at %d, no need to resize", currentNBuffers);
+ PG_RETURN_BOOL(true);
+ }
+
+ if (!pg_atomic_test_set_flag(&ShmemCtrl->resize_in_progress))
+ {
+ elog(LOG, "shared buffer resizing already in progress");
+ PG_RETURN_BOOL(false);
+ }
+
+ /*
+ * TODO: NBuffersPending may change after it was sampled above, thus
+ * leading to wrong memory size estimates. Find a way to pass
+ * targetNBuffers value to BufferManagerShmemSize().
+ */
+ BufferManagerShmemSize(mapping_sizes);
+ /* Round it off to a multiple of a typical page size */
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ /* Structures in main memory segment are never resized. */
+ if (i == MAIN_SHMEM_SEGMENT)
+ continue;
+
+ round_off_mapping_sizes(&mapping_sizes[i]);
+ }
+
+ /*
+ * TODO: What if the NBuffersPending value seen here is not the desired
+ * one because somebody did a pg_reload_conf() between the last
+ * pg_reload_conf() and execution of this function?
+ */
+ MarkBufferResizingStart(targetNBuffers, currentNBuffers);
+ elog(LOG, "resizing shared buffers from %d to %d", currentNBuffers, targetNBuffers);
+
+ INJECTION_POINT("pg-resize-shared-buffers-flag-set", NULL);
+
+ /* Phase 1: SHBUF_SHRINK - Only for shrinking buffer pool */
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * Phase 1: Shrinking - send SHBUF_SHRINK barrier Every backend sets
+ * activeNBuffers = NewNBuffers to restrict buffer pool allocations to
+ * the new size
+ */
+ elog(LOG, "Phase 1: Shrinking buffer pool, restricting allocations to %d buffers", targetNBuffers);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_SHRINK, CppAsString(PROCSIGNAL_BARRIER_SHBUF_SHRINK));
+
+ /* Evict buffers in the area being shrunk */
+ elog(LOG, "evicting buffers %u..%u", targetNBuffers + 1, currentNBuffers);
+ if (!EvictExtraBuffers(targetNBuffers, currentNBuffers))
+ {
+ elog(WARNING, "failed to evict extra buffers during shrinking");
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED, CppAsString(PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED));
+ MarkBufferResizingEnd(currentNBuffers);
+ pg_atomic_clear_flag(&ShmemCtrl->resize_in_progress);
+ PG_RETURN_BOOL(false);
+ }
+
+ /*
+ * Shrink buffer manager structures before shrinking the shared
+ * memory.
+ */
+ BufferManagerShmemResize(currentNBuffers, targetNBuffers);
+
+ /* Update the current NBuffers. */
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, targetNBuffers);
+ }
+
+ /* Phase 2: SHBUF_RESIZE_MAP_AND_MEM - Both expanding and shrinking */
+ elog(LOG, "Phase 2: Remapping shared memory segments and updating structures");
+ for (int i = 0; i < NUM_MEMORY_MAPPINGS; i++)
+ {
+ /* Structures in the main memory segment are never resized. */
+ if (i == MAIN_SHMEM_SEGMENT)
+ continue;
+
+ if (!PGSharedMemoryResize(i, &mapping_sizes[i]))
+ {
+ /*
+ * This should never fail since address map should already be
+ * reserved. So the failure should be treated as PANIC.
+ */
+ elog(PANIC, "failed to resize anonymous shared memory");
+ }
+ }
+
+ INJECTION_POINT("pgrsb-after-shmem-resize", NULL);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM, CppAsString(PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM));
+
+ /* Phase 3: SHBUF_EXPAND - Only for expanding buffer pool */
+ if (targetNBuffers > currentNBuffers)
+ {
+ /* Expand buffer manager structures after expanding the shared memory. */
+ BufferManagerShmemResize(currentNBuffers, targetNBuffers);
+
+ /*
+ * Phase 3: Expanding - send SHBUF_EXPAND barrier Backends set
+ * activeNBuffers = NewNBuffers and start allocating buffers from the
+ * expanded range
+ */
+ elog(LOG, "Phase 3: Expanding buffer pool, enabling allocations up to %d buffers", targetNBuffers);
+ pg_atomic_write_u32(&ShmemCtrl->currentNBuffers, targetNBuffers);
+
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_EXPAND, CppAsString(PROCSIGNAL_BARRIER_SHBUF_EXPAND));
+ }
+
+ /*
+ * Reset buffer resize control area.
+ */
+ MarkBufferResizingEnd(targetNBuffers);
+
+ pg_atomic_clear_flag(&ShmemCtrl->resize_in_progress);
+
+ elog(LOG, "successfully resized shared buffers to %d", targetNBuffers);
+
+ PG_RETURN_BOOL(result);
+}
+
+bool
+ProcessBarrierShmemShrink(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * Delay adjusting the new active size of buffer pool till this process
+ * becomes ready to resize buffers.
+ */
+ if (delay_shmem_resize)
+ {
+ elog(LOG, "Phase 1: Delaying SHBUF_SHRINK barrier - restricting allocations to %d buffers, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+
+ return false;
+ }
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation
+ * statistics and the strategy control together so that background
+ * writer doesn't go out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody
+ * would reset the strategy control area. So we can't rely on
+ * background worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, targetNBuffers);
+ /* Reset strategy control to new size */
+ StrategyReset(targetNBuffers);
+ }
+
+ elog(LOG, "Phase 1: Processing SHBUF_SHRINK barrier - target buffer pool size = %d, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemResizeMapAndMem(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * If buffer pool is being shrunk, we are already working with a smaller
+ * buffer pool, so shrinking address space and shared structures should
+ * not be a problem. When expanding, expanding the address space and
+ * shared structures beyond the current boundaries is not going to be a
+ * problem since we are not accessing that memory yet. So there is no
+ * reason to delay processing this barrier.
+ */
+
+ /*
+ * Coordinator has already adjusted its address map and also updated sizes
+ * of the shared buffer structures, no further validation needed.
+ */
+ if (ShmemCtrl->coordinator == MyProcPid)
+ return true;
+
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * When shrinking, shared data structures have been resized at this
+ * point. Validate that their pointers to shared buffer structures
+ * are still valid and have the correct size after resizing.
+ *
+ * TODO: Do want to do this only in assert enabled builds?
+ */
+ BufferManagerShmemValidate(targetNBuffers);
+ elog(LOG, "Backend %d successfully validated structure pointers after resize", MyProcPid);
+ }
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemExpand(void)
+{
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ /*
+ * Delay adjusting the new active size of buffer pool till this process
+ * becomes ready to resize buffers.
+ */
+ if (delay_shmem_resize)
+ {
+ elog(LOG, "Phase 3: delaying SHBUF_EXPAND barrier - enabling allocations up to %d buffers, coordinator is %d",
+ targetNBuffers, ShmemCtrl->coordinator);
+ return false;
+ }
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation
+ * statistics and the strategy control together so that background
+ * writer doesn't go out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody
+ * would reset the strategy control area. So we can't rely on
+ * background worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, targetNBuffers);
+ StrategyReset(targetNBuffers);
+ }
+
+ /*
+ * Shared data structures must have been resized by now. Validate that
+ * their pointers to shared buffer structures are still valid and have the
+ * correct size after resizing.
+ *
+ * TODO: Do want to do this only in assert enabled builds?
+ */
+ BufferManagerShmemValidate(targetNBuffers);
+ elog(LOG, "Backend %d successfully validated structure pointers after resize", MyProcPid);
+
+ elog(LOG, "Phase 3: Processing SHBUF_EXPAND barrier - targetNBuffers = %d, ShmemCtrl->coordinator = %d", targetNBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
+
+bool
+ProcessBarrierShmemResizeFailed(void)
+{
+ int currentNBuffers = pg_atomic_read_u32(&ShmemCtrl->currentNBuffers);
+ int targetNBuffers = pg_atomic_read_u32(&ShmemCtrl->targetNBuffers);
+
+ Assert(!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress));
+
+ if (MyBackendType == B_BG_WRITER)
+ {
+ /*
+ * We have to reset the background writer's buffer allocation
+ * statistics and the strategy control together so that background
+ * writer doesn't go out of sync with ClockSweepTick().
+ *
+ * TODO: But in case the background writer is not running, nobody
+ * would reset the strategy control area. So we can't rely on
+ * background worker to do that. So find a better way.
+ */
+ BgBufferSyncReset(NBuffers, currentNBuffers);
+ /* Reset strategy control to new size */
+ StrategyReset(currentNBuffers);
+ }
+
+ elog(LOG, "received proc signal indicating failure to resize shared buffers from %d to %d, restoring to %d, coordinator is %d",
+ currentNBuffers, targetNBuffers, currentNBuffers, ShmemCtrl->coordinator);
+
+ return true;
+}
+
+/*
+ * TODO: add progress report facility if required.
+ */
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index a33786a460b..0dfdffb25c8 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -41,7 +41,8 @@ static HTAB *SharedBufHash;
/*
* Estimate space needed for mapping hashtable
- * size is the desired hash table size (possibly more than NBuffers)
+ * size is the desired hash table size (possibly more than the size of buffer
+ * pool)
*/
Size
BufTableShmemSize(int size)
@@ -65,6 +66,13 @@ InitBufTable(int size)
info.entrysize = sizeof(BufferLookupEnt);
info.num_partitions = NUM_BUFFER_PARTITIONS;
+ /*
+ * The shared buffer look up table is set up only once with maximum
+ * possible entries considering maximum size of the buffer pool. It is not
+ * resized after that even if the buffer pool is resized. Hence it is
+ * allocated in the main shared memory segment and not in a resizeable
+ * shared memory segment.
+ */
SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table",
size, size,
&info,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 7241477cac0..be4cd69f2b4 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -57,6 +57,7 @@
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proc.h"
#include "storage/proclist.h"
#include "storage/read_stream.h"
@@ -223,11 +224,11 @@ static BufferDesc *PinCountWaitBuf = NULL;
* and, if so, in what mode.
*
*
- * To avoid - as we used to - requiring an array with NBuffers entries to keep
- * track of local buffers, we use a small sequentially searched array
- * (PrivateRefCountArrayKeys, with the corresponding data stored in
- * PrivateRefCountArray) and an overflow hash table (PrivateRefCountHash) to
- * keep track of backend local pins.
+ * To avoid - as we used to - requiring an array, with as many entries as the
+ * size of buffer pool, to keep track of local buffers, we use a small
+ * sequentially searched array (PrivateRefCountArrayKeys, with the corresponding
+ * data stored in PrivateRefCountArray) and an overflow hash table
+ * (PrivateRefCountHash) to keep track of backend local pins.
*
* Until no more than REFCOUNT_ARRAY_ENTRIES buffers are pinned at once, all
* refcounts are kept track of in the array; after that, new array entries
@@ -3523,7 +3524,7 @@ BufferSync(int flags)
set_bits, 0,
0);
- /* Check for barrier events in case NBuffers is large. */
+ /* Check for barrier events in case the buffer pool is large. */
if (ProcSignalBarrierPending)
ProcessProcSignalBarrier();
}
@@ -3720,6 +3721,32 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_DONE(NBuffers, num_written, num_to_scan);
}
+/*
+ * Information saved between BgBufferSync() calls so we can determine the
+ * strategy point's advance rate and avoid scanning already-cleaned buffers. The
+ * variables are global instead of static local so that BgBufferSyncReset() can
+ * adjust it when resizing shared buffers.
+ */
+static bool saved_info_valid = false;
+static int prev_strategy_buf_id;
+static uint32 prev_strategy_passes;
+static int next_to_clean;
+static uint32 next_passes;
+
+/* Moving averages of allocation rate and clean-buffer density */
+static float smoothed_alloc = 0;
+static float smoothed_density = 10.0;
+
+void
+BgBufferSyncReset(int currentNBuffers, int targetNBuffers)
+{
+ saved_info_valid = false;
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer status after resizing buffers from %d to %d",
+ currentNBuffers, targetNBuffers);
+#endif
+}
+
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
@@ -3739,20 +3766,6 @@ BgBufferSync(WritebackContext *wb_context)
uint32 strategy_passes;
uint32 recent_alloc;
- /*
- * Information saved between calls so we can determine the strategy
- * point's advance rate and avoid scanning already-cleaned buffers.
- */
- static bool saved_info_valid = false;
- static int prev_strategy_buf_id;
- static uint32 prev_strategy_passes;
- static int next_to_clean;
- static uint32 next_passes;
-
- /* Moving averages of allocation rate and clean-buffer density */
- static float smoothed_alloc = 0;
- static float smoothed_density = 10.0;
-
/* Potentially these could be tunables, but for now, not */
float smoothing_samples = 16;
float scan_whole_pool_milliseconds = 120000.0;
@@ -3775,6 +3788,25 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ /*
+ * If buffer pool is being shrunk the buffer being written out may not
+ * remain valid. If the buffer pool is being expanded, more buffers will
+ * become available without even this function writing out any. Hence wait
+ * till buffer resizing finishes i.e. go into hibernation mode.
+ *
+ * TODO: We may not need this synchronization if background worker itself
+ * becomes the coordinator.
+ */
+ if (!pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress))
+ return true;
+
+ /*
+ * Resizing shared buffers while this function is performing an LRU scan
+ * on them may lead to wrong results. Indicate that the resizing should
+ * wait for the LRU scan to complete.
+ */
+ delay_shmem_resize = true;
+
/*
* Find out where the clock-sweep currently is, and how many buffer
* allocations have happened since our last call.
@@ -3792,6 +3824,7 @@ BgBufferSync(WritebackContext *wb_context)
if (bgwriter_lru_maxpages <= 0)
{
saved_info_valid = false;
+ delay_shmem_resize = false;
return true;
}
@@ -3951,8 +3984,17 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
- while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
+ /*
+ * Execute the LRU scan.
+ *
+ * If buffer pool is being shrunk, the buffer being written may not remain
+ * valid. If the buffer pool is being expanded, more buffers will become
+ * available without even this function writing any. Hence stop what we
+ * are doing. This also unblocks other processes that are waiting for
+ * buffer resizing to finish.
+ */
+ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est &&
+ !pg_atomic_unlocked_test_flag(&ShmemCtrl->resize_in_progress))
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
@@ -4011,6 +4053,9 @@ BgBufferSync(WritebackContext *wb_context)
#endif
}
+ /* Let the resizing commence. */
+ delay_shmem_resize = false;
+
/* Return true if OK to hibernate */
return (bufs_to_lap == 0 && recent_alloc == 0);
}
@@ -4341,7 +4386,29 @@ DebugPrintBufferRefcount(Buffer buffer)
void
CheckPointBuffers(int flags)
{
+ /*
+ * Mark that buffer sync is in progress - delay any shared memory
+ * resizing.
+ */
+ /*
+ * TODO: We need to assess whether we should allow checkpoint and buffer
+ * resizing to run in parallel. When expanding buffers it may be fine to
+ * let the checkpointer run in RESIZE_MAP_AND_MEM phase but delay phase
+ * EXPAND phase till the checkpoint finishes, at the same time not allow
+ * checkpoint to run during expansion phase. When shrinking the buffers,
+ * we should delay SHRINK phase till checkpoint finishes and not allow to
+ * start checkpoint till SHRINK phase is done, but allow it to run in
+ * RESIZE_MAP_AND_MEM phase. This needs careful analysis and testing.
+ */
+ delay_shmem_resize = true;
+
BufferSync(flags);
+
+ /*
+ * Mark that buffer sync is no longer in progress - allow shared memory
+ * resizing
+ */
+ delay_shmem_resize = false;
}
/*
@@ -8538,3 +8605,70 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ */
+bool
+EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
+{
+ bool result = true;
+
+ Assert(targetNBuffers < currentNBuffers);
+
+ /*
+ * If the buffer being evicated is locked, this function will need to
+ * wait. This function should not be called from a Postmaster since it can
+ * not wait on a lock.
+ */
+ Assert(IsUnderPostmaster);
+
+ /*
+ * TODO: Before evicting any buffer, we should check whether any of the
+ * buffers are pinned. If we find that a buffer is pinned after evicting
+ * most of them, that will impact performance since all those evicted
+ * buffers might need to be read again.
+ */
+ for (Buffer buf = targetNBuffers + 1; buf <= currentNBuffers; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint64 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u64(&desc->state);
+
+ /*
+ * Nobody is expected to touch the buffers while resizing is going one
+ * hence unlocked precheck should be safe and saves some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ /*
+ * XXX: Looks like CurrentResourceOwner can be NULL here, find another
+ * one in that case?
+ */
+ if (CurrentResourceOwner)
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 403890055be..6221e035024 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -33,10 +33,16 @@ typedef struct
/* Spinlock: protects the values below */
slock_t buffer_strategy_lock;
+ /*
+ * Number of active buffers that can be allocated. During buffer resizing,
+ * this may be different from the actual size of the buffer pool.
+ */
+ pg_atomic_uint32 activeNBuffers;
+
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo NBuffers.
+ * get an actual buffer, it needs to be used modulo activeNBuffers.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -101,21 +107,27 @@ static inline uint32
ClockSweepTick(void)
{
uint32 victim;
+ int activeBuffers;
/*
* Atomically move hand ahead one buffer - if there's several processes
* doing this, this can lead to buffers being returned slightly out of
- * apparent order.
+ * apparent order. We need to read both the current position of hand and
+ * the current buffer allocation limit together consistently. They may be
+ * reset by concurrent resize.
*/
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
victim =
pg_atomic_fetch_add_u32(&StrategyControl->nextVictimBuffer, 1);
+ activeBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
- if (victim >= NBuffers)
+ if (victim >= activeBuffers)
{
uint32 originalVictim = victim;
/* always wrap what we look up in BufferDescriptors */
- victim = victim % NBuffers;
+ victim = victim % activeBuffers;
/*
* If we're the one that just caused a wraparound, force
@@ -143,7 +155,7 @@ ClockSweepTick(void)
*/
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
- wrapped = expected % NBuffers;
+ wrapped = expected % activeBuffers;
success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
&expected, wrapped);
@@ -228,7 +240,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
pg_atomic_fetch_add_u32(&StrategyControl->numBufferAllocs, 1);
/* Use the "clock sweep" algorithm to find a free buffer */
- trycounter = NBuffers;
+ trycounter = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+
for (;;)
{
uint64 old_buf_state;
@@ -281,7 +294,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
local_buf_state))
{
- trycounter = NBuffers;
+ trycounter = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
break;
}
}
@@ -323,10 +336,12 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
{
uint32 nextVictimBuffer;
int result;
+ uint32 activeNBuffers;
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
- result = nextVictimBuffer % NBuffers;
+ activeNBuffers = pg_atomic_read_u32(&StrategyControl->activeNBuffers);
+ result = nextVictimBuffer % activeNBuffers;
if (complete_passes)
{
@@ -336,7 +351,7 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
* Additionally add the number of wraparounds that happened before
* completePasses could be incremented. C.f. ClockSweepTick().
*/
- *complete_passes += nextVictimBuffer / NBuffers;
+ *complete_passes += nextVictimBuffer / activeNBuffers;
}
if (num_buf_alloc)
@@ -383,7 +398,7 @@ StrategyShmemSize(void)
Size size = 0;
/* size of lookup hash table ... see comment in StrategyInitialize */
- size = add_size(size, BufTableShmemSize(NBuffers + NUM_BUFFER_PARTITIONS));
+ size = add_size(size, BufTableShmemSize(MaxNBuffers + NUM_BUFFER_PARTITIONS));
/* size of the shared replacement strategy control block */
size = add_size(size, MAXALIGN(sizeof(BufferStrategyControl)));
@@ -391,6 +406,32 @@ StrategyShmemSize(void)
return size;
}
+void
+StrategyReset(int activeNBuffers)
+{
+ Assert(StrategyControl);
+
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ /* Update the active buffer count for the strategy */
+ pg_atomic_write_u32(&StrategyControl->activeNBuffers, activeNBuffers);
+
+ /* Reset the clock-sweep pointer to start from beginning */
+ pg_atomic_write_u32(&StrategyControl->nextVictimBuffer, 0);
+
+ /*
+ * The statistics is viewed in the context of the number of shared
+ * buffers. Reset it as the size of active number of shared buffers
+ * changes.
+ */
+ StrategyControl->completePasses = 0;
+ pg_atomic_write_u32(&StrategyControl->numBufferAllocs, 0);
+
+ /* TODO: Do we need to seset background writer notifications? */
+ StrategyControl->bgwprocno = -1;
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+}
+
/*
* StrategyInitialize -- initialize the buffer cache replacement
* strategy.
@@ -408,12 +449,21 @@ StrategyInitialize(bool init)
*
* Since we can't tolerate running out of lookup table entries, we must be
* sure to specify an adequate table size here. The maximum steady-state
- * usage is of course NBuffers entries, but BufferAlloc() tries to insert
- * a new entry before deleting the old. In principle this could be
- * happening in each partition concurrently, so we could need as many as
- * NBuffers + NUM_BUFFER_PARTITIONS entries.
+ * usage is of course is as many number of entries as the number of
+ * buffers in the buffer pool. Right now there is no way to free shared
+ * memory. Even if we shrink the buffer lookup table when shrinking the
+ * buffer pool the unused hash table entries can not be freed. When we
+ * expand the buffer pool, more entries can be allocated but we can not
+ * resize the hash table directory without rehashing all the entries. Just
+ * allocating more entries will lead to more contention. Hence we setup
+ * the buffer lookup table considering the maximum possible size of the
+ * buffer pool which is MaxNBuffers.
+ *
+ * Additionally BufferAlloc() tries to insert a new entry before deleting
+ * the old. In principle this could be happening in each partition
+ * concurrently, so we need extra NUM_BUFFER_PARTITIONS entries.
*/
- InitBufTable(NBuffers + NUM_BUFFER_PARTITIONS);
+ InitBufTable(MaxNBuffers + NUM_BUFFER_PARTITIONS);
/*
* Get or create the shared strategy control block
@@ -432,6 +482,8 @@ StrategyInitialize(bool init)
SpinLockInit(&StrategyControl->buffer_strategy_lock);
+ /* Initialize the active buffer count */
+ pg_atomic_init_u32(&StrategyControl->activeNBuffers, NBuffersPending);
/* Initialize the clock-sweep pointer */
pg_atomic_init_u32(&StrategyControl->nextVictimBuffer, 0);
@@ -669,12 +721,25 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
+ * If the slot hasn't been filled yet or the buffer in the slot has been
+ * invalidated when buffer pool was shrunk, tell the caller to allocate a
+ * new buffer with the normal allocation strategy. He will then fill this
* slot by calling AddBufferToRing with the new buffer.
+ *
+ * TODO: buffer ids in the ring will never be greater than the size of
+ * buffer pool, except maybe the first time ring is accessed after
+ * shrinking the buffer pool. Checking the upper bound on buffer id always
+ * may mask a bug bugs that introduces buffer ids higher than the size of
+ * buffer pool in the ring. But performing that check only once after
+ * shrinking seems impossible. The BufferAccessStrategy objects are not
+ * accessible outside the ScanState. Hence we can not purge the buffers
+ * while evicting the buffers. After the resizing is finished, it's not
+ * possible to notice when we touch the first of those objects and the
+ * last of objects. See if this can fixed.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer ||
+ bufnum > pg_atomic_read_u32(&StrategyControl->activeNBuffers))
return NULL;
buf = GetBufferDescriptor(bufnum - 1);
diff --git a/src/backend/storage/buffer/meson.build b/src/backend/storage/buffer/meson.build
index ed84bf08971..f219e29d5ef 100644
--- a/src/backend/storage/buffer/meson.build
+++ b/src/backend/storage/buffer/meson.build
@@ -6,4 +6,5 @@ backend_sources += files(
'bufmgr.c',
'freelist.c',
'localbuf.c',
+ 'buf_resize.c',
)
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 6d2c4520c43..a65436bd7e5 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -156,6 +156,14 @@ CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
size = add_size(size, WaitLSNShmemSize());
size = add_size(size, LogicalDecodingCtlShmemSize());
+ /*
+ * XXX: For some reason slightly more memory is needed for larger
+ * shared_buffers, but this size is enough for any large value I've tested
+ * with. Is it a mistake in how slots are split, or there was a hidden
+ * inconsistency in shmem calculation?
+ */
+ size = add_size(size, 1024 * 1024 * 100);
+
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -170,8 +178,7 @@ CalculateShmemSize(MemoryMappingSizes *mapping_sizes)
/* might as well round it off to a multiple of a typical page size */
for (int segment = 0; segment < NUM_MEMORY_MAPPINGS; segment++)
{
- mapping_sizes[segment].shmem_req_size = add_size(mapping_sizes[segment].shmem_req_size, 8192 - (mapping_sizes[segment].shmem_req_size % 8192));
- mapping_sizes[segment].shmem_reserved = add_size(mapping_sizes[segment].shmem_reserved, 8192 - (mapping_sizes[segment].shmem_reserved % 8192));
+ round_off_mapping_sizes(&mapping_sizes[segment]);
/* Compute the total size of all segments */
size = size + mapping_sizes[segment].shmem_req_size;
}
@@ -323,6 +330,8 @@ CreateOrAttachShmemStructs(void)
CommitTsShmemInit();
SUBTRANSShmemInit();
MultiXactShmemInit();
+ /* TODO: This should be part of BufferManagerShmemInit() */
+ ShmemControlInit();
BufferManagerShmemInit();
/*
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 8e56922dcea..df9bb7a0f1a 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -24,9 +24,11 @@
#include "port/pg_bitutils.h"
#include "replication/logicalworker.h"
#include "replication/walsender.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
#include "storage/sinval.h"
#include "storage/smgr.h"
@@ -109,6 +111,10 @@ static bool CheckProcSignal(ProcSignalReason reason);
static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
+#ifdef DEBUG_SHMEM_RESIZE
+bool delay_proc_signal_init = false;
+#endif
+
/*
* ProcSignalShmemSize
* Compute space needed for ProcSignal's shared memory
@@ -170,6 +176,44 @@ ProcSignalInit(const uint8 *cancel_key, int cancel_key_len)
uint32 old_pss_pid;
Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH);
+
+#ifdef DEBUG_SHMEM_RESIZE
+
+ /*
+ * Introduced for debugging purposes. You can change the variable at
+ * runtime using gdb, then start new backends with delayed ProcSignal
+ * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt
+ * needed for testing. Taken from pg_sleep;
+ */
+ if (delay_proc_signal_init)
+ {
+#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0)
+ float8 endtime = GetNowFloat() + 5;
+
+ for (;;)
+ {
+ float8 delay;
+ long delay_ms;
+
+ CHECK_FOR_INTERRUPTS();
+
+ delay = endtime - GetNowFloat();
+ if (delay >= 600.0)
+ delay_ms = 600000;
+ else if (delay > 0.0)
+ delay_ms = (long) (delay * 1000.0);
+ else
+ break;
+
+ (void) WaitLatch(MyLatch,
+ WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+ delay_ms,
+ WAIT_EVENT_PG_SLEEP);
+ ResetLatch(MyLatch);
+ }
+ }
+#endif
+
if (MyProcNumber < 0)
elog(ERROR, "MyProcNumber not set");
if (MyProcNumber >= NumProcSignalSlots)
@@ -579,6 +623,18 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_UPDATE_XLOG_LOGICAL_INFO:
processed = ProcessBarrierUpdateXLogLogicalInfo();
break;
+ case PROCSIGNAL_BARRIER_SHBUF_SHRINK:
+ processed = ProcessBarrierShmemShrink();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM:
+ processed = ProcessBarrierShmemResizeMapAndMem();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_EXPAND:
+ processed = ProcessBarrierShmemExpand();
+ break;
+ case PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED:
+ processed = ProcessBarrierShmemResizeFailed();
+ break;
}
/*
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 5cb59d97871..aa6a44e9755 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -71,11 +71,19 @@
#include "funcapi.h"
#include "miscadmin.h"
#include "port/pg_numa.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
#include "storage/lwlock.h"
#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
#include "storage/shmem.h"
#include "storage/spin.h"
#include "utils/builtins.h"
+#include "utils/injection_point.h"
+#include "utils/wait_event.h"
/*
* This is the first data structure stored in the shared memory segment, at
@@ -498,8 +506,7 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, int segmen
{
/*
* Structure is in the shmem index so someone else has allocated it
- * already. The size better be the same as the size we are trying to
- * initialize to, or there is a name conflict (or worse).
+ * already. The size better be the same as the size we are trying to
*/
if (result->size != size)
{
@@ -509,6 +516,7 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, int segmen
" \"%s\": expected %zu, actual %zu",
name, size, result->size)));
}
+
structPtr = result->location;
}
else
@@ -543,6 +551,97 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, int segmen
return structPtr;
}
+/*
+ * ShmemResizeStructInSegment -- Resize the given structure in shared memory.
+ *
+ * This function resizes an existing shared memory structure while preserving
+ * the existing memory location.
+ *
+ * Returns: pointer to the existing structure location, if the resize is
+ * successful, otherwise NULL.
+ */
+void *
+ShmemResizeStructInSegment(const char *name, Size size, bool *foundPtr,
+ int segment_id)
+{
+ ShmemIndexEnt *result;
+ void *structPtr;
+ ShmemSegment *segment;
+ PGShmemHeader *shmhdr;
+ ShmemAllocatorData *ShmemAllocator;
+ Size allocated_size;
+ Size newFree;
+
+ Assert(segment_id >= 0 && segment_id < NUM_MEMORY_MAPPINGS);
+ Assert(segment_id != MAIN_SHMEM_SEGMENT); /* main segment structures not
+ * resizable */
+ Assert(ShmemIndex);
+ Assert(size > 0);
+ segment = &Segments[segment_id];
+ shmhdr = segment->ShmemSegHdr;
+ ShmemAllocator = segment->ShmemAllocator;
+ Assert(shmhdr != NULL);
+
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+ /* Look up the structure in the shmem index */
+ result = (ShmemIndexEnt *)
+ hash_search(ShmemIndex, name, HASH_FIND, foundPtr);
+
+ Assert(*foundPtr);
+ Assert(result);
+ Assert(result->segment_id == segment_id);
+
+ /* Save the existing structure pointer to be returned. */
+ structPtr = result->location;
+
+ /* Cachealign new size */
+ allocated_size = CACHELINEALIGN(size);
+
+ if (allocated_size == result->allocated_size)
+ {
+ result->size = size;
+ /* No need to resize if the existing allocated size is sufficient */
+ LWLockRelease(ShmemIndexLock);
+ return structPtr;
+ }
+
+ SpinLockAcquire(&ShmemAllocator->shmem_lock);
+
+ /*
+ * The resizable structures are placed in their own segment after the
+ * header and the spinlock. Hence the memory location where they end are
+ * same as the start of free memory in that segment.
+ */
+ Assert((char *) segment->ShmemBase + ShmemAllocator->free_offset == (char *) result->location + result->allocated_size);
+ newFree = ShmemAllocator->free_offset + (allocated_size - result->allocated_size);
+ if (newFree > shmhdr->totalsize)
+ {
+ structPtr = NULL;
+ }
+ else
+ {
+ ShmemAllocator->free_offset = newFree;
+ result->size = size;
+ result->allocated_size = allocated_size;
+ }
+
+ /*
+ * End of the structure should still be same as the start of free memory
+ * in the segment
+ */
+ Assert((char *) segment->ShmemBase + ShmemAllocator->free_offset == (char *) result->location + result->allocated_size);
+
+ SpinLockRelease(&ShmemAllocator->shmem_lock);
+ LWLockRelease(ShmemIndexLock);
+
+ /* note this assert is okay with structPtr == NULL */
+ Assert(structPtr == (void *) CACHELINEALIGN(structPtr));
+
+ return structPtr;
+}
+
+
+
/*
* Add two Size values, checking for overflow
*/
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 02e9aaa6bca..7334bfd15e7 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -63,6 +63,7 @@
#include "rewrite/rewriteHandler.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
#include "storage/pmsignal.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -4107,6 +4108,9 @@ PostgresSingleUserMain(int argc, char *argv[],
/* Initialize size of fast-path lock cache. */
InitializeFastPathLocks();
+ /* Initialize MaxNBuffers for buffer pool resizing. */
+ InitializeMaxNBuffers();
+
/*
* Give preloaded libraries a chance to request additional shared memory.
*/
@@ -4297,6 +4301,13 @@ PostgresMain(const char *dbname, const char *username)
*/
BeginReportingGUCOptions();
+ /*
+ * TODO: The new backend should fetch the shared buffers status. If the
+ * resizing is going on, it should bring itself upto speed with it. If
+ * not, simply fetch the latest pointers are sizes. Is this the right
+ * place to do that?
+ */
+
/*
* Also set up handler to log session end; we have to wait till now to be
* sure Log_disconnections has its final value.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 4aa864fe3c3..bfe5b9c409a 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -163,6 +163,7 @@ WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit."
WAL_RECEIVER_WAIT_START "Waiting for startup process to send initial data for streaming replication."
WAL_SUMMARY_READY "Waiting for a new WAL summary to be generated."
XACT_GROUP_UPDATE "Waiting for the group leader to update transaction status at transaction end."
+PM_BUFFER_RESIZE_WAIT "Waiting for the postmaster to complete shared buffer pool resize operations."
ABI_compatibility:
@@ -366,6 +367,7 @@ SerialControl "Waiting to read or update shared <filename>pg_serial</filename> s
AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
+ShmemResize "Waiting to resize shared memory."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index 36ad708b360..be9d3909c0f 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -139,7 +139,10 @@ int max_parallel_maintenance_workers = 2;
* MaxBackends is computed by PostmasterMain after modules have had a chance to
* register background workers.
*/
-int NBuffers = 16384;
+int NBuffers = 0;
+int NBuffersPending = 16384;
+bool finalMaxNBuffers = false;
+int MaxNBuffers = 0;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 3f401faf3de..3869ed5201c 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -595,6 +595,55 @@ InitializeFastPathLocks(void)
pg_nextpower2_32(FastPathLockGroupsPerBackend));
}
+/*
+ * Initialize MaxNBuffers variable with validation.
+ *
+ * This must be called after GUCs have been loaded but before shared memory size
+ * is determined.
+ *
+ * Since MaxNBuffers limits the size of the buffer pool, it must be at least as
+ * much as NBuffersPending. If MaxNBuffers is 0 (default), set it to
+ * NBuffersPending. Otherwise, validate that MaxNBuffers is not less than
+ * NBuffersPending.
+ */
+void
+InitializeMaxNBuffers(void)
+{
+ if (MaxNBuffers == 0) /* default/boot value */
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", NBuffersPending);
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+
+ /*
+ * We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
+ * However, if the DBA explicitly set max_shared_buffers = 0 in the
+ * config file, then PGC_S_DYNAMIC_DEFAULT will fail to override that
+ * and we must force the matter with PGC_S_OVERRIDE.
+ */
+ if (MaxNBuffers == 0) /* failed to apply it? */
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+ else
+ {
+ if (MaxNBuffers < NBuffersPending)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("max_shared_buffers (%d) cannot be less than current shared_buffers (%d)",
+ MaxNBuffers, NBuffersPending),
+ errhint("Increase max_shared_buffers or decrease shared_buffers.")));
+ }
+ }
+
+ Assert(MaxNBuffers > 0);
+ Assert(!finalMaxNBuffers);
+ finalMaxNBuffers = true;
+}
+
/*
* Early initialization of a backend (either standalone or under postmaster).
* This happens even before InitPostgres.
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index ae9d5f3fb70..6bc35a1132b 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -2599,7 +2599,7 @@ convert_to_base_unit(double value, const char *unit,
* the value without loss. For example, if the base unit is GUC_UNIT_KB, 1024
* is converted to 1 MB, but 1025 is represented as 1025 kB.
*/
-static void
+void
convert_int_from_base_unit(int64 base_value, int base_unit,
int64 *value, const char **unit)
{
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index c1f1603cd39..422cdb53336 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2030,6 +2030,15 @@
max => 'MAX_BACKENDS /* XXX? */',
},
+{ name => "max_shared_buffers", type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+ short_desc => 'Sets the upper limit for the shared_buffers value.',
+ flags => 'GUC_UNIT_BLOCKS',
+ variable => 'MaxNBuffers',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX / 2',
+},
+
{ name => 'max_slot_wal_keep_size', type => 'int', context => 'PGC_SIGHUP', group => 'REPLICATION_SENDING',
short_desc => 'Sets the maximum WAL size that can be reserved by replication slots.',
long_desc => 'Replication slots will be marked as failed, and segments released for deletion or recycling, if this much space is occupied by WAL on disk. -1 means no maximum.',
@@ -2598,13 +2607,15 @@
# We sometimes multiply the number of shared buffers by two without
# checking for overflow, so we mustn't allow more than INT_MAX / 2.
-{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
- variable => 'NBuffers',
+ variable => 'NBuffersPending',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
+ check_hook => 'check_shared_buffers',
+ show_hook => 'show_shared_buffers',
},
{ name => 'shared_memory_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 4b27f2a245e..3390b77921c 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12703,4 +12703,10 @@
proname => 'hashoid8extended', prorettype => 'int8',
proargtypes => 'oid8 int8', prosrc => 'hashoid8extended' },
+{ oid => '9999', descr => 'resize shared buffers according to the value of GUC `shared_buffers`',
+ proname => 'pg_resize_shared_buffers',
+ provolatile => 'v',
+ prorettype => 'bool',
+ proargtypes => '',
+ prosrc => 'pg_resize_shared_buffers'},
]
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index db559b39c4d..df6d3d8f4dd 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -173,7 +173,14 @@ extern PGDLLIMPORT bool ExitOnAnyError;
extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
+/*
+ * TODO: This is no more a GUC variable and does not track the size of the shared
+ * buffer pool; should be removed.
+ */
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersPending;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
@@ -502,6 +509,7 @@ extern PGDLLIMPORT ProcessingMode Mode;
extern void pg_split_opts(char **argv, int *argcp, const char *optstr);
extern void InitializeMaxBackends(void);
extern void InitializeFastPathLocks(void);
+extern void InitializeMaxNBuffers(void);
extern void InitPostgres(const char *in_dbname, Oid dboid,
const char *username, Oid useroid,
bits32 flags,
diff --git a/src/include/storage/buf.h b/src/include/storage/buf.h
index b21445522b1..a12d6b9082a 100644
--- a/src/include/storage/buf.h
+++ b/src/include/storage/buf.h
@@ -17,7 +17,7 @@
/*
* Buffer identifiers.
*
- * Zero is invalid, positive is the index of a shared buffer (1..NBuffers),
+ * Zero is invalid, positive is the index of a shared buffer (1..{size of shared buffer pool}),
* negative is the index of a local buffer (-1 .. -NLocBuffer).
*/
typedef int Buffer;
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index f4e9e703b8b..91e28ce19d2 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -137,8 +137,8 @@ StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2),
/*
* The maximum allowed value of usage_count represents a tradeoff between
- * accuracy and speed of the clock-sweep buffer management algorithm. A
- * large value (comparable to NBuffers) would approximate LRU semantics.
+ * accuracy and speed of the clock-sweep buffer management algorithm. A large
+ * value (comparable to the size of buffer pool) would approximate LRU semantics.
* But it can take as many as BM_MAX_USAGE_COUNT+1 complete cycles of the
* clock-sweep hand to find a free buffer, so in practice we don't want the
* value to be very large.
@@ -574,6 +574,7 @@ extern void StrategyNotifyBgWriter(int bgwprocno);
extern Size StrategyShmemSize(void);
extern void StrategyInitialize(bool init);
+extern void StrategyReset(int activeNBuffers);
/* buf_table.c */
extern Size BufTableShmemSize(int size);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 93348a34378..649a1a35105 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -21,6 +21,7 @@
#include "storage/bufpage.h"
#include "storage/pg_shmem.h"
#include "storage/relfilelocator.h"
+#include "utils/guc.h"
#include "utils/relcache.h"
#include "utils/snapmgr.h"
@@ -159,6 +160,7 @@ typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersPending;
/* in bufmgr.c */
extern PGDLLIMPORT bool zero_damaged_pages;
@@ -221,6 +223,11 @@ typedef enum BufferLockMode
BUFFER_LOCK_EXCLUSIVE,
} BufferLockMode;
+/*
+ * prototypes for functions in buf_init.c
+ */
+extern const char *show_shared_buffers(void);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
/*
* prototypes for functions in bufmgr.c
@@ -341,6 +348,7 @@ extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
extern bool BgBufferSync(WritebackContext *wb_context);
+extern void BgBufferSyncReset(int currentNBuffers, int targetNBuffers);
extern uint32 GetPinLimit(void);
extern uint32 GetLocalPinLimit(void);
@@ -365,10 +373,13 @@ extern void MarkDirtyRelUnpinnedBuffers(Relation rel,
extern void MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
int32 *buffers_already_dirty,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int targetNBuffers, int currentNBuffers);
/* in buf_init.c */
extern void BufferManagerShmemInit(void);
extern Size BufferManagerShmemSize(MemoryMappingSizes *mapping_sizes);
+extern void BufferManagerShmemResize(int currentNBuffers, int targetNBuffers);
+extern void BufferManagerShmemValidate(int targetNBuffers);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
@@ -417,7 +428,7 @@ extern void FreeAccessStrategy(BufferAccessStrategy strategy);
static inline bool
BufferIsValid(Buffer bufnum)
{
- Assert(bufnum <= NBuffers);
+ Assert(bufnum <= (Buffer) pg_atomic_read_u32(&ShmemCtrl->currentNBuffers));
Assert(bufnum >= -NLocBuffer);
return bufnum != InvalidBuffer;
@@ -471,4 +482,11 @@ BufferGetPage(Buffer buffer)
#endif /* FRONTEND */
+/* buf_resize.c */
+extern Datum pg_resize_shared_buffers(PG_FUNCTION_ARGS);
+extern bool ProcessBarrierShmemShrink(void);
+extern bool ProcessBarrierShmemResizeMapAndMem(void);
+extern bool ProcessBarrierShmemExpand(void);
+extern bool ProcessBarrierShmemResizeFailed(void);
+
#endif /* BUFMGR_H */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index f1d0802d048..23003412c9a 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -66,6 +66,7 @@ typedef void (*shmem_startup_hook_type) (void);
/* ipc.c */
extern PGDLLIMPORT bool proc_exit_inprogress;
extern PGDLLIMPORT bool shmem_exit_inprogress;
+extern PGDLLIMPORT volatile bool delay_shmem_resize;
pg_noreturn extern void proc_exit(int code);
extern void shmem_exit(int code);
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index e94ebce95b9..cb454f4c81d 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -87,6 +87,7 @@ PG_LWLOCK(52, SerialControl)
PG_LWLOCK(53, AioWorkerSubmissionQueue)
PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
+PG_LWLOCK(56, ShmemResize)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index eafab1dae52..aec8299d2b3 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -24,7 +24,12 @@
#ifndef PG_SHMEM_H
#define PG_SHMEM_H
+#include "port/atomics.h"
+#include "storage/barrier.h"
#include "storage/dsm_impl.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
+#include "utils/guc.h"
typedef struct PGShmemHeader /* standard header for all Postgres shmem */
@@ -98,11 +103,39 @@ typedef struct MemoryMappingSizes
extern PGDLLIMPORT PGUsedShmemInfo UsedShmemInfo[NUM_MEMORY_MAPPINGS];
+/*
+ * ShmemControl is shared between backends and helps to coordinate shared
+ * memory resize.
+ *
+ * TODO: I think we need a lock to protect this structure. If we do so, do we
+ * need to use atomic integers?
+ *
+ * TODO: Merge this structure into StrategyControl?
+ */
+typedef struct
+{
+ pg_atomic_flag resize_in_progress; /* true if resizing is in progress.
+ * false otherwise. */
+ pg_atomic_uint32 currentNBuffers; /* Original NBuffers value before
+ * resize started */
+ pg_atomic_uint32 targetNBuffers;
+ pid_t coordinator;
+} ShmemControl;
+
+extern PGDLLIMPORT ShmemControl *ShmemCtrl;
+
+/* The phases for shared memory resizing, used by for ProcSignal barrier. */
+#define SHMEM_RESIZE_REQUESTED 0
+#define SHMEM_RESIZE_START 1
+#define SHMEM_RESIZE_DONE 2
+
/* GUC variables */
extern PGDLLIMPORT int shared_memory_type;
extern PGDLLIMPORT int huge_pages;
extern PGDLLIMPORT int huge_page_size;
extern PGDLLIMPORT int huge_pages_status;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
/* Possible values for huge_pages and huge_pages_status */
typedef enum
@@ -134,13 +167,15 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping_sizes,
- PGShmemHeader **shim);
-extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
-extern void PGSharedMemoryDetach(void);
-extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
- int *memfd_flags);
-extern void PrepareHugePages(void);
+/*
+ * round off mapping size to a multiple of a typical page size.
+ */
+static inline void
+round_off_mapping_sizes(MemoryMappingSizes *mapping_sizes)
+{
+ mapping_sizes->shmem_req_size = add_size(mapping_sizes->shmem_req_size, 8192 - (mapping_sizes->shmem_req_size % 8192));
+ mapping_sizes->shmem_reserved = add_size(mapping_sizes->shmem_reserved, 8192 - (mapping_sizes->shmem_reserved % 8192));
+}
static inline const char *
MappingName(int segment_id)
@@ -156,5 +191,18 @@ MappingName(int segment_id)
}
}
+extern PGShmemHeader *PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *mapping_sizes,
+ PGShmemHeader **shim);
+extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
+extern void PGSharedMemoryDetach(void);
+extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
+ int *memfd_flags);
+extern bool PGSharedMemoryResize(int segment_id, MemoryMappingSizes *mapping_sizes);
+
+extern void PrepareHugePages(void);
+extern const char *show_shared_buffers(void);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
+extern void ShmemControlInit(void);
+
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index e52b8eb7697..9dc9819a72b 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -56,6 +56,15 @@ typedef enum
PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */
PROCSIGNAL_BARRIER_UPDATE_XLOG_LOGICAL_INFO, /* ask to update
* XLogLogicalInfo */
+ PROCSIGNAL_BARRIER_SHBUF_SHRINK, /* shrink buffer pool - restrict
+ * allocations to new size */
+ PROCSIGNAL_BARRIER_SHBUF_RESIZE_MAP_AND_MEM, /* remap shared memory
+ * segments and update
+ * structure pointers */
+ PROCSIGNAL_BARRIER_SHBUF_EXPAND, /* expand buffer pool - enable
+ * allocations in new range */
+ PROCSIGNAL_BARRIER_SHBUF_RESIZE_FAILED, /* signal backends that the shared
+ * buffer resizing failed. */
} ProcSignalBarrierType;
/*
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 5d58b3b39e6..39da4911bcf 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -39,11 +39,14 @@ extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
extern void *ShmemInitStructInSegment(const char *name, Size size,
bool *foundPtr, int segment_id);
+extern void *ShmemResizeStructInSegment(const char *name, Size size,
+ bool *foundPtr, int segment_id);
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
extern PGDLLIMPORT Size pg_get_shmem_pagesize(void);
+
/* ipci.c */
extern void RequestAddinShmemSpace(Size size);
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index bf39878c43e..72d8ba9e59e 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -459,6 +459,8 @@ extern config_handle *get_config_handle(const char *name);
extern void AlterSystemSetConfigFile(AlterSystemStmt *altersysstmt);
extern char *GetConfigOptionByName(const char *name, const char **varname,
bool missing_ok);
+extern void convert_int_from_base_unit(int64 base_value, int base_unit,
+ int64 *value, const char **unit);
extern void TransformGUCArray(ArrayType *array, List **names,
List **values);
diff --git a/src/test/Makefile b/src/test/Makefile
index 3eb0a06abb4..7a0d74086c1 100644
--- a/src/test/Makefile
+++ b/src/test/Makefile
@@ -20,7 +20,8 @@ SUBDIRS = \
postmaster \
recovery \
regress \
- subscription
+ subscription \
+ buffermgr
ifeq ($(with_icu),yes)
SUBDIRS += icu
diff --git a/src/test/README b/src/test/README
index afdc7676519..77f11607ff7 100644
--- a/src/test/README
+++ b/src/test/README
@@ -15,6 +15,9 @@ examples/
Demonstration programs for libpq that double as regression tests via
"make check"
+buffermgr/
+ Tests for resizing buffer pool without restarting the server
+
isolation/
Tests for concurrent behavior at the SQL level
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
new file mode 100644
index 00000000000..eb275027fa6
--- /dev/null
+++ b/src/test/buffermgr/Makefile
@@ -0,0 +1,30 @@
+#-------------------------------------------------------------------------
+#
+# Makefile for src/test/buffermgr
+#
+# Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+# Portions Copyright (c) 1994, Regents of the University of California
+#
+# src/test/buffermgr/Makefile
+#
+#-------------------------------------------------------------------------
+
+EXTRA_INSTALL = contrib/pg_buffercache
+
+REGRESS = buffer_resize
+
+# Custom configuration for buffer manager tests
+TEMP_CONFIG = $(srcdir)/buffermgr_test.conf
+
+subdir = src/test/buffermgr
+top_builddir = ../../..
+include $(top_builddir)/src/Makefile.global
+
+check:
+ $(prove_check)
+
+installcheck:
+ $(prove_installcheck)
+
+clean distclean:
+ rm -rf tmp_check
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
new file mode 100644
index 00000000000..c375ad80989
--- /dev/null
+++ b/src/test/buffermgr/README
@@ -0,0 +1,26 @@
+src/test/buffermgr/README
+
+Regression tests for buffer manager
+===================================
+
+This directory contains a test suite for resizing buffer manager without restarting the server.
+
+
+Running the tests
+=================
+
+NOTE: You must have given the --enable-tap-tests argument to configure.
+
+Run
+ make check
+or
+ make installcheck
+You can use "make installcheck" if you previously did "make install".
+In that case, the code in the installation tree is tested. With
+"make check", a temporary installation tree is built from the current
+sources and then tested.
+
+Either way, this test initializes, starts, and stops a test Postgres
+cluster.
+
+See src/test/perl/README for more info about running these tests.
diff --git a/src/test/buffermgr/buffermgr_test.conf b/src/test/buffermgr/buffermgr_test.conf
new file mode 100644
index 00000000000..a15f3e442a5
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.conf
@@ -0,0 +1,11 @@
+# Configuration for buffer manager regression tests
+
+# Even if max_shared_buffers is set multiple times only the last one is used to
+# as the limit on shared_buffers.
+max_shared_buffers = 128kB
+# Set initial shared_buffers as expected by test
+shared_buffers = 128MB
+# Set a larger value for max_shared_buffers to allow testing resize operations
+max_shared_buffers = 300MB
+# Turn huge pages off, since that affects the size of memory segments
+huge_pages = off
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
new file mode 100644
index 00000000000..47ffa69f3dd
--- /dev/null
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -0,0 +1,312 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+-- TODO: The test sets shared_buffers values in MBs. Instead it could use values
+-- in kBs so that the test runs on very small machines.
+-- TODO: test the actual memory allocated in the shared memory segments.
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+---------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | main | 2457600 | 2457600
+ Buffer IO Condition Variables | main | 614400 | 614400
+ Checkpoint BufferIds | main | 768000 | 768000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+---------+-----------+---------------
+ buffers | 134225920 | 314580992
+(1 row)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+---------+-----------+----------------
+ Buffer Blocks | buffers | 134221824 | 134221824
+ Buffer Descriptors | main | 2457600 | 2457600
+ Buffer IO Condition Variables | main | 614400 | 614400
+ Checkpoint BufferIds | main | 768000 | 768000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+---------+-----------+---------------
+ buffers | 134225920 | 314580992
+(1 row)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 128MB (pending: 64MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 64MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+---------+----------+----------------
+ Buffer Blocks | buffers | 67112960 | 67112960
+ Buffer Descriptors | main | 2457600 | 2457600
+ Buffer IO Condition Variables | main | 614400 | 614400
+ Checkpoint BufferIds | main | 768000 | 768000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+---------+----------+---------------
+ buffers | 67117056 | 314580992
+(1 row)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 8192
+(1 row)
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 64MB (pending: 256MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+---------+-----------+----------------
+ Buffer Blocks | buffers | 268439552 | 268439552
+ Buffer Descriptors | main | 2457600 | 2457600
+ Buffer IO Condition Variables | main | 614400 | 614400
+ Checkpoint BufferIds | main | 768000 | 768000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+---------+-----------+---------------
+ buffers | 268443648 | 314580992
+(1 row)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 32768
+(1 row)
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 256MB (pending: 100MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 100MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+---------+-----------+----------------
+ Buffer Blocks | buffers | 104861696 | 104861696
+ Buffer Descriptors | main | 2457600 | 2457600
+ Buffer IO Condition Variables | main | 614400 | 614400
+ Checkpoint BufferIds | main | 768000 | 768000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+---------+-----------+---------------
+ buffers | 104865792 | 314580992
+(1 row)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 12800
+(1 row)
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 100MB (pending: 128kB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | segment | size | allocated_size
+-------------------------------+---------+---------+----------------
+ Buffer Blocks | buffers | 135168 | 135168
+ Buffer Descriptors | main | 2457600 | 2457600
+ Buffer IO Condition Variables | main | 614400 | 614400
+ Checkpoint BufferIds | main | 768000 | 768000
+(4 rows)
+
+SELECT * FROM buffer_segments;
+ name | size | reserved_size
+---------+--------+---------------
+ buffers | 139264 | 314580992
+(1 row)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16
+(1 row)
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+ERROR: invalid value for parameter "shared_buffers": 51200
+DETAIL: "shared_buffers" must be less than "max_shared_buffers".
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
new file mode 100644
index 00000000000..c24bff721e6
--- /dev/null
+++ b/src/test/buffermgr/meson.build
@@ -0,0 +1,23 @@
+# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+
+tests += {
+ 'name': 'buffermgr',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'buffer_resize',
+ ],
+ 'regress_args': ['--temp-config', files('buffermgr_test.conf')],
+ },
+ 'tap': {
+ 'env': {
+ 'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
+ },
+ 'tests': [
+ 't/001_resize_buffer.pl',
+ 't/003_parallel_resize_buffer.pl',
+ 't/004_client_join_buffer_resize.pl',
+ ],
+ },
+}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
new file mode 100644
index 00000000000..fc27522f097
--- /dev/null
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -0,0 +1,97 @@
+-- Test buffer pool resizing and shared memory allocation tracking
+-- This test resizes the buffer pool multiple times and monitors
+-- shared memory allocations related to buffer management
+-- TODO: The test sets shared_buffers values in MBs. Instead it could use values
+-- in kBs so that the test runs on very small machines.
+
+-- TODO: test the actual memory allocated in the shared memory segments.
+
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, segment, size, allocated_size
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+
+-- Note: We exclude the 'main' segment even if it contains the shared buffer
+-- lookup table because it contains other shared structures whose total sizes
+-- may vary as the code changes.
+CREATE VIEW buffer_segments AS
+SELECT name, size, reserved_size
+FROM pg_shmem_segments
+WHERE name <> 'main'
+ORDER BY name;
+
+-- Enable pg_buffercache for buffer count verification
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT * FROM buffer_segments;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+SHOW max_shared_buffers;
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
new file mode 100644
index 00000000000..30e8bfea9cf
--- /dev/null
+++ b/src/test/buffermgr/t/001_resize_buffer.pl
@@ -0,0 +1,143 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Minimal test testing shared_buffer resizing under load
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Function to resize buffer pool and verify the change.
+sub apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ # Use the new pg_resize_shared_buffers() interface which handles everything synchronously
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ # If resize function fails, try a few times before giving up
+ my $max_retries = 5;
+ my $retry_delay = 1; # seconds
+ my $success = 0;
+ for my $attempt (1..$max_retries) {
+ my $result = $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()");
+ if ($result eq 't') {
+ $success = 1;
+ last;
+ }
+
+ # If not the last attempt, wait before retrying
+ if ($attempt < $max_retries) {
+ note "Resizing buffer pool to $new_size, attempt $attempt failed, retrying after $retry_delay seconds...";
+ sleep($retry_delay);
+ }
+ }
+
+ is($success, 1, 'resizing to ' . $new_size . ' succeeded after retries');
+ is($node->safe_psql('postgres', "SHOW shared_buffers"), $new_size,
+ 'SHOW after resizing to '. $new_size . ' succeeded');
+}
+
+# Initialize a cluster and start pgbench in the background for concurrent load.
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Permit resizing up to 1GB for this test and let the server start with 128MB.
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 160
+shared_buffers = 16
+log_statement = none
+});
+
+$node->start;
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+my $pgb_scale = 1;
+my $pgb_duration = 120;
+my $pgb_num_clients = 3;
+$node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$pgb_scale --quiet",
+ 0,
+ [qr{^$}],
+ [ # stderr patterns to verify initialization stages
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$pgb_scale)"
+);
+my ($pgbench_stdin, $pgbench_stdout, $pgbench_stderr) = ('', '', '');
+# Use --exit-on-abort so that the test stops on the first server crash or error,
+# thus making it easy to debug the failure. Use -C to increase the chances of a
+# new backend being created while resizing the buffer pool.
+my $pgbench_process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-T', $pgb_duration,
+ '-c', $pgb_num_clients,
+ '-C',
+ '--exit-on-abort',
+ 'postgres'
+ ],
+ '<' => \$pgbench_stdin,
+ '>' => \$pgbench_stdout,
+ '2>' => \$pgbench_stderr
+);
+
+ok($pgbench_process, "pgbench started successfully");
+
+# Allow pgbench to establish connections and start generating load.
+#
+# TODO: When creating new backends is known to work well with buffer pool
+# resizing, this wait should be removed.
+sleep(1);
+
+# Resize buffer pool to various sizes while pgbench is running in the
+# background. We use smaller sizes to induce frequent buffer eviction and
+# allocation. Also smaller buffer pool means frequent wraparound in background
+# writer, default buffer allocation strategy and checkpointer.
+#
+# TODO: These are pseudo-randomly picked sizes, but we can do better.
+my $tests_completed = 0;
+my @buffer_sizes = (32, 24, 29, 40, 29, 20, 16, 24);
+for my $target_size (@buffer_sizes)
+{
+ # Convert number of buffers to a string that will be reported by SHOW
+ # shared_buffers. This simple calculation works for sizes smaller than 128
+ # beyond which the unit changes to MB.
+ $target_size = $target_size * 8;
+ $target_size = $target_size . 'kB';
+
+ # Verify workload generator is still running
+ if (!$pgbench_process->pumpable) {
+ ok(0, "pgbench is still running");
+ last;
+ }
+
+ apply_and_verify_buffer_change($node, $target_size);
+ $tests_completed++;
+
+ # Wait for the resized buffer pool to stabilize. If the resized buffer pool
+ # is utilized fully, it might hit any wrongly initialized areas of shared
+ # memory.
+ sleep(2);
+}
+is($tests_completed, scalar(@buffer_sizes), "All buffer sizes were tested");
+
+# Make sure that pgbench can end normally.
+$pgbench_process->signal('TERM');
+IPC::Run::finish $pgbench_process;
+ok(grep { $pgbench_process->result == $_ } (0, 15), "pgbench exited gracefully");
+
+# Log any error output from pgbench for debugging
+diag("pgbench stderr:\n$pgbench_stderr");
+diag("pgbench stdout:\n$pgbench_stdout");
+
+# Ensure database is still functional after all the buffer changes
+$node->connect_ok("dbname=postgres",
+ "Database remains accessible after $tests_completed buffer resize operations");
+
+done_testing();
diff --git a/src/test/buffermgr/t/003_parallel_resize_buffer.pl b/src/test/buffermgr/t/003_parallel_resize_buffer.pl
new file mode 100644
index 00000000000..40d4bfde437
--- /dev/null
+++ b/src/test/buffermgr/t/003_parallel_resize_buffer.pl
@@ -0,0 +1,71 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test that only one pg_resize_shared_buffers() call succeeds when multiple
+# sessions attempt to resize buffers concurrently
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Initialize a cluster
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', 'shared_buffers = 128kB');
+$node->append_conf('postgresql.conf', 'max_shared_buffers = 256kB');
+$node->start;
+
+# Load injection points extension for test coordination
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Test 1: Two concurrent pg_resize_shared_buffers() calls
+# Set up injection point to pause the first resize call
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('pg-resize-shared-buffers-flag-set', 'wait')");
+
+# Change shared_buffers for the resize operation
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '144kB'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+# Start first resize session (will pause at injection point)
+my $session1 = $node->background_psql('postgres');
+$session1->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+);
+
+# Wait until session actually reaches the injection point
+$node->wait_for_event('client backend', 'pg-resize-shared-buffers-flag-set');
+
+# Start second resize session (should fail immediately since resize is in progress)
+my $result2 = $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()");
+
+# The second call should return false (already in progress)
+is($result2, 'f', 'Second concurrent resize call returns false');
+
+# Wake up the first session
+$node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('pg-resize-shared-buffers-flag-set')");
+
+# The pg_resize_shared_buffers() in session1 should now complete successfully
+# We can't easily capture the return value from query_until, but we can
+# verify the session completes without error and the resize actually happened
+$session1->quit;
+
+# Detach injection point
+$node->safe_psql('postgres',
+ "SELECT injection_points_detach('pg-resize-shared-buffers-flag-set')");
+
+done_testing();
diff --git a/src/test/buffermgr/t/004_client_join_buffer_resize.pl b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
new file mode 100644
index 00000000000..072eee535f6
--- /dev/null
+++ b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
@@ -0,0 +1,243 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test shared_buffer resizing coordination with client connections joining using injection points
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+use Time::HiRes qw(sleep);
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Function to calculate the size of test table required to fill up maximum
+# buffer pool when populating it.
+sub calculate_test_sizes
+{
+ my ($node, $block_size) = @_;
+
+ # Get the maximum buffer pool size from configuration
+ my $max_shared_buffers = $node->safe_psql('postgres', "SHOW max_shared_buffers");
+ my ($max_val, $max_unit) = ($max_shared_buffers =~ /(\d+)(\w+)/);
+ my $max_size_bytes;
+ if (lc($max_unit) eq 'kb') {
+ $max_size_bytes = $max_val * 1024;
+ } elsif (lc($max_unit) eq 'mb') {
+ $max_size_bytes = $max_val * 1024 * 1024;
+ } elsif (lc($max_unit) eq 'gb') {
+ $max_size_bytes = $max_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $max_size_bytes = $max_val * 1024;
+ }
+
+ # Fill more pages than minimally required to increase the chances of pages
+ # from the test table filling the buffer cache.
+ $max_size_bytes = $max_size_bytes;
+ my $pages_needed = int($max_size_bytes / $block_size) + 10; # Add some extra to ensure buffers are filled
+ my $rows_to_insert = $pages_needed * 100; # Assuming roughly 100 rows per page for our table structure
+ return ($max_size_bytes, $pages_needed, $rows_to_insert);
+}
+
+# Function to calculate expected buffer count from size string
+sub calculate_buffer_count
+{
+ my ($size_string, $block_size) = @_;
+ # Parse size and convert to bytes
+ my ($size_val, $unit) = ($size_string =~ /(\d+)(\w+)/);
+ my $size_bytes;
+ if (lc($unit) eq 'kb') {
+ $size_bytes = $size_val * 1024;
+ } elsif (lc($unit) eq 'mb') {
+ $size_bytes = $size_val * 1024 * 1024;
+ } elsif (lc($unit) eq 'gb') {
+ $size_bytes = $size_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $size_bytes = $size_val * 1024;
+ }
+ return int($size_bytes / $block_size);
+}
+
+# Initialize cluster with very small buffer sizes for testing
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Configure for buffer resizing with very small buffer pool sizes for faster tests.
+# TODO: for some reason parallel workers try to load default number of shared_buffers which doesn't work with lower max_shared_buffers. We need to fix that - somewhere it's picking default value of shared buffers. For now disable parallelism
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 512kB
+shared_buffers = 320kB
+max_parallel_workers_per_gather = 0
+});
+$node->start;
+
+# Enable injection points
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Get the block size (this is fixed for the binary)
+my $block_size = $node->safe_psql('postgres', "SHOW block_size");
+
+# Try to create pg_buffercache extension for buffer analysis
+eval {
+ $node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+};
+if ($@) {
+ $node->stop;
+ plan skip_all => 'pg_buffercache extension not available - cannot verify buffer usage';
+}
+
+# Create a small test table, and fetch its properties for later reference if required.
+$node->safe_psql('postgres', qq{
+ CREATE TABLE client_test (c1 int, data char(50));
+});
+my $table_oid = $node->safe_psql('postgres', "SELECT oid FROM pg_class WHERE relname = 'client_test'");
+my $table_relfilenode = $node->safe_psql('postgres', "SELECT relfilenode FROM pg_class WHERE relname = 'client_test'");
+note("Test table client_test: OID = $table_oid, relfilenode = $table_relfilenode");
+my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $block_size);
+
+# Create dedicated sessions for injection point handling and test queries,
+# so that we don't create new backends for test operations after starting
+# resize operation. Only one backend, which tests new backend synchronization
+# with resizing operation, should start after resizing has commenced.
+my $injection_session = $node->background_psql('postgres');
+my $query_session = $node->background_psql('postgres');
+my $resize_session = $node->background_psql('postgres');
+
+# Function to run a single injection point test
+sub run_injection_point_test
+{
+ my ($test_name, $injection_point, $target_size, $operation_type) = @_;
+
+ # Silence the logging of the statements we run to avoid
+ # unnecessarily bloating the test logs. This runs before the
+ # upgrade we're testing, so the details should not be very
+ # interesting for debugging. But if needed, you can make it more
+ # verbose by setting this.
+ my $verbose = 0;
+
+ note("Test with $test_name ($operation_type)");
+
+ # Calculate test parameters before starting resize
+ my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $target_size, $block_size);
+
+ # Update buffer pool size and wait for it to reflect pending state
+ $resize_session->query_safe("ALTER SYSTEM SET shared_buffers = '$target_size'", verbose => $verbose);
+ $resize_session->query_safe("SELECT pg_reload_conf()", verbose => $verbose);
+ my $pending_size_str = "pending: $target_size";
+ $resize_session->poll_query_until("SELECT substring(current_setting('shared_buffers'), '$pending_size_str')", $pending_size_str, verbose => $verbose);
+
+ # Set up injection point in injection session
+ $injection_session->query_safe("SELECT injection_points_attach('$injection_point', 'wait')", verbose => $verbose);
+
+ # Trigger resize
+ $resize_session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+ );
+
+ # Wait until resize actually reaches the injection point using the query session
+ $query_session->wait_for_event('client backend', $injection_point, verbose => $verbose);
+
+ # Start a client while resize is paused
+ my $client = $node->background_psql('postgres');
+ note("Background client backend PID: " . $client->query_safe("SELECT pg_backend_pid()", verbose => $verbose));
+
+ # Wake up the injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_wakeup('$injection_point')", verbose => $verbose);
+
+ # Test buffer functionality immediately after waking up injection point
+ # Insert data to test buffer pool functionality during/after resize
+ $client->query_safe("INSERT INTO client_test SELECT i, 'test_data_' || i FROM generate_series(1, $rows_to_insert) i", verbose => $verbose);
+ # Verify the data was inserted correctly and can be read back
+ is($client->query_safe("SELECT COUNT(*) FROM client_test", verbose => $verbose), $rows_to_insert, "inserted $rows_to_insert during $test_name ($operation_type) successful");
+
+ # Verify table size is reasonable (should be substantial for testing)
+ ok($query_session->query_safe("SELECT pg_total_relation_size('client_test')", verbose => $verbose) >= $max_size_bytes,"table size is large enough to overflow buffer pool in test $test_name ($operation_type)");
+
+ # Wait for the resize operation to complete. There is no direct way to do so
+ # in background_psql. Hence fire a psql command and wait for it to finish
+ $resize_session->query(q(\echo 'done'), verbose => $verbose);
+
+ # Detach injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_detach('$injection_point')", verbose => $verbose);
+
+ # Verify resize completed successfully
+ is($query_session->query_safe("SELECT current_setting('shared_buffers')", verbose => $verbose), $target_size,
+ "resize completed successfully to $target_size");
+
+ # Check buffer pool size using pg_buffercache after resize completion
+ is($query_session->query_safe("SELECT COUNT(*) FROM pg_buffercache", verbose => $verbose), calculate_buffer_count($target_size, $block_size), "all buffers in the buffer pool used in $test_name ($operation_type)");
+
+ # Wait for client to complete
+ ok($client->quit, "client succeeded during $test_name ($operation_type)");
+
+ # Clean up for next test
+ $query_session->query_safe("DELETE FROM client_test", verbose => $verbose);
+}
+
+# Test injection points during buffer resize with client connections
+my @common_injection_tests = (
+ {
+ name => 'flag setting phase',
+ injection_point => 'pg-resize-shared-buffers-flag-set',
+ },
+ {
+ name => 'memory remap phase',
+ injection_point => 'pgrsb-after-shmem-resize',
+ },
+ {
+ name => 'resize map barrier complete',
+ injection_point => 'pgrsb-resize-barrier-sent',
+ },
+);
+
+# Test common injection points for both shrinking and expanding
+foreach my $test (@common_injection_tests)
+{
+ # Test shrinking scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '272kB', 'shrinking');
+
+ # Test expanding scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '400kB', 'expanding');
+}
+
+my @shrink_only_tests = (
+ {
+ name => 'shrink barrier complete',
+ injection_point => 'pgrsb-shrink-barrier-sent',
+ size => '200kB',
+ }
+);
+foreach my $test (@shrink_only_tests)
+{
+ run_injection_point_test($test->{name}, $test->{injection_point}, $test->{size}, 'shrinking only');
+}
+
+my @expand_only_tests = (
+ {
+ name => 'expand barrier complete',
+ injection_point => 'pgrsb-expand-barrier-sent',
+ size => '416kB',
+ }
+);
+
+foreach my $test (@expand_only_tests)
+{
+ run_injection_point_test($test->{name}, $test->{injection_point}, $test->{size}, 'expanding only');
+}
+
+$injection_session->quit;
+$query_session->quit;
+$resize_session->quit;
+
+done_testing();
diff --git a/src/test/meson.build b/src/test/meson.build
index cd45cbf57fb..e9550933063 100644
--- a/src/test/meson.build
+++ b/src/test/meson.build
@@ -4,6 +4,7 @@ subdir('regress')
subdir('isolation')
subdir('authentication')
+subdir('buffermgr')
subdir('postmaster')
subdir('recovery')
subdir('subscription')
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index c6ff2dbde4c..079abde29d9 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -61,6 +61,7 @@ use Config;
use IPC::Run;
use PostgreSQL::Test::Utils qw(pump_until);
use Test::More;
+use Time::HiRes qw(usleep);
=pod
@@ -374,4 +375,79 @@ sub set_query_timer_restart
return $self->{query_timer_restart};
}
+=pod
+
+=item $session->poll_query_until($query [, $expected ])
+
+Run B<$query> repeatedly in this background session, until it returns the
+B<$expected> result ('t', or SQL boolean true, by default).
+Continues polling if the query returns an error result.
+Times out after a reasonable number of attempts.
+Returns 1 if successful, 0 if timed out.
+
+=cut
+
+sub poll_query_until
+{
+ my ($self, $query, $expected, %params) = @_;
+
+ $expected = 't' unless defined($expected); # default value
+
+ my $max_attempts = 10 * $PostgreSQL::Test::Utils::timeout_default;
+ my $attempts = 0;
+ my ($stdout, $stderr_flag);
+
+ while ($attempts < $max_attempts)
+ {
+ ($stdout, $stderr_flag) = $self->query($query, %params);
+
+ chomp($stdout);
+
+ # If query succeeded and returned expected result
+ if (!$stderr_flag && $stdout eq $expected)
+ {
+ return 1;
+ }
+
+ # Wait 0.1 second before retrying.
+ usleep(100_000);
+
+ $attempts++;
+ }
+
+ # Give up. Print the output from the last attempt, hopefully that's useful
+ # for debugging.
+ my $stderr_output = $stderr_flag ? $self->{stderr} : '';
+ diag qq(poll_query_until timed out executing this query:
+$query
+expecting this output:
+$expected
+last actual query output:
+$stdout
+with stderr:
+$stderr_output);
+ return 0;
+}
+
+=item $session->wait_for_event(backend_type, wait_event_name)
+
+Poll pg_stat_activity until backend_type reaches wait_event_name using this
+background session.
+
+=cut
+
+sub wait_for_event
+{
+ my ($self, $backend_type, $wait_event_name, %params) = @_;
+
+ $self->poll_query_until(qq[
+ SELECT count(*) > 0 FROM pg_stat_activity
+ WHERE backend_type = '$backend_type' AND wait_event = '$wait_event_name'
+ ], undef, %params)
+ or die
+ qq(timed out when waiting for $backend_type to reach wait event '$wait_event_name');
+
+ return;
+}
+
1;
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index e4dbe1b787b..2a698b1f4e6 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2810,6 +2810,7 @@ ShippableCacheEntry
ShmemAllocatorData
ShippableCacheKey
ShmemIndexEnt
+ShmemControl
ShmemSegment
ShutdownForeignScan_function
ShutdownInformation
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-09 22:14 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-10 15:23 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 20:42 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
1 sibling, 2 replies; 167+ messages in thread
From: Heikki Linnakangas @ 2026-02-09 22:14 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On 09/02/2026 17:15, Ashutosh Bapat wrote:
> On Sun, Feb 8, 2026 at 5:14 AM Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>> Could you write a standalone test module in src/test/modules to
>> demonstrate how to use the resizable shmem segments, please? That'd
>> allow focusing on that interface without worrying all the other
>> complexities of shared buffers.
>
> From your writeup it seems like you are leaning towards creating the
> shared memory segments on-demand rather than having predefined
> segments as done in the patch.
Right. Same as with normal shared memory structs: there are
ShmemInitStruct calls scattered throughout the codebase, there's no need
to predefine them in a single header file or anything like that.
(Note: "on-demand" doesn't mean "on the fly", i.e. each shared memory
area still needs to be registered at postmaster startup.)
> I think your proposal is interesting
> and might create a possibility for extensions to be able to create
> resizable shared memory structures.
Ah yes, good point. This facility should certainly be open for
extensions too. I didn't even realize that the predefined segments
scheme would not allow that.
> Since the shared memory segments are predefined, a test module, as you
> suggest, does not have a segment that it can use. That led me to think
> that you are imagining some kind of on-demand shared memory segments
> that extensions or test modules can use. I think that will be a useful
> feature by itself. Just as an example, imagine an extension to provide
> shared plan cache which resizes the plan cache as needed. However, we
> need to make sure that these on-demand shared memory segments work
> well with CalculateShmemSize(), PGSharedMemoryDetach(),
> AnonymousShmemDetach() etc.
+1
> 2. There is no predefined limit on the number of segments. When
> CalculateShmemSize() is called, BufferManagerShmemSize() outputs the
> shared memory required in the main segment and total of shared memory
> required in the other segments. Similarly shmem_request_hook of each
> external module outputs the shared memory required in the main segment
> and the total of shared memory required in the other segments. The
> main segment is created in CreateSharedMemoryAndSemaphores() directly
> using the requested size. The other output size is merely used for
> reporting purposes in InitializeShmemGUCs(). When allocating shared
> data structures, the main shared memory data structures are allocated
> using ShmemInitStruct(), whereas the data structures in other memory
> segments are allocated using ShmemInitStructExt() which takes
> immediate allocation size and size of address space to be reserved as
> arguments. It does not require segment_id though. For every new data
> structure that ShmemInitStructExt() encounters, it a. creates a new
> shared memory segment using PGSharedMemoryCreate(), b. allocates that
> structure with the initial size. This function also needs to create
> the PGShmemInfo and AnonShmemData entries corresponding to new
> segments and make them available to PGSharedMemoryDetach(),
> AnonymousShmemDetach() etc. For that we create a shared hash table in
> the main shared memory segment (just like ShmemIndex) where we store
> the metadata against each segment name (by coining it from the name of
> the structure). We expect ShmemInitStructExt() to allocate structures
> and segments only in the Postmaster and only at the beginning. When a
> backend starts, it pulls the segment metadata from the shared hash
> table which is ultimately used by PGSharedMemoryDetach(),
> AnonymousShmemDetach(). At run time any backend which has access to
> the shared memory should be able to call ShmemResizeStruct() given a
> resizable structure and new size. Some higher level synchronization is
> needed to ensure that the same structure is not resized simultaneously
> by two backends.
>
> The first approach is simple but has limited use given the fixed
> number of segments. Second is more flexible but that's some work. I am
> not sure whether it's worth doing all that if there are hardly any
> extensions which could use resizable shared data structures.
It doesn't seem *that* much more work.
Putting this patch aside for a moment, I don't much like our current
interface for defining shared memory structs anyway. The
[SubSystem]ShmemSize() functions feel too detached from the
ShmemInitStruct() calls. For example, it's an easy mistake to return a
slightly different size in the FoobarShmemSize() call than what you use
in the ShmemInitStruct() call. And then there's the fact that the
initialization functions run both at postmaster startup but also at
backend startup in EXEC_BACKEND mode. There were reasons for that when
EXEC_BACKEND was introduced, but it's always felt awkward to me.
I feel that it'd be good to have a single definition of each shmem
struct, and derive all the other things from there.
Attached is a proof-of-concept of what I have in mind. Don't look too
closely at how it's implemented, it's very hacky and EXEC_BACKEND mode
is slightly broken, for example. The point is to demonstrate what the
callers would look like. I converted only a few subsystems to use the
new API, the rest still use ShmemInitStruct() and ShmemInitHash().
With this, initialization of a subsystem that defines a shared memory
area looks like this:
--------------
/* This struct lives in shared memory */
typedef struct
{
int field;
} FoobarSharedCtlData;
static void FoobarShmemInit(void *arg);
/* Descriptor for the shared memory area */
ShmemStructDesc FoobarShmemDesc = {
.name = "Foobar subsystem",
.size = sizeof(FoobarSharedCtlData),
.init_fn = FoobarShmemInit,
};
/* Pointer to the shared memory struct */
#define FoobarCtl ((FoobarSharedCtlData *) FoobarShmemDesc.ptr)
/*
* Register the shared memory struct. This is called once at
* postmaster startup, before the shared memory segment is allocated,
* and in EXEC_BACKEND mode also early at backend startup.
*
* For core subsystems, there's a list of all these functions in core
* in ipci.c, similar to all the *ShmemSize() and *ShmemInit() functions
* today. In an extension, this would be done in _PG_init() or in
* the shmem_request_hook, replacing the RequestAddinShmemSpace calls
* we have today.
*/
void
FoobarShmemRegister(void)
{
ShmemRegisterStruct(&FoobarShmemDesc);
}
/*
* This callback is called once at postmaster startup, to initialize
* the shared memory struct. FoobarShmemDesc.ptr has already been
* set when this is called.
*/
static void
FoobarShmemInit(void *arg)
{
memset(FoobarCtl, 0, sizeof(FoobarSharedCtlData));
FoobarCtl->field = 123;
}
--------------
The ShmemStructDesc provides room for extending the facility in the
future. For example, you could specify alignment there, or an additional
"attach" callback when you need to do more per-backend initialization in
EXEC_BACKEND mode. And with the resizeable shared memory, a max size.
Thoughts?
- Heikki
Attachments:
[text/x-patch] 0001-wip-Introduce-a-new-way-of-registering-shared-memory.patch (53.7K, ../../91265854-b3ba-45c6-aa44-7e8dcdd51470@iki.fi/2-0001-wip-Introduce-a-new-way-of-registering-shared-memory.patch)
download | inline diff:
From c3c90517778b47156787f05a529096416dd67727 Mon Sep 17 00:00:00 2001
From: Heikki Linnakangas <heikki.linnakangas@iki.fi>
Date: Mon, 9 Feb 2026 22:28:23 +0200
Subject: [PATCH 1/1] wip: Introduce a new way of registering shared memory
structs
---
.../pg_stat_statements/pg_stat_statements.c | 112 ++++-----
src/backend/access/transam/varsup.c | 32 +--
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/postmaster/launch_backend.c | 11 +-
src/backend/postmaster/postmaster.c | 2 +
src/backend/storage/ipc/dsm.c | 46 ++--
src/backend/storage/ipc/dsm_registry.c | 34 ++-
src/backend/storage/ipc/ipci.c | 51 ++--
src/backend/storage/ipc/pmsignal.c | 53 ++--
src/backend/storage/ipc/procarray.c | 127 +++++-----
src/backend/storage/ipc/procsignal.c | 63 +++--
src/backend/storage/ipc/shmem.c | 230 +++++++++++++++++-
src/backend/storage/ipc/sinvaladt.c | 39 +--
src/backend/storage/lmgr/proc.c | 156 ++++++------
src/backend/tcop/postgres.c | 2 +
src/include/access/transam.h | 12 +-
src/include/storage/dsm_registry.h | 3 +-
src/include/storage/ipc.h | 1 +
src/include/storage/pmsignal.h | 3 +-
src/include/storage/proc.h | 5 +-
src/include/storage/procarray.h | 3 +-
src/include/storage/procsignal.h | 3 +-
src/include/storage/shmem.h | 57 +++++
src/include/storage/sinvaladt.h | 3 +-
24 files changed, 662 insertions(+), 388 deletions(-)
diff --git a/contrib/pg_stat_statements/pg_stat_statements.c b/contrib/pg_stat_statements/pg_stat_statements.c
index 4a427533bd8..71debc8b47f 100644
--- a/contrib/pg_stat_statements/pg_stat_statements.c
+++ b/contrib/pg_stat_statements/pg_stat_statements.c
@@ -258,6 +258,25 @@ typedef struct pgssSharedState
pgssGlobalStats stats; /* global statistics for pgss */
} pgssSharedState;
+static void pgss_shmem_init(void *arg);
+
+static ShmemStructDesc pgssSharedStateShmemDesc = {
+ .name = "pg_stat_statements",
+ .size = sizeof(pgssSharedState),
+ .init_fn = pgss_shmem_init,
+};
+
+static ShmemHashDesc pgssSharedHashDesc = {
+ .name = "pg_stat_statements hash",
+ .init_size = 0, /* set from 'pgss_max' */
+ .max_size = 0, /* set from 'pgss_max' */
+};
+
+/* Links to shared memory state */
+#define pgss ((pgssSharedState *) pgssSharedStateShmemDesc.ptr)
+#define pgss_hash (pgssSharedHashDesc.ptr)
+
+
/*---- Local variables ----*/
/* Current nesting depth of planner/ExecutorRun/ProcessUtility calls */
@@ -274,10 +293,6 @@ static ExecutorFinish_hook_type prev_ExecutorFinish = NULL;
static ExecutorEnd_hook_type prev_ExecutorEnd = NULL;
static ProcessUtility_hook_type prev_ProcessUtility = NULL;
-/* Links to shared memory state */
-static pgssSharedState *pgss = NULL;
-static HTAB *pgss_hash = NULL;
-
/*---- GUC variables ----*/
typedef enum
@@ -365,7 +380,6 @@ static void pgss_store(const char *query, int64 queryId,
static void pg_stat_statements_internal(FunctionCallInfo fcinfo,
pgssVersion api_version,
bool showtext);
-static Size pgss_memsize(void);
static pgssEntry *entry_alloc(pgssHashKey *key, Size query_offset, int query_len,
int encoding, bool sticky);
static void entry_dealloc(void);
@@ -500,11 +514,39 @@ _PG_init(void)
static void
pgss_shmem_request(void)
{
+ HASHCTL info;
+
if (prev_shmem_request_hook)
prev_shmem_request_hook();
- RequestAddinShmemSpace(pgss_memsize());
RequestNamedLWLockTranche("pg_stat_statements", 1);
+
+ /*
+ * Register our shared memory state, including hash table
+ */
+ ShmemRegisterStruct(&pgssSharedStateShmemDesc);
+
+ info.keysize = sizeof(pgssHashKey);
+ info.entrysize = sizeof(pgssEntry);
+ pgssSharedHashDesc.init_size = pgss_max;
+ pgssSharedHashDesc.max_size = pgss_max;
+ ShmemRegisterHash(&pgssSharedHashDesc,
+ &info,
+ HASH_ELEM | HASH_BLOBS);
+}
+
+static void
+pgss_shmem_init(void *arg)
+{
+ pgss->lock = &(GetNamedLWLockTranche("pg_stat_statements"))->lock;
+ pgss->cur_median_usage = ASSUMED_MEDIAN_INIT;
+ pgss->mean_query_len = ASSUMED_LENGTH_INIT;
+ SpinLockInit(&pgss->mutex);
+ pgss->extent = 0;
+ pgss->n_writers = 0;
+ pgss->gc_count = 0;
+ pgss->stats.dealloc = 0;
+ pgss->stats.stats_reset = GetCurrentTimestamp();
}
/*
@@ -516,8 +558,6 @@ pgss_shmem_request(void)
static void
pgss_shmem_startup(void)
{
- bool found;
- HASHCTL info;
FILE *file = NULL;
FILE *qfile = NULL;
uint32 header;
@@ -530,42 +570,6 @@ pgss_shmem_startup(void)
if (prev_shmem_startup_hook)
prev_shmem_startup_hook();
- /* reset in case this is a restart within the postmaster */
- pgss = NULL;
- pgss_hash = NULL;
-
- /*
- * Create or attach to the shared memory state, including hash table
- */
- LWLockAcquire(AddinShmemInitLock, LW_EXCLUSIVE);
-
- pgss = ShmemInitStruct("pg_stat_statements",
- sizeof(pgssSharedState),
- &found);
-
- if (!found)
- {
- /* First time through ... */
- pgss->lock = &(GetNamedLWLockTranche("pg_stat_statements"))->lock;
- pgss->cur_median_usage = ASSUMED_MEDIAN_INIT;
- pgss->mean_query_len = ASSUMED_LENGTH_INIT;
- SpinLockInit(&pgss->mutex);
- pgss->extent = 0;
- pgss->n_writers = 0;
- pgss->gc_count = 0;
- pgss->stats.dealloc = 0;
- pgss->stats.stats_reset = GetCurrentTimestamp();
- }
-
- info.keysize = sizeof(pgssHashKey);
- info.entrysize = sizeof(pgssEntry);
- pgss_hash = ShmemInitHash("pg_stat_statements hash",
- pgss_max, pgss_max,
- &info,
- HASH_ELEM | HASH_BLOBS);
-
- LWLockRelease(AddinShmemInitLock);
-
/*
* If we're in the postmaster (or a standalone backend...), set up a shmem
* exit hook to dump the statistics to disk.
@@ -573,12 +577,6 @@ pgss_shmem_startup(void)
if (!IsUnderPostmaster)
on_shmem_exit(pgss_shmem_shutdown, (Datum) 0);
- /*
- * Done if some other process already completed our initialization.
- */
- if (found)
- return;
-
/*
* Note: we don't bother with locks here, because there should be no other
* processes running when this code is reached.
@@ -2082,20 +2080,6 @@ pg_stat_statements_info(PG_FUNCTION_ARGS)
PG_RETURN_DATUM(HeapTupleGetDatum(heap_form_tuple(tupdesc, values, nulls)));
}
-/*
- * Estimate shared memory space needed.
- */
-static Size
-pgss_memsize(void)
-{
- Size size;
-
- size = MAXALIGN(sizeof(pgssSharedState));
- size = add_size(size, hash_estimate_size(pgss_max, sizeof(pgssEntry)));
-
- return size;
-}
-
/*
* Allocate a new hashtable entry.
* caller must hold an exclusive lock on pgss->lock
diff --git a/src/backend/access/transam/varsup.c b/src/backend/access/transam/varsup.c
index 3e95d4cfd16..11ad90e7372 100644
--- a/src/backend/access/transam/varsup.c
+++ b/src/backend/access/transam/varsup.c
@@ -30,35 +30,27 @@
/* Number of OIDs to prefetch (preallocate) per XLOG write */
#define VAR_OID_PREFETCH 8192
-/* pointer to variables struct in shared memory */
-TransamVariablesData *TransamVariables = NULL;
+static void VarsupShmemInit(void *arg);
+ShmemStructDesc TransamVariablesShmemDesc = {
+ .name = "TransamVariables",
+ .size = sizeof(TransamVariablesData),
+ .init_fn = VarsupShmemInit,
+};
/*
* Initialization of shared memory for TransamVariables.
*/
-Size
-VarsupShmemSize(void)
+void
+VarsupShmemRegister(void)
{
- return sizeof(TransamVariablesData);
+ ShmemRegisterStruct(&TransamVariablesShmemDesc);
}
-void
-VarsupShmemInit(void)
+static void
+VarsupShmemInit(void *arg)
{
- bool found;
-
- /* Initialize our shared state struct */
- TransamVariables = ShmemInitStruct("TransamVariables",
- sizeof(TransamVariablesData),
- &found);
- if (!IsUnderPostmaster)
- {
- Assert(!found);
- memset(TransamVariables, 0, sizeof(TransamVariablesData));
- }
- else
- Assert(found);
+ memset(TransamVariables, 0, sizeof(TransamVariablesData));
}
/*
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index 7d32cd0e159..0ded7018e86 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -337,6 +337,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ RegisterShmemStructs();
+
CreateSharedMemoryAndSemaphores();
/*
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index 05b1feef3cf..5d397e3c619 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -49,6 +49,7 @@
#include "replication/walreceiver.h"
#include "storage/dsm.h"
#include "storage/io_worker.h"
+#include "storage/ipc.h"
#include "storage/pg_shmem.h"
#include "tcop/backend_startup.h"
#include "utils/memutils.h"
@@ -104,12 +105,10 @@ typedef struct
char **LWLockTrancheNames;
int *LWLockCounter;
LWLockPadded *MainLWLockArray;
- slock_t *ProcStructLock;
PROC_HDR *ProcGlobal;
PGPROC *AuxiliaryProcs;
PGPROC *PreparedXactProcs;
volatile PMSignalData *PMSignalState;
- ProcSignalHeader *ProcSignal;
pid_t PostmasterPid;
TimestampTz PgStartTime;
TimestampTz PgReloadTime;
@@ -678,8 +677,12 @@ SubPostmasterMain(int argc, char *argv[])
/* Restore basic shared memory pointers */
if (UsedShmemSegAddr != NULL)
+ {
InitShmemAllocator(UsedShmemSegAddr);
+ RegisterShmemStructs();
+ }
+
/*
* Run the appropriate Main function
*/
@@ -735,12 +738,10 @@ save_backend_variables(BackendParameters *param,
param->LWLockTrancheNames = LWLockTrancheNames;
param->LWLockCounter = LWLockCounter;
param->MainLWLockArray = MainLWLockArray;
- param->ProcStructLock = ProcStructLock;
param->ProcGlobal = ProcGlobal;
param->AuxiliaryProcs = AuxiliaryProcs;
param->PreparedXactProcs = PreparedXactProcs;
param->PMSignalState = PMSignalState;
- param->ProcSignal = ProcSignal;
param->PostmasterPid = PostmasterPid;
param->PgStartTime = PgStartTime;
@@ -995,12 +996,10 @@ restore_backend_variables(BackendParameters *param)
LWLockTrancheNames = param->LWLockTrancheNames;
LWLockCounter = param->LWLockCounter;
MainLWLockArray = param->MainLWLockArray;
- ProcStructLock = param->ProcStructLock;
ProcGlobal = param->ProcGlobal;
AuxiliaryProcs = param->AuxiliaryProcs;
PreparedXactProcs = param->PreparedXactProcs;
PMSignalState = param->PMSignalState;
- ProcSignal = param->ProcSignal;
PostmasterPid = param->PostmasterPid;
PgStartTime = param->PgStartTime;
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index d6133bfebc6..f6d3369f917 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -968,6 +968,8 @@ PostmasterMain(int argc, char *argv[])
* shared memory, determine the value of any runtime-computed GUCs that
* depend on the amount of shared memory required.
*/
+ RegisterShmemStructs();
+
InitializeShmemGUCs();
/*
diff --git a/src/backend/storage/ipc/dsm.c b/src/backend/storage/ipc/dsm.c
index 6a5b16392f7..55f46c7687e 100644
--- a/src/backend/storage/ipc/dsm.c
+++ b/src/backend/storage/ipc/dsm.c
@@ -108,7 +108,15 @@ static inline bool is_main_region_dsm_handle(dsm_handle handle);
static bool dsm_init_done = false;
/* Preallocated DSM space in the main shared memory region. */
-static void *dsm_main_space_begin = NULL;
+static void dsm_main_space_init(void *);
+
+static ShmemStructDesc dsm_main_space_shmem_desc = {
+ .name = "Preallocated DSM",
+ .size = 0, /* dynamic */
+ .init_fn = dsm_main_space_init,
+};
+
+#define dsm_main_space_begin (dsm_main_space_shmem_desc.ptr)
/*
* List of dynamic shared memory segments used by this backend.
@@ -479,27 +487,29 @@ void
dsm_shmem_init(void)
{
size_t size = dsm_estimate_size();
- bool found;
if (size == 0)
return;
- dsm_main_space_begin = ShmemInitStruct("Preallocated DSM", size, &found);
- if (!found)
- {
- FreePageManager *fpm = (FreePageManager *) dsm_main_space_begin;
- size_t first_page = 0;
- size_t pages;
-
- /* Reserve space for the FreePageManager. */
- while (first_page * FPM_PAGE_SIZE < sizeof(FreePageManager))
- ++first_page;
-
- /* Initialize it and give it all the rest of the space. */
- FreePageManagerInitialize(fpm, dsm_main_space_begin);
- pages = (size / FPM_PAGE_SIZE) - first_page;
- FreePageManagerPut(fpm, first_page, pages);
- }
+ ShmemRegisterStruct(&dsm_main_space_shmem_desc);
+}
+
+static void
+dsm_main_space_init(void *arg)
+{
+ size_t size = dsm_main_space_shmem_desc.size;
+ FreePageManager *fpm = (FreePageManager *) dsm_main_space_begin;
+ size_t first_page = 0;
+ size_t pages;
+
+ /* Reserve space for the FreePageManager. */
+ while (first_page * FPM_PAGE_SIZE < sizeof(FreePageManager))
+ ++first_page;
+
+ /* Initialize it and give it all the rest of the space. */
+ FreePageManagerInitialize(fpm, dsm_main_space_begin);
+ pages = (size / FPM_PAGE_SIZE) - first_page;
+ FreePageManagerPut(fpm, first_page, pages);
}
/*
diff --git a/src/backend/storage/ipc/dsm_registry.c b/src/backend/storage/ipc/dsm_registry.c
index 068c1577b12..882af83b7b2 100644
--- a/src/backend/storage/ipc/dsm_registry.c
+++ b/src/backend/storage/ipc/dsm_registry.c
@@ -54,7 +54,15 @@ typedef struct DSMRegistryCtxStruct
dshash_table_handle dshh;
} DSMRegistryCtxStruct;
-static DSMRegistryCtxStruct *DSMRegistryCtx;
+static void DSMRegistryCtxShmemInit(void *arg);
+
+static ShmemStructDesc DSMRegistryCtxShmemDesc = {
+ .name = "DSM Registry Data",
+ .size = sizeof(DSMRegistryCtxStruct),
+ .init_fn = DSMRegistryCtxShmemInit,
+};
+
+#define DSMRegistryCtx ((DSMRegistryCtxStruct *) DSMRegistryCtxShmemDesc.ptr)
typedef struct NamedDSMState
{
@@ -113,27 +121,17 @@ static const dshash_parameters dsh_params = {
static dsa_area *dsm_registry_dsa;
static dshash_table *dsm_registry_table;
-Size
-DSMRegistryShmemSize(void)
+void
+DSMRegistryShmemRegister(void)
{
- return MAXALIGN(sizeof(DSMRegistryCtxStruct));
+ ShmemRegisterStruct(&DSMRegistryCtxShmemDesc);
}
-void
-DSMRegistryShmemInit(void)
+static void
+DSMRegistryCtxShmemInit(void *)
{
- bool found;
-
- DSMRegistryCtx = (DSMRegistryCtxStruct *)
- ShmemInitStruct("DSM Registry Data",
- DSMRegistryShmemSize(),
- &found);
-
- if (!found)
- {
- DSMRegistryCtx->dsah = DSA_HANDLE_INVALID;
- DSMRegistryCtx->dshh = DSHASH_HANDLE_INVALID;
- }
+ DSMRegistryCtx->dsah = DSA_HANDLE_INVALID;
+ DSMRegistryCtx->dshh = DSHASH_HANDLE_INVALID;
}
/*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 1f7e933d500..952988645d0 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -101,13 +101,14 @@ CalculateShmemSize(void)
size = add_size(size, hash_estimate_size(SHMEM_INDEX_SIZE,
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
- size = add_size(size, DSMRegistryShmemSize());
+
+ size = add_size(size, ShmemRegisteredSize());
+
+ /* legacy subsystmes */
size = add_size(size, BufferManagerShmemSize());
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
- size = add_size(size, ProcGlobalShmemSize());
size = add_size(size, XLogPrefetchShmemSize());
- size = add_size(size, VarsupShmemSize());
size = add_size(size, XLOGShmemSize());
size = add_size(size, XLogRecoveryShmemSize());
size = add_size(size, CLOGShmemSize());
@@ -117,11 +118,7 @@ CalculateShmemSize(void)
size = add_size(size, BackgroundWorkerShmemSize());
size = add_size(size, MultiXactShmemSize());
size = add_size(size, LWLockShmemSize());
- size = add_size(size, ProcArrayShmemSize());
size = add_size(size, BackendStatusShmemSize());
- size = add_size(size, SharedInvalShmemSize());
- size = add_size(size, PMSignalShmemSize());
- size = add_size(size, ProcSignalShmemSize());
size = add_size(size, CheckpointerShmemSize());
size = add_size(size, AutoVacuumShmemSize());
size = add_size(size, ReplicationSlotsShmemSize());
@@ -217,6 +214,10 @@ CreateSharedMemoryAndSemaphores(void)
*/
InitShmemAllocator(seghdr);
+ /* Reserve space for semaphores. */
+ if (!IsUnderPostmaster)
+ PGReserveSemaphores(ProcGlobalSemas());
+
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -230,6 +231,19 @@ CreateSharedMemoryAndSemaphores(void)
shmem_startup_hook();
}
+void
+RegisterShmemStructs(void)
+{
+ DSMRegistryShmemRegister();
+
+ ProcGlobalShmemRegister();
+ VarsupShmemRegister();
+ ProcArrayShmemRegister();
+ SharedInvalShmemRegister();
+ PMSignalShmemRegister();
+ ProcSignalShmemRegister();
+}
+
/*
* Initialize various subsystems, setting up their data structures in
* shared memory.
@@ -259,14 +273,23 @@ CreateOrAttachShmemStructs(void)
*/
InitShmemIndex();
+#ifdef EXEC_BACKEND
+ if (IsUnderPostmaster)
+ ShmemAttachRegistered();
+ else
+#endif
+ {
+ ShmemInitRegistered();
+ }
+
dsm_shmem_init();
- DSMRegistryShmemInit();
+ //DSMRegistryShmemInit();
/*
* Set up xlog, clog, and buffers
*/
- VarsupShmemInit();
XLOGShmemInit();
+
XLogPrefetchShmemInit();
XLogRecoveryShmemInit();
CLOGShmemInit();
@@ -288,23 +311,13 @@ CreateOrAttachShmemStructs(void)
/*
* Set up process table
*/
- if (!IsUnderPostmaster)
- InitProcGlobal();
- ProcArrayShmemInit();
BackendStatusShmemInit();
TwoPhaseShmemInit();
BackgroundWorkerShmemInit();
- /*
- * Set up shared-inval messaging
- */
- SharedInvalShmemInit();
-
/*
* Set up interprocess signaling mechanisms
*/
- PMSignalShmemInit();
- ProcSignalShmemInit();
CheckpointerShmemInit();
AutoVacuumShmemInit();
ReplicationSlotsShmemInit();
diff --git a/src/backend/storage/ipc/pmsignal.c b/src/backend/storage/ipc/pmsignal.c
index 4618820b337..23752500d16 100644
--- a/src/backend/storage/ipc/pmsignal.c
+++ b/src/backend/storage/ipc/pmsignal.c
@@ -80,9 +80,24 @@ struct PMSignalData
sig_atomic_t PMChildFlags[FLEXIBLE_ARRAY_MEMBER];
};
-/* PMSignalState pointer is valid in both postmaster and child processes */
+static void PMSignalShmemInit(void *);
+
+static ShmemStructDesc PMSignalShmemDesc = {
+ .name = "PMSignalState",
+ .size = 0, /* dynamic */
+ .init_fn = PMSignalShmemInit,
+};
+
+/*
+ * PMSignalState pointer is valid in both postmaster and child processes
+ *
+ * This is a stand-alone variable rather than just a #define over
+ * PMSignalShmemDesc.ptr because it is needed early at backend startup and
+ * passed as a backend parameter in EXEC_BACKEND mode
+ */
NON_EXEC_STATIC volatile PMSignalData *PMSignalState = NULL;
+
/*
* Local copy of PMSignalState->num_child_flags, only valid in the
* postmaster. Postmaster keeps a local copy so that it doesn't need to
@@ -123,39 +138,28 @@ postmaster_death_handler(SIGNAL_ARGS)
static void MarkPostmasterChildInactive(int code, Datum arg);
/*
- * PMSignalShmemSize
- * Compute space needed for pmsignal.c's shared memory
+ * PMSignalShmemRegister - Register our shared memory
*/
-Size
-PMSignalShmemSize(void)
+void
+PMSignalShmemRegister(void)
{
Size size;
size = offsetof(PMSignalData, PMChildFlags);
size = add_size(size, mul_size(MaxLivePostmasterChildren(),
sizeof(sig_atomic_t)));
-
- return size;
+ PMSignalShmemDesc.size = size;
+ ShmemRegisterStruct(&PMSignalShmemDesc);
}
-/*
- * PMSignalShmemInit - initialize during shared-memory creation
- */
-void
-PMSignalShmemInit(void)
+static void
+PMSignalShmemInit(void *arg)
{
- bool found;
-
- PMSignalState = (PMSignalData *)
- ShmemInitStruct("PMSignalState", PMSignalShmemSize(), &found);
-
- if (!found)
- {
- /* initialize all flags to zeroes */
- MemSet(unvolatize(PMSignalData *, PMSignalState), 0, PMSignalShmemSize());
- num_child_flags = MaxLivePostmasterChildren();
- PMSignalState->num_child_flags = num_child_flags;
- }
+ /* initialize all flags to zeroes */
+ PMSignalState = PMSignalShmemDesc.ptr;
+ MemSet(unvolatize(PMSignalData *, PMSignalState), 0, PMSignalShmemDesc.size);
+ num_child_flags = MaxLivePostmasterChildren();
+ PMSignalState->num_child_flags = num_child_flags;
}
/*
@@ -291,6 +295,7 @@ RegisterPostmasterChildActive(void)
{
int slot = MyPMChildSlot;
+ Assert(PMSignalState);
Assert(slot > 0 && slot <= PMSignalState->num_child_flags);
slot--;
Assert(PMSignalState->PMChildFlags[slot] == PM_CHILD_ASSIGNED);
diff --git a/src/backend/storage/ipc/procarray.c b/src/backend/storage/ipc/procarray.c
index 301f54fb5a8..08c63bcb2a7 100644
--- a/src/backend/storage/ipc/procarray.c
+++ b/src/backend/storage/ipc/procarray.c
@@ -101,6 +101,18 @@ typedef struct ProcArrayStruct
int pgprocnos[FLEXIBLE_ARRAY_MEMBER];
} ProcArrayStruct;
+static void ProcArrayShmemInit(void *arg);
+static void ProcArrayShmemAttach(void *arg);
+
+static ShmemStructDesc ProcArrayShmemDesc = {
+ .name = "Proc Array",
+ .size = 0, /* dynamic */
+ .init_fn = ProcArrayShmemInit,
+ .attach_fn = ProcArrayShmemAttach,
+};
+
+#define procArray ((ProcArrayStruct *) ProcArrayShmemDesc.ptr)
+
/*
* State for the GlobalVisTest* family of functions. Those functions can
* e.g. be used to decide if a deleted row can be removed without violating
@@ -267,9 +279,6 @@ typedef enum KAXCompressReason
KAX_STARTUP_PROCESS_IDLE, /* startup process is about to sleep */
} KAXCompressReason;
-
-static ProcArrayStruct *procArray;
-
static PGPROC *allProcs;
/*
@@ -280,8 +289,23 @@ static TransactionId cachedXidIsNotInProgress = InvalidTransactionId;
/*
* Bookkeeping for tracking emulated transactions in recovery
*/
-static TransactionId *KnownAssignedXids;
-static bool *KnownAssignedXidsValid;
+
+static ShmemStructDesc KnownAssignedXidsShmemDesc = {
+ .name = "KnownAssignedXids",
+ .size = 0, /* dynamic */
+ .init_fn = NULL,
+};
+
+#define KnownAssignedXids ((TransactionId *) KnownAssignedXidsShmemDesc.ptr)
+
+static ShmemStructDesc KnownAssignedXidsValidShmemDesc = {
+ .name = "KnownAssignedXidsValid",
+ .size = 0, /* dynamic */
+ .init_fn = NULL,
+};
+
+#define KnownAssignedXidsValid ((bool *) KnownAssignedXidsValidShmemDesc.ptr)
+
static TransactionId latestObservedXid = InvalidTransactionId;
/*
@@ -372,18 +396,19 @@ static inline FullTransactionId FullXidRelativeTo(FullTransactionId rel,
static void GlobalVisUpdateApply(ComputeXidHorizonsResult *horizons);
/*
- * Report shared-memory space needed by ProcArrayShmemInit
+ * Register the shared PGPROC array during postmaster startup.
*/
-Size
-ProcArrayShmemSize(void)
+void
+ProcArrayShmemRegister(void)
{
- Size size;
-
- /* Size of the ProcArray structure itself */
#define PROCARRAY_MAXPROCS (MaxBackends + max_prepared_xacts)
- size = offsetof(ProcArrayStruct, pgprocnos);
- size = add_size(size, mul_size(sizeof(int), PROCARRAY_MAXPROCS));
+ /* Create or attach to the ProcArray shared structure */
+ ProcArrayShmemDesc.size =
+ add_size(offsetof(ProcArrayStruct, pgprocnos),
+ mul_size(sizeof(int),
+ PROCARRAY_MAXPROCS));
+ ShmemRegisterStruct(&ProcArrayShmemDesc);
/*
* During Hot Standby processing we have a data structure called
@@ -403,64 +428,38 @@ ProcArrayShmemSize(void)
if (EnableHotStandby)
{
- size = add_size(size,
- mul_size(sizeof(TransactionId),
- TOTAL_MAX_CACHED_SUBXIDS));
- size = add_size(size,
- mul_size(sizeof(bool), TOTAL_MAX_CACHED_SUBXIDS));
+ KnownAssignedXidsShmemDesc.size =
+ mul_size(sizeof(TransactionId),
+ TOTAL_MAX_CACHED_SUBXIDS);
+ ShmemRegisterStruct(&KnownAssignedXidsShmemDesc);
+
+ KnownAssignedXidsValidShmemDesc.size =
+ mul_size(sizeof(bool), TOTAL_MAX_CACHED_SUBXIDS);
+ ShmemRegisterStruct(&KnownAssignedXidsValidShmemDesc);
}
-
- return size;
}
-/*
- * Initialize the shared PGPROC array during postmaster startup.
- */
-void
-ProcArrayShmemInit(void)
+static void
+ProcArrayShmemInit(void *arg)
{
- bool found;
-
- /* Create or attach to the ProcArray shared structure */
- procArray = (ProcArrayStruct *)
- ShmemInitStruct("Proc Array",
- add_size(offsetof(ProcArrayStruct, pgprocnos),
- mul_size(sizeof(int),
- PROCARRAY_MAXPROCS)),
- &found);
-
- if (!found)
- {
- /*
- * We're the first - initialize.
- */
- procArray->numProcs = 0;
- procArray->maxProcs = PROCARRAY_MAXPROCS;
- procArray->maxKnownAssignedXids = TOTAL_MAX_CACHED_SUBXIDS;
- procArray->numKnownAssignedXids = 0;
- procArray->tailKnownAssignedXids = 0;
- procArray->headKnownAssignedXids = 0;
- procArray->lastOverflowedXid = InvalidTransactionId;
- procArray->replication_slot_xmin = InvalidTransactionId;
- procArray->replication_slot_catalog_xmin = InvalidTransactionId;
- TransamVariables->xactCompletionCount = 1;
- }
+ procArray->numProcs = 0;
+ procArray->maxProcs = PROCARRAY_MAXPROCS;
+ procArray->maxKnownAssignedXids = TOTAL_MAX_CACHED_SUBXIDS;
+ procArray->numKnownAssignedXids = 0;
+ procArray->tailKnownAssignedXids = 0;
+ procArray->headKnownAssignedXids = 0;
+ procArray->lastOverflowedXid = InvalidTransactionId;
+ procArray->replication_slot_xmin = InvalidTransactionId;
+ procArray->replication_slot_catalog_xmin = InvalidTransactionId;
+ TransamVariables->xactCompletionCount = 1;
allProcs = ProcGlobal->allProcs;
+}
- /* Create or attach to the KnownAssignedXids arrays too, if needed */
- if (EnableHotStandby)
- {
- KnownAssignedXids = (TransactionId *)
- ShmemInitStruct("KnownAssignedXids",
- mul_size(sizeof(TransactionId),
- TOTAL_MAX_CACHED_SUBXIDS),
- &found);
- KnownAssignedXidsValid = (bool *)
- ShmemInitStruct("KnownAssignedXidsValid",
- mul_size(sizeof(bool), TOTAL_MAX_CACHED_SUBXIDS),
- &found);
- }
+static void
+ProcArrayShmemAttach(void *arg)
+{
+ allProcs = ProcGlobal->allProcs;
}
/*
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 8e56922dcea..5743f088324 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -102,7 +102,16 @@ struct ProcSignalHeader
#define BARRIER_CLEAR_BIT(flags, type) \
((flags) &= ~(((uint32) 1) << (uint32) (type)))
-NON_EXEC_STATIC ProcSignalHeader *ProcSignal = NULL;
+static void ProcSignalShmemInit(void *arg);
+
+static ShmemStructDesc ProcSignalShmemDesc = {
+ .name = "ProcSignal",
+ .size = 0, /* dynamic */
+ .init_fn = ProcSignalShmemInit,
+};
+
+#define ProcSignal ((ProcSignalHeader *) ProcSignalShmemDesc.ptr)
+
static ProcSignalSlot *MyProcSignalSlot = NULL;
static bool CheckProcSignal(ProcSignalReason reason);
@@ -110,51 +119,37 @@ static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
/*
- * ProcSignalShmemSize
- * Compute space needed for ProcSignal's shared memory
+ * ProcSignalShmemRegister
+ * Register ProcSignal's shared memory needs at postmaster startup
*/
-Size
-ProcSignalShmemSize(void)
+void
+ProcSignalShmemRegister(void)
{
Size size;
size = mul_size(NumProcSignalSlots, sizeof(ProcSignalSlot));
size = add_size(size, offsetof(ProcSignalHeader, psh_slot));
- return size;
+
+ ProcSignalShmemDesc.size = size;
+ ShmemRegisterStruct(&ProcSignalShmemDesc);
}
-/*
- * ProcSignalShmemInit
- * Allocate and initialize ProcSignal's shared memory
- */
-void
-ProcSignalShmemInit(void)
+static void
+ProcSignalShmemInit(void *arg)
{
- Size size = ProcSignalShmemSize();
- bool found;
+ pg_atomic_init_u64(&ProcSignal->psh_barrierGeneration, 0);
- ProcSignal = (ProcSignalHeader *)
- ShmemInitStruct("ProcSignal", size, &found);
-
- /* If we're first, initialize. */
- if (!found)
+ for (int i = 0; i < NumProcSignalSlots; ++i)
{
- int i;
-
- pg_atomic_init_u64(&ProcSignal->psh_barrierGeneration, 0);
+ ProcSignalSlot *slot = &ProcSignal->psh_slot[i];
- for (i = 0; i < NumProcSignalSlots; ++i)
- {
- ProcSignalSlot *slot = &ProcSignal->psh_slot[i];
-
- SpinLockInit(&slot->pss_mutex);
- pg_atomic_init_u32(&slot->pss_pid, 0);
- slot->pss_cancel_key_len = 0;
- MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
- pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
- pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
- ConditionVariableInit(&slot->pss_barrierCV);
- }
+ SpinLockInit(&slot->pss_mutex);
+ pg_atomic_init_u32(&slot->pss_pid, 0);
+ slot->pss_cancel_key_len = 0;
+ MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
+ pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
+ ConditionVariableInit(&slot->pss_barrierCV);
}
}
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9f362ce8641..2ba6385ffc6 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -19,6 +19,8 @@
* methods). The routines in this file are used for allocating and
* binding to shared memory data structures.
*
+ * FIXME: NOTES below are outdated
+ *
* NOTES:
* (a) There are three kinds of shared memory data structures
* available to POSTGRES: fixed-size structures, queues and hash
@@ -76,6 +78,16 @@
#include "storage/spin.h"
#include "utils/builtins.h"
+/* size constants for the shmem index table */
+ /* max size of data structure string name */
+#define SHMEM_INDEX_KEYSIZE (48)
+ /* estimated size of the shmem index table (not a hard limit) */
+#define SHMEM_INDEX_SIZE (64)
+
+/* these are in postmaster private memory */
+static ShmemStructDesc *registry[SHMEM_INDEX_SIZE];
+static int num_registrations = 0;
+
/*
* This is the first data structure stored in the shared memory segment, at
* the offset that PGShmemHeader->content_offset points to. Allocations by
@@ -95,6 +107,9 @@ typedef struct ShmemAllocatorData
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void shmem_hash_init(void *arg);
+static void shmem_hash_attach(void *arg);
+
/* shared memory global variables */
static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
@@ -103,13 +118,134 @@ static void *ShmemEnd; /* end+1 address of shared memory */
static ShmemAllocatorData *ShmemAllocator;
slock_t *ShmemLock; /* points to ShmemAllocator->shmem_lock */
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+
+
+static ShmemHashDesc ShmemIndexHashDesc = {
+ .name = "ShmemIndex",
+ .init_size = SHMEM_INDEX_SIZE,
+ .max_size = SHMEM_INDEX_SIZE,
+};
+
+ /* primary index hashtable for shmem */
+#define ShmemIndex (ShmemIndexHashDesc.ptr)
+
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
Datum pg_numa_available(PG_FUNCTION_ARGS);
+
+void
+ShmemRegisterStruct(ShmemStructDesc *desc)
+{
+ elog(DEBUG2, "REGISTER: %s with size %zd", desc->name, desc->size);
+
+ registry[num_registrations++] = desc;
+}
+
+size_t
+ShmemRegisteredSize(void)
+{
+ size_t size;
+
+ size = 0;
+ for (int i = 0; i < num_registrations; i++)
+ {
+ size = add_size(size, registry[i]->size);
+ size = add_size(size, registry[i]->extra_size);
+ }
+
+ elog(DEBUG2, "SIZE: total %zd", size);
+
+ return size;
+}
+
+void
+ShmemInitRegistered(void)
+{
+ for (int i = 0; i < num_registrations; i++)
+ {
+ size_t allocated_size;
+ void *structPtr;
+ bool found;
+ ShmemIndexEnt *result;
+
+ elog(DEBUG2, "INIT [%d/%d]: %s", i, num_registrations, registry[i]->name);
+
+ /* look it up in the shmem index */
+ result = (ShmemIndexEnt *)
+ hash_search(ShmemIndex, registry[i]->name, HASH_ENTER_NULL, &found);
+ if (!result)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("could not create ShmemIndex entry for data structure \"%s\"",
+ registry[i]->name)));
+ }
+ if (found)
+ elog(ERROR, "shmem struct \"%s\" is already initialized", registry[i]->name);
+
+ /* allocate and initialize it */
+ structPtr = ShmemAllocRaw(registry[i]->size, &allocated_size);
+ if (structPtr == NULL)
+ {
+ /* out of memory; remove the failed ShmemIndex entry */
+ hash_search(ShmemIndex, registry[i]->name, HASH_REMOVE, NULL);
+ ereport(ERROR,
+ (errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("not enough shared memory for data structure"
+ " \"%s\" (%zu bytes requested)",
+ registry[i]->name, registry[i]->size)));
+ }
+ result->size = registry[i]->size;
+ result->allocated_size = allocated_size;
+ result->location = structPtr;
+
+ registry[i]->ptr = structPtr;
+ if (registry[i]->init_fn)
+ registry[i]->init_fn(registry[i]->init_fn_arg);
+ }
+}
+
+#ifdef EXEC_BACKEND
+void
+ShmemAttachRegistered(void)
+{
+ /* Must be initializing a (non-standalone) backend */
+ Assert(IsUnderPostmaster);
+ Assert(ShmemAllocator->index != NULL);
+
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+
+ for (int i = 0; i < num_registrations; i++)
+ {
+ bool found;
+ ShmemIndexEnt *result;
+
+ elog(LOG, "ATTACH [%d/%d]: %s", i, num_registrations, registry[i]->name);
+
+ /* look it up in the shmem index */
+ result = (ShmemIndexEnt *)
+ hash_search(ShmemIndex, registry[i]->name, HASH_FIND, &found);
+ if (!found)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("could not find ShmemIndex entry for data structure \"%s\"",
+ registry[i]->name)));
+ }
+
+ registry[i]->ptr = result->location;
+
+ if (registry[i]->attach_fn)
+ registry[i]->attach_fn(registry[i]->attach_fn_arg);
+ }
+
+ LWLockRelease(ShmemIndexLock);
+}
+#endif
+
/*
* InitShmemAllocator() --- set up basic pointers to shared memory.
*
@@ -292,6 +428,98 @@ InitShmemIndex(void)
HASH_ELEM | HASH_STRINGS);
}
+/*
+ * ShmemInitHash -- Create and initialize, or attach to, a
+ * shared memory hash table.
+ *
+ * We assume caller is doing some kind of synchronization
+ * so that two processes don't try to create/initialize the same
+ * table at once. (In practice, all creations are done in the postmaster
+ * process; child processes should always be attaching to existing tables.)
+ *
+ * max_size is the estimated maximum number of hashtable entries. This is
+ * not a hard limit, but the access efficiency will degrade if it is
+ * exceeded substantially (since it's used to compute directory size and
+ * the hash table buckets will get overfull).
+ *
+ * init_size is the number of hashtable entries to preallocate. For a table
+ * whose maximum size is certain, this should be equal to max_size; that
+ * ensures that no run-time out-of-shared-memory failures can occur.
+ *
+ * *infoP and hash_flags must specify at least the entry sizes and key
+ * comparison semantics (see hash_create()). Flag bits and values specific
+ * to shared-memory hash tables are added here, except that callers may
+ * choose to specify HASH_PARTITION and/or HASH_FIXED_SIZE.
+ *
+ * Note: before Postgres 9.0, this function returned NULL for some failure
+ * cases. Now, it always throws error instead, so callers need not check
+ * for NULL.
+ */
+void
+ShmemRegisterHash(ShmemHashDesc *desc, /* configuration */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags) /* info about infoP */
+{
+ /*
+ * Hash tables allocated in shared memory have a fixed directory; it can't
+ * grow or other backends wouldn't be able to find it. So, make sure we
+ * make it big enough to start with.
+ *
+ * The shared memory allocator must be specified too.
+ */
+ infoP->dsize = infoP->max_dsize = hash_select_dirsize(desc->max_size);
+ infoP->alloc = ShmemAllocNoError;
+ hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
+
+ /* look it up in the shmem index */
+ memset(&desc->base_desc, 0, sizeof(desc->base_desc));
+ desc->base_desc.name = desc->name;
+ desc->base_desc.size = hash_get_shared_size(infoP, hash_flags);
+ desc->base_desc.init_fn = shmem_hash_init;
+ desc->base_desc.init_fn_arg = desc;
+ desc->base_desc.attach_fn = shmem_hash_attach;
+ desc->base_desc.attach_fn_arg = desc;
+
+ desc->base_desc.extra_size = hash_estimate_size(desc->max_size, infoP->entrysize) - desc->base_desc.size;
+
+ desc->hash_flags = hash_flags;
+ desc->infoP = MemoryContextAlloc(TopMemoryContext, sizeof(HASHCTL));
+ memcpy(desc->infoP, infoP, sizeof(HASHCTL));
+
+ ShmemRegisterStruct(&desc->base_desc);
+}
+
+static void
+shmem_hash_init(void *arg)
+{
+ ShmemHashDesc *desc = (ShmemHashDesc *) arg;
+ int hash_flags = desc->hash_flags;
+
+ /* Pass location of hashtable header to hash_create */
+ desc->ptr = desc->base_desc.ptr;
+ desc->infoP->hctl = (HASHHDR *) desc->ptr;
+
+ desc->ptr = hash_create(desc->name, desc->init_size, desc->infoP, hash_flags);
+}
+
+static void
+shmem_hash_attach(void *arg)
+{
+ ShmemHashDesc *desc = (ShmemHashDesc *) arg;
+ int hash_flags = desc->hash_flags;
+
+ /*
+ * if it already exists, attach to it rather than allocate and initialize
+ * new space
+ */
+ hash_flags |= HASH_ATTACH;
+
+ /* Pass location of hashtable header to hash_create */
+ desc->infoP->hctl = (HASHHDR *) desc->ptr;
+
+ desc->ptr = hash_create(desc->name, desc->init_size, desc->infoP, hash_flags);
+}
+
/*
* ShmemInitHash -- Create and initialize, or attach to, a
* shared memory hash table.
diff --git a/src/backend/storage/ipc/sinvaladt.c b/src/backend/storage/ipc/sinvaladt.c
index a7a7cc4f0a9..0fe0f256971 100644
--- a/src/backend/storage/ipc/sinvaladt.c
+++ b/src/backend/storage/ipc/sinvaladt.c
@@ -203,7 +203,16 @@ typedef struct SISeg
*/
#define NumProcStateSlots (MaxBackends + NUM_AUXILIARY_PROCS)
-static SISeg *shmInvalBuffer; /* pointer to the shared inval buffer */
+static void SharedInvalShmemInit(void *arg);
+
+static ShmemStructDesc SharedInvalShmemDesc = {
+ .name = "shmInvalBuffer",
+ .size = 0, /* dynamic */
+ .init_fn = SharedInvalShmemInit,
+};
+
+/* pointer to the shared inval buffer */
+#define shmInvalBuffer ((SISeg *) SharedInvalShmemDesc.ptr)
static LocalTransactionId nextLocalTransactionId;
@@ -212,10 +221,11 @@ static void CleanupInvalidationState(int status, Datum arg);
/*
- * SharedInvalShmemSize --- return shared-memory space needed
+ * SharedInvalShmemRegister
+ * Register shared memory needs for the SI message buffer
*/
-Size
-SharedInvalShmemSize(void)
+void
+SharedInvalShmemRegister(void)
{
Size size;
@@ -223,26 +233,17 @@ SharedInvalShmemSize(void)
size = add_size(size, mul_size(sizeof(ProcState), NumProcStateSlots)); /* procState */
size = add_size(size, mul_size(sizeof(int), NumProcStateSlots)); /* pgprocnos */
- return size;
+ /* Allocate space in shared memory */
+ SharedInvalShmemDesc.size = size;
+ ShmemRegisterStruct(&SharedInvalShmemDesc);
}
-/*
- * SharedInvalShmemInit
- * Create and initialize the SI message buffer
- */
-void
-SharedInvalShmemInit(void)
+static void
+SharedInvalShmemInit(void *arg)
{
int i;
- bool found;
-
- /* Allocate space in shared memory */
- shmInvalBuffer = (SISeg *)
- ShmemInitStruct("shmInvalBuffer", SharedInvalShmemSize(), &found);
- if (found)
- return;
- /* Clear message counters, save size of procState array, init spinlock */
+ /* Clear message counters, save size of procState array FIXME, init spinlock */
shmInvalBuffer->minMsgNum = 0;
shmInvalBuffer->maxMsgNum = 0;
shmInvalBuffer->nextThreshold = CLEANUP_MIN;
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 8560a903bc8..96432c633bf 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -73,13 +73,33 @@ PGPROC *MyProc = NULL;
* relatively infrequently (only at backend startup or shutdown) and not for
* very long, so a spinlock is okay.
*/
-NON_EXEC_STATIC slock_t *ProcStructLock = NULL;
+#define ProcStructLock (&ProcGlobal->freeProcsLock)
+
+static void ProcGlobalShmemInit(void *arg);
+
+static ShmemStructDesc ProcGlobalShmemDesc = {
+ .name = "Proc Header",
+ .size = sizeof(PROC_HDR),
+ .init_fn = ProcGlobalShmemInit,
+};
+
+static ShmemStructDesc ProcGlobalAllProcsShmemDesc = {
+ .name = "PGPROC structures",
+ .size = 0, /* dynamic */
+};
+
+static ShmemStructDesc FastPathLockArrayShmemDesc = {
+ .name = "Fast-Path Lock Array",
+ .size = 0, /* dynamic */
+};
/* Pointers to shared-memory structures */
PROC_HDR *ProcGlobal = NULL;
NON_EXEC_STATIC PGPROC *AuxiliaryProcs = NULL;
PGPROC *PreparedXactProcs = NULL;
+static uint32 TotalProcs;
+
/* Is a deadlock check pending? */
static volatile sig_atomic_t got_deadlock_timeout;
@@ -89,24 +109,6 @@ static void AuxiliaryProcKill(int code, Datum arg);
static DeadLockState CheckDeadLock(void);
-/*
- * Report shared-memory space needed by PGPROC.
- */
-static Size
-PGProcShmemSize(void)
-{
- Size size = 0;
- Size TotalProcs =
- add_size(MaxBackends, add_size(NUM_AUXILIARY_PROCS, max_prepared_xacts));
-
- size = add_size(size, mul_size(TotalProcs, sizeof(PGPROC)));
- size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->xids)));
- size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->subxidStates)));
- size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->statusFlags)));
-
- return size;
-}
-
/*
* Report shared-memory space needed by Fast-Path locks.
*/
@@ -114,8 +116,6 @@ static Size
FastPathLockShmemSize(void)
{
Size size = 0;
- Size TotalProcs =
- add_size(MaxBackends, add_size(NUM_AUXILIARY_PROCS, max_prepared_xacts));
Size fpLockBitsSize,
fpRelIdSize;
@@ -131,25 +131,6 @@ FastPathLockShmemSize(void)
return size;
}
-/*
- * Report shared-memory space needed by InitProcGlobal.
- */
-Size
-ProcGlobalShmemSize(void)
-{
- Size size = 0;
-
- /* ProcGlobal */
- size = add_size(size, sizeof(PROC_HDR));
- size = add_size(size, sizeof(slock_t));
-
- size = add_size(size, PGSemaphoreShmemSize(ProcGlobalSemas()));
- size = add_size(size, PGProcShmemSize());
- size = add_size(size, FastPathLockShmemSize());
-
- return size;
-}
-
/*
* Report number of semaphores needed by InitProcGlobal.
*/
@@ -184,35 +165,63 @@ ProcGlobalSemas(void)
* implementation typically requires us to create semaphores in the
* postmaster, not in backends.
*
- * Note: this is NOT called by individual backends under a postmaster,
+ * Note: this is NOT called by individual backends under a postmaster, XXX
* not even in the EXEC_BACKEND case. The ProcGlobal and AuxiliaryProcs
* pointers must be propagated specially for EXEC_BACKEND operation.
*/
void
-InitProcGlobal(void)
+ProcGlobalShmemRegister(void)
+{
+ Size size = 0;
+
+ /*
+ * Reserve all the PGPROC structures we'll need. There are
+ * six separate consumers: (1) normal backends, (2) autovacuum workers and
+ * special workers, (3) background workers, (4) walsenders, (5) auxiliary
+ * processes, and (6) prepared transactions. (For largely-historical
+ * reasons, we combine autovacuum and special workers into one category
+ * with a single freelist.) Each PGPROC structure is dedicated to exactly
+ * one of these purposes, and they do not move between groups.
+ */
+ TotalProcs =
+ add_size(MaxBackends, add_size(NUM_AUXILIARY_PROCS, max_prepared_xacts));
+
+ size = add_size(size, mul_size(TotalProcs, sizeof(PGPROC)));
+
+ /* FIXME: the sizeofs look dangerous because ProcGlobal is not initialized yet */
+ size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->xids)));
+ size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->subxidStates)));
+ size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->statusFlags)));
+
+ ProcGlobalAllProcsShmemDesc.size = size;
+ ShmemRegisterStruct(&ProcGlobalAllProcsShmemDesc);
+
+ FastPathLockArrayShmemDesc.size = FastPathLockShmemSize();
+ ShmemRegisterStruct(&FastPathLockArrayShmemDesc);
+
+ /*
+ * Create the ProcGlobal shared structure last. Its init callback
+ * initializes the others too.
+ */
+ ShmemRegisterStruct(&ProcGlobalShmemDesc);
+}
+
+static void
+ProcGlobalShmemInit(void *arg)
{
+ char *ptr;
+ size_t requestSize;
PGPROC *procs;
int i,
j;
- bool found;
- uint32 TotalProcs = MaxBackends + NUM_AUXILIARY_PROCS + max_prepared_xacts;
-
/* Used for setup of per-backend fast-path slots. */
char *fpPtr,
*fpEndPtr PG_USED_FOR_ASSERTS_ONLY;
Size fpLockBitsSize,
fpRelIdSize;
- Size requestSize;
- char *ptr;
- /* Create the ProcGlobal shared structure */
- ProcGlobal = (PROC_HDR *)
- ShmemInitStruct("Proc Header", sizeof(PROC_HDR), &found);
- Assert(!found);
+ ProcGlobal = ProcGlobalShmemDesc.ptr;
- /*
- * Initialize the data structures.
- */
ProcGlobal->spins_per_delay = DEFAULT_SPINS_PER_DELAY;
dlist_init(&ProcGlobal->freeProcs);
dlist_init(&ProcGlobal->autovacFreeProcs);
@@ -223,23 +232,11 @@ InitProcGlobal(void)
ProcGlobal->checkpointerProc = INVALID_PROC_NUMBER;
pg_atomic_init_u32(&ProcGlobal->procArrayGroupFirst, INVALID_PROC_NUMBER);
pg_atomic_init_u32(&ProcGlobal->clogGroupFirst, INVALID_PROC_NUMBER);
+ SpinLockInit(ProcStructLock);
- /*
- * Create and initialize all the PGPROC structures we'll need. There are
- * six separate consumers: (1) normal backends, (2) autovacuum workers and
- * special workers, (3) background workers, (4) walsenders, (5) auxiliary
- * processes, and (6) prepared transactions. (For largely-historical
- * reasons, we combine autovacuum and special workers into one category
- * with a single freelist.) Each PGPROC structure is dedicated to exactly
- * one of these purposes, and they do not move between groups.
- */
- requestSize = PGProcShmemSize();
-
- ptr = ShmemInitStruct("PGPROC structures",
- requestSize,
- &found);
-
- MemSet(ptr, 0, requestSize);
+ ptr = ProcGlobalAllProcsShmemDesc.ptr;
+ requestSize = ProcGlobalAllProcsShmemDesc.size;
+ memset(ptr, 0, requestSize);
procs = (PGPROC *) ptr;
ptr = ptr + TotalProcs * sizeof(PGPROC);
@@ -275,20 +272,13 @@ InitProcGlobal(void)
fpLockBitsSize = MAXALIGN(FastPathLockGroupsPerBackend * sizeof(uint64));
fpRelIdSize = MAXALIGN(FastPathLockSlotsPerBackend() * sizeof(Oid));
- requestSize = FastPathLockShmemSize();
-
- fpPtr = ShmemInitStruct("Fast-Path Lock Array",
- requestSize,
- &found);
-
- MemSet(fpPtr, 0, requestSize);
+ fpPtr = FastPathLockArrayShmemDesc.ptr;
+ requestSize = FastPathLockArrayShmemDesc.size;
+ memset(fpPtr, 0, requestSize);
/* For asserts checking we did not overflow. */
fpEndPtr = fpPtr + requestSize;
- /* Reserve space for semaphores. */
- PGReserveSemaphores(ProcGlobalSemas());
-
for (i = 0; i < TotalProcs; i++)
{
PGPROC *proc = &procs[i];
@@ -378,12 +368,6 @@ InitProcGlobal(void)
*/
AuxiliaryProcs = &procs[MaxBackends];
PreparedXactProcs = &procs[MaxBackends + NUM_AUXILIARY_PROCS];
-
- /* Create ProcStructLock spinlock, too */
- ProcStructLock = (slock_t *) ShmemInitStruct("ProcStructLock spinlock",
- sizeof(slock_t),
- &found);
- SpinLockInit(ProcStructLock);
}
/*
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 02e9aaa6bca..eed188416ee 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4117,6 +4117,8 @@ PostgresSingleUserMain(int argc, char *argv[],
* shared memory, determine the value of any runtime-computed GUCs that
* depend on the amount of shared memory required.
*/
+ RegisterShmemStructs();
+
InitializeShmemGUCs();
/*
diff --git a/src/include/access/transam.h b/src/include/access/transam.h
index 6fa91bfcdc0..49d476e9d5c 100644
--- a/src/include/access/transam.h
+++ b/src/include/access/transam.h
@@ -15,7 +15,9 @@
#define TRANSAM_H
#include "access/xlogdefs.h"
-
+#ifndef FRONTEND
+#include "storage/shmem.h"
+#endif
/* ----------------
* Special transaction ID values
@@ -330,7 +332,10 @@ TransactionIdFollowsOrEquals(TransactionId id1, TransactionId id2)
extern bool TransactionStartedDuringRecovery(void);
/* in transam/varsup.c */
-extern PGDLLIMPORT TransamVariablesData *TransamVariables;
+#ifndef FRONTEND
+extern PGDLLIMPORT struct ShmemStructDesc TransamVariablesShmemDesc;
+#define TransamVariables ((TransamVariablesData *) TransamVariablesShmemDesc.ptr)
+#endif
/*
* prototypes for functions in transam/transam.c
@@ -345,8 +350,7 @@ extern TransactionId TransactionIdLatest(TransactionId mainxid,
extern XLogRecPtr TransactionIdGetCommitLSN(TransactionId xid);
/* in transam/varsup.c */
-extern Size VarsupShmemSize(void);
-extern void VarsupShmemInit(void);
+extern void VarsupShmemRegister(void);
extern FullTransactionId GetNewTransactionId(bool isSubXact);
extern void AdvanceNextFullTransactionIdPastXid(TransactionId xid);
extern FullTransactionId ReadNextFullTransactionId(void);
diff --git a/src/include/storage/dsm_registry.h b/src/include/storage/dsm_registry.h
index 506fae2c9ca..9a1b4d982af 100644
--- a/src/include/storage/dsm_registry.h
+++ b/src/include/storage/dsm_registry.h
@@ -22,7 +22,6 @@ extern dsa_area *GetNamedDSA(const char *name, bool *found);
extern dshash_table *GetNamedDSHash(const char *name,
const dshash_parameters *params,
bool *found);
-extern Size DSMRegistryShmemSize(void);
-extern void DSMRegistryShmemInit(void);
+extern void DSMRegistryShmemRegister(void);
#endif /* DSM_REGISTRY_H */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index da32787ab51..8a3b71ad5d3 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,6 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
+extern void RegisterShmemStructs(void);
extern Size CalculateShmemSize(void);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 206fb78f8a5..7cdc4852334 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -66,8 +66,7 @@ extern PGDLLIMPORT volatile PMSignalData *PMSignalState;
/*
* prototypes for functions in pmsignal.c
*/
-extern Size PMSignalShmemSize(void);
-extern void PMSignalShmemInit(void);
+extern void PMSignalShmemRegister(void);
extern void SendPostmasterSignal(PMSignalReason reason);
extern bool CheckPostmasterSignal(PMSignalReason reason);
extern void SetQuitSignalReason(QuitSignalReason reason);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 679f0624f92..37023e1a93f 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -418,6 +418,9 @@ typedef struct PROC_HDR
dlist_head bgworkerFreeProcs;
/* Head of list of walsender free PGPROC structures */
dlist_head walsenderFreeProcs;
+
+ slock_t freeProcsLock;
+
/* First pgproc waiting for group XID clear */
pg_atomic_uint32 procArrayGroupFirst;
/* First pgproc waiting for group transaction status update */
@@ -488,7 +491,7 @@ extern PGDLLIMPORT PGPROC *AuxiliaryProcs;
* Function Prototypes
*/
extern int ProcGlobalSemas(void);
-extern Size ProcGlobalShmemSize(void);
+extern void ProcGlobalShmemRegister(void);
extern void InitProcGlobal(void);
extern void InitProcess(void);
extern void InitProcessPhase2(void);
diff --git a/src/include/storage/procarray.h b/src/include/storage/procarray.h
index 3a8593f87ba..41753c3a630 100644
--- a/src/include/storage/procarray.h
+++ b/src/include/storage/procarray.h
@@ -20,8 +20,7 @@
#include "utils/snapshot.h"
-extern Size ProcArrayShmemSize(void);
-extern void ProcArrayShmemInit(void);
+extern void ProcArrayShmemRegister(void);
extern void ProcArrayAdd(PGPROC *proc);
extern void ProcArrayRemove(PGPROC *proc, TransactionId latestXid);
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index e52b8eb7697..f2df1f30c5f 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -71,8 +71,7 @@ typedef enum
/*
* prototypes for functions in procsignal.c
*/
-extern Size ProcSignalShmemSize(void);
-extern void ProcSignalShmemInit(void);
+extern void ProcSignalShmemRegister(void);
extern void ProcSignalInit(const uint8 *cancel_key, int cancel_key_len);
extern int SendProcSignal(pid_t pid, ProcSignalReason reason,
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 89d45287c17..40e2fc17056 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -24,6 +24,53 @@
#include "storage/spin.h"
#include "utils/hsearch.h"
+typedef void (*ShmemInitCallback) (void *arg);
+typedef void (*ShmemAttachCallback) (void *arg);
+
+/*
+ * Descriptor for a named area or struct in shared memory
+ */
+typedef struct ShmemStructDesc
+{
+ /* Name of the shared memory area. Must be unique across the system */
+ const char *name;
+
+ size_t size;
+
+ size_t alignment;
+ ShmemInitCallback init_fn;
+ ShmemInitCallback attach_fn;
+ void *init_fn_arg;
+ void *attach_fn_arg;
+
+ /*
+ * Extra space to allocated in the shared memory segment, but it's not
+ * part of the struct itself. This is used for shared memory hash tables
+ * that can grow beyond the initial size when more buckets are allocated.
+ */
+ size_t extra_size;
+
+ /* Pointer to the shared memory area, when it's allocated. */
+ void *ptr;
+} ShmemStructDesc;
+
+/*
+ * Descriptor for shared memory hash table
+ */
+typedef struct ShmemHashDesc
+{
+ const char *name;
+
+ int hash_flags;
+
+ size_t init_size; /* initial number of entries */
+ size_t max_size; /* max number of entries */
+ HASHCTL *infoP;
+
+ HTAB *ptr;
+
+ ShmemStructDesc base_desc;
+} ShmemHashDesc;
/* shmem.c */
extern PGDLLIMPORT slock_t *ShmemLock;
@@ -34,9 +81,19 @@ extern void *ShmemAlloc(Size size);
extern void *ShmemAllocNoError(Size size);
extern bool ShmemAddrIsValid(const void *addr);
extern void InitShmemIndex(void);
+
+extern void ShmemRegisterHash(ShmemHashDesc *desc, HASHCTL *infoP, int hash_flags);
+extern void ShmemRegisterStruct(ShmemStructDesc *desc);
+
+/* Legacy functions */
extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
HASHCTL *infoP, int hash_flags);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+
+extern size_t ShmemRegisteredSize(void);
+extern void ShmemInitRegistered(void);
+extern void ShmemAttachRegistered(void);
+
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
diff --git a/src/include/storage/sinvaladt.h b/src/include/storage/sinvaladt.h
index a1694500a85..4edba2936e6 100644
--- a/src/include/storage/sinvaladt.h
+++ b/src/include/storage/sinvaladt.h
@@ -28,8 +28,7 @@
/*
* prototypes for functions in sinvaladt.c
*/
-extern Size SharedInvalShmemSize(void);
-extern void SharedInvalShmemInit(void);
+extern void SharedInvalShmemRegister(void);
extern void SharedInvalBackendInit(bool sendOnly);
extern void SIInsertDataEntries(const SharedInvalidationMessage *data, int n);
--
2.47.3
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:14 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
@ 2026-02-10 15:23 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 15:39 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2026-02-10 15:23 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On Tue, Feb 10, 2026 at 3:44 AM Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>
> On 09/02/2026 17:15, Ashutosh Bapat wrote:
> > On Sun, Feb 8, 2026 at 5:14 AM Heikki Linnakangas <hlinnaka@iki.fi> wrote:
> >> Could you write a standalone test module in src/test/modules to
> >> demonstrate how to use the resizable shmem segments, please? That'd
> >> allow focusing on that interface without worrying all the other
> >> complexities of shared buffers.
> >
> > From your writeup it seems like you are leaning towards creating the
> > shared memory segments on-demand rather than having predefined
> > segments as done in the patch.
>
> Right. Same as with normal shared memory structs: there are
> ShmemInitStruct calls scattered throughout the codebase, there's no need
> to predefine them in a single header file or anything like that.
>
> (Note: "on-demand" doesn't mean "on the fly", i.e. each shared memory
> area still needs to be registered at postmaster startup.)
Yes, that's what I have in mind.
>
> > I think your proposal is interesting
> > and might create a possibility for extensions to be able to create
> > resizable shared memory structures.
>
> Ah yes, good point. This facility should certainly be open for
> extensions too. I didn't even realize that the predefined segments
> scheme would not allow that.
>
> > Since the shared memory segments are predefined, a test module, as you
> > suggest, does not have a segment that it can use. That led me to think
> > that you are imagining some kind of on-demand shared memory segments
> > that extensions or test modules can use. I think that will be a useful
> > feature by itself. Just as an example, imagine an extension to provide
> > shared plan cache which resizes the plan cache as needed. However, we
> > need to make sure that these on-demand shared memory segments work
> > well with CalculateShmemSize(), PGSharedMemoryDetach(),
> > AnonymousShmemDetach() etc.
>
> +1
>
> > 2. There is no predefined limit on the number of segments. When
> > CalculateShmemSize() is called, BufferManagerShmemSize() outputs the
> > shared memory required in the main segment and total of shared memory
> > required in the other segments. Similarly shmem_request_hook of each
> > external module outputs the shared memory required in the main segment
> > and the total of shared memory required in the other segments. The
> > main segment is created in CreateSharedMemoryAndSemaphores() directly
> > using the requested size. The other output size is merely used for
> > reporting purposes in InitializeShmemGUCs(). When allocating shared
> > data structures, the main shared memory data structures are allocated
> > using ShmemInitStruct(), whereas the data structures in other memory
> > segments are allocated using ShmemInitStructExt() which takes
> > immediate allocation size and size of address space to be reserved as
> > arguments. It does not require segment_id though. For every new data
> > structure that ShmemInitStructExt() encounters, it a. creates a new
> > shared memory segment using PGSharedMemoryCreate(), b. allocates that
> > structure with the initial size. This function also needs to create
> > the PGShmemInfo and AnonShmemData entries corresponding to new
> > segments and make them available to PGSharedMemoryDetach(),
> > AnonymousShmemDetach() etc. For that we create a shared hash table in
> > the main shared memory segment (just like ShmemIndex) where we store
> > the metadata against each segment name (by coining it from the name of
> > the structure). We expect ShmemInitStructExt() to allocate structures
> > and segments only in the Postmaster and only at the beginning. When a
> > backend starts, it pulls the segment metadata from the shared hash
> > table which is ultimately used by PGSharedMemoryDetach(),
> > AnonymousShmemDetach(). At run time any backend which has access to
> > the shared memory should be able to call ShmemResizeStruct() given a
> > resizable structure and new size. Some higher level synchronization is
> > needed to ensure that the same structure is not resized simultaneously
> > by two backends.
> >
> > The first approach is simple but has limited use given the fixed
> > number of segments. Second is more flexible but that's some work. I am
> > not sure whether it's worth doing all that if there are hardly any
> > extensions which could use resizable shared data structures.
>
> It doesn't seem *that* much more work.
>
> Putting this patch aside for a moment, I don't much like our current
> interface for defining shared memory structs anyway. The
> [SubSystem]ShmemSize() functions feel too detached from the
> ShmemInitStruct() calls. For example, it's an easy mistake to return a
> slightly different size in the FoobarShmemSize() call than what you use
> in the ShmemInitStruct() call. And then there's the fact that the
> initialization functions run both at postmaster startup but also at
> backend startup in EXEC_BACKEND mode. There were reasons for that when
> EXEC_BACKEND was introduced, but it's always felt awkward to me.
>
> I feel that it'd be good to have a single definition of each shmem
> struct, and derive all the other things from there.
>
> Attached is a proof-of-concept of what I have in mind. Don't look too
> closely at how it's implemented, it's very hacky and EXEC_BACKEND mode
> is slightly broken, for example. The point is to demonstrate what the
> callers would look like. I converted only a few subsystems to use the
> new API, the rest still use ShmemInitStruct() and ShmemInitHash().
>
> With this, initialization of a subsystem that defines a shared memory
> area looks like this:
I think this is much better than what we have today. I haven't tried
adding APIs to specify resizable structures, but I will try that
tomorrow unless you want me to wait for this to mature.
I didn't look too deeply but what's broken in the EXEC_BACKEND case?
>
> --------------
>
> /* This struct lives in shared memory */
> typedef struct
> {
> int field;
> } FoobarSharedCtlData;
>
> static void FoobarShmemInit(void *arg);
>
> /* Descriptor for the shared memory area */
> ShmemStructDesc FoobarShmemDesc = {
> .name = "Foobar subsystem",
> .size = sizeof(FoobarSharedCtlData),
> .init_fn = FoobarShmemInit,
> };
>
> /* Pointer to the shared memory struct */
> #define FoobarCtl ((FoobarSharedCtlData *) FoobarShmemDesc.ptr)
>
I don't like this much since it limits the ability to debug. A macro
is not present in the symbol table. How about something like attached?
Once we do that we can make the ShmemStructDesc local to
ShmemRegisterStruct() calls and construct them on the fly. There's no
other use for them.
> /*
> * Register the shared memory struct. This is called once at
> * postmaster startup, before the shared memory segment is allocated,
> * and in EXEC_BACKEND mode also early at backend startup.
> *
> * For core subsystems, there's a list of all these functions in core
> * in ipci.c, similar to all the *ShmemSize() and *ShmemInit() functions
> * today. In an extension, this would be done in _PG_init() or in
> * the shmem_request_hook, replacing the RequestAddinShmemSpace calls
> * we have today.
> */
> void
> FoobarShmemRegister(void)
> {
> ShmemRegisterStruct(&FoobarShmemDesc);
> }
>
> /*
> * This callback is called once at postmaster startup, to initialize
> * the shared memory struct. FoobarShmemDesc.ptr has already been
> * set when this is called.
> */
> static void
> FoobarShmemInit(void *arg)
> {
> memset(FoobarCtl, 0, sizeof(FoobarSharedCtlData));
> FoobarCtl->field = 123;
> }
>
> --------------
>
> The ShmemStructDesc provides room for extending the facility in the
> future. For example, you could specify alignment there, or an additional
> "attach" callback when you need to do more per-backend initialization in
> EXEC_BACKEND mode. And with the resizeable shared memory, a max size.
>
> Thoughts?
LGTM.
One more comment.
/* estimated size of the shmem index table (not a hard limit) */
#define SHMEM_INDEX_SIZE (64)
What do you mean by "not a hard limit"?
We will be limited by the number of entries in the array; a limit that
doesn't exist in the earlier implementation. I mean one can easily
have 100s of extensions loaded, each claiming a couple shared memory
structures. How large do we expect the array to be? Maybe we should
create a linked list which is converted to an array at the end of
registration. The array and its size can be easily passed to the child
through the launch backend as compared to the list. Looking at
launch_backend changes, even that doesn't seem to be required since
SubPostmasterMain() calls RegisterShmemStructs().
Also, do we want to discuss this in a thread of its own?
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] 0002-Get-rid-of-global-shared-memory-pointer-mac-20260210.patch (15.0K, ../../CAExHW5uVQGGn8p3Dq0uTSi8-+Tb83mfkhTzQLj_CSQJqdQ8p+g@mail.gmail.com/2-0002-Get-rid-of-global-shared-memory-pointer-mac-20260210.patch)
download | inline diff:
From 395a95e9934286869b1fe8d45dc5a155ea9be030 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 10 Feb 2026 20:26:31 +0530
Subject: [PATCH 2/2] Get rid of global shared memory pointer macro
declarations
---
.../pg_stat_statements/pg_stat_statements.c | 10 ++++---
src/backend/access/transam/varsup.c | 5 ++++
src/backend/storage/ipc/dsm.c | 5 ++--
src/backend/storage/ipc/dsm_registry.c | 4 ++-
src/backend/storage/ipc/pmsignal.c | 18 +++++-------
src/backend/storage/ipc/procarray.c | 13 +++++----
src/backend/storage/ipc/procsignal.c | 4 ++-
src/backend/storage/ipc/shmem.c | 29 ++++++++++++-------
src/backend/storage/ipc/sinvaladt.c | 7 +++--
src/backend/storage/lmgr/proc.c | 25 ++++++++++------
src/include/access/transam.h | 2 +-
src/include/storage/shmem.h | 6 ++--
12 files changed, 77 insertions(+), 51 deletions(-)
diff --git a/contrib/pg_stat_statements/pg_stat_statements.c b/contrib/pg_stat_statements/pg_stat_statements.c
index 71debc8b47f..73fdf561419 100644
--- a/contrib/pg_stat_statements/pg_stat_statements.c
+++ b/contrib/pg_stat_statements/pg_stat_statements.c
@@ -258,24 +258,26 @@ typedef struct pgssSharedState
pgssGlobalStats stats; /* global statistics for pgss */
} pgssSharedState;
+/* Links to shared memory state */
+pgssSharedState *pgss = NULL;
+HTAB *pgss_hash = NULL;
+
static void pgss_shmem_init(void *arg);
static ShmemStructDesc pgssSharedStateShmemDesc = {
.name = "pg_stat_statements",
.size = sizeof(pgssSharedState),
.init_fn = pgss_shmem_init,
+ .ptr = (void *) &pgss,
};
static ShmemHashDesc pgssSharedHashDesc = {
.name = "pg_stat_statements hash",
.init_size = 0, /* set from 'pgss_max' */
.max_size = 0, /* set from 'pgss_max' */
+ .ptr = &pgss_hash,
};
-/* Links to shared memory state */
-#define pgss ((pgssSharedState *) pgssSharedStateShmemDesc.ptr)
-#define pgss_hash (pgssSharedHashDesc.ptr)
-
/*---- Local variables ----*/
diff --git a/src/backend/access/transam/varsup.c b/src/backend/access/transam/varsup.c
index 11ad90e7372..3dfda875e80 100644
--- a/src/backend/access/transam/varsup.c
+++ b/src/backend/access/transam/varsup.c
@@ -32,10 +32,14 @@
static void VarsupShmemInit(void *arg);
+/* pointer to variables struct in shared memory */
+TransamVariablesData *TransamVariables = NULL;
+
ShmemStructDesc TransamVariablesShmemDesc = {
.name = "TransamVariables",
.size = sizeof(TransamVariablesData),
.init_fn = VarsupShmemInit,
+ .ptr = (void **) &TransamVariables,
};
/*
@@ -49,6 +53,7 @@ VarsupShmemRegister(void)
static void
VarsupShmemInit(void *arg)
+
{
memset(TransamVariables, 0, sizeof(TransamVariablesData));
}
diff --git a/src/backend/storage/ipc/dsm.c b/src/backend/storage/ipc/dsm.c
index 55f46c7687e..73644ec3bbb 100644
--- a/src/backend/storage/ipc/dsm.c
+++ b/src/backend/storage/ipc/dsm.c
@@ -110,14 +110,15 @@ static bool dsm_init_done = false;
/* Preallocated DSM space in the main shared memory region. */
static void dsm_main_space_init(void *);
+static void *dsm_main_space_begin = NULL;
+
static ShmemStructDesc dsm_main_space_shmem_desc = {
.name = "Preallocated DSM",
.size = 0, /* dynamic */
.init_fn = dsm_main_space_init,
+ .ptr = &dsm_main_space_begin,
};
-#define dsm_main_space_begin (dsm_main_space_shmem_desc.ptr)
-
/*
* List of dynamic shared memory segments used by this backend.
*
diff --git a/src/backend/storage/ipc/dsm_registry.c b/src/backend/storage/ipc/dsm_registry.c
index 882af83b7b2..1659e1dd71d 100644
--- a/src/backend/storage/ipc/dsm_registry.c
+++ b/src/backend/storage/ipc/dsm_registry.c
@@ -56,13 +56,15 @@ typedef struct DSMRegistryCtxStruct
static void DSMRegistryCtxShmemInit(void *arg);
+DSMRegistryCtxStruct *DSMRegistryCtx = NULL;
+
static ShmemStructDesc DSMRegistryCtxShmemDesc = {
.name = "DSM Registry Data",
.size = sizeof(DSMRegistryCtxStruct),
.init_fn = DSMRegistryCtxShmemInit,
+ .ptr = (void **) &DSMRegistryCtx,
};
-#define DSMRegistryCtx ((DSMRegistryCtxStruct *) DSMRegistryCtxShmemDesc.ptr)
typedef struct NamedDSMState
{
diff --git a/src/backend/storage/ipc/pmsignal.c b/src/backend/storage/ipc/pmsignal.c
index 23752500d16..3aa0380eadd 100644
--- a/src/backend/storage/ipc/pmsignal.c
+++ b/src/backend/storage/ipc/pmsignal.c
@@ -82,21 +82,17 @@ struct PMSignalData
static void PMSignalShmemInit(void *);
-static ShmemStructDesc PMSignalShmemDesc = {
- .name = "PMSignalState",
- .size = 0, /* dynamic */
- .init_fn = PMSignalShmemInit,
-};
-
/*
* PMSignalState pointer is valid in both postmaster and child processes
- *
- * This is a stand-alone variable rather than just a #define over
- * PMSignalShmemDesc.ptr because it is needed early at backend startup and
- * passed as a backend parameter in EXEC_BACKEND mode
*/
NON_EXEC_STATIC volatile PMSignalData *PMSignalState = NULL;
+static ShmemStructDesc PMSignalShmemDesc = {
+ .name = "PMSignalState",
+ .size = 0, /* dynamic */
+ .init_fn = PMSignalShmemInit,
+ .ptr = (void **) &PMSignalState,
+};
/*
* Local copy of PMSignalState->num_child_flags, only valid in the
@@ -156,7 +152,7 @@ static void
PMSignalShmemInit(void *arg)
{
/* initialize all flags to zeroes */
- PMSignalState = PMSignalShmemDesc.ptr;
+ Assert(PMSignalState);
MemSet(unvolatize(PMSignalData *, PMSignalState), 0, PMSignalShmemDesc.size);
num_child_flags = MaxLivePostmasterChildren();
PMSignalState->num_child_flags = num_child_flags;
diff --git a/src/backend/storage/ipc/procarray.c b/src/backend/storage/ipc/procarray.c
index 08c63bcb2a7..736504d3a3e 100644
--- a/src/backend/storage/ipc/procarray.c
+++ b/src/backend/storage/ipc/procarray.c
@@ -104,15 +104,16 @@ typedef struct ProcArrayStruct
static void ProcArrayShmemInit(void *arg);
static void ProcArrayShmemAttach(void *arg);
+ProcArrayStruct *procArray = NULL;
+
static ShmemStructDesc ProcArrayShmemDesc = {
.name = "Proc Array",
.size = 0, /* dynamic */
.init_fn = ProcArrayShmemInit,
.attach_fn = ProcArrayShmemAttach,
+ .ptr = (void **) &procArray,
};
-#define procArray ((ProcArrayStruct *) ProcArrayShmemDesc.ptr)
-
/*
* State for the GlobalVisTest* family of functions. Those functions can
* e.g. be used to decide if a deleted row can be removed without violating
@@ -290,22 +291,24 @@ static TransactionId cachedXidIsNotInProgress = InvalidTransactionId;
* Bookkeeping for tracking emulated transactions in recovery
*/
+TransactionId *KnownAssignedXids = NULL;
+
static ShmemStructDesc KnownAssignedXidsShmemDesc = {
.name = "KnownAssignedXids",
.size = 0, /* dynamic */
.init_fn = NULL,
+ .ptr = (void **) &KnownAssignedXids,
};
-#define KnownAssignedXids ((TransactionId *) KnownAssignedXidsShmemDesc.ptr)
+bool *KnownAssignedXidsValid = NULL;
static ShmemStructDesc KnownAssignedXidsValidShmemDesc = {
.name = "KnownAssignedXidsValid",
.size = 0, /* dynamic */
.init_fn = NULL,
+ .ptr = (void **) &KnownAssignedXidsValid,
};
-#define KnownAssignedXidsValid ((bool *) KnownAssignedXidsValidShmemDesc.ptr)
-
static TransactionId latestObservedXid = InvalidTransactionId;
/*
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 5743f088324..eec04eae3f4 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -104,13 +104,15 @@ struct ProcSignalHeader
static void ProcSignalShmemInit(void *arg);
+ProcSignalHeader *ProcSignal = NULL;
+
static ShmemStructDesc ProcSignalShmemDesc = {
.name = "ProcSignal",
.size = 0, /* dynamic */
.init_fn = ProcSignalShmemInit,
+ .ptr = (void **) &ProcSignal,
};
-#define ProcSignal ((ProcSignalHeader *) ProcSignalShmemDesc.ptr)
static ProcSignalSlot *MyProcSignalSlot = NULL;
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index faa0fcbd21e..e73ac489b2b 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -118,17 +118,18 @@ static void *ShmemEnd; /* end+1 address of shared memory */
static ShmemAllocatorData *ShmemAllocator;
slock_t *ShmemLock; /* points to ShmemAllocator->shmem_lock */
+ /* primary index hashtable for shmem */
+HTAB *ShmemIndex = NULL;
+
static ShmemHashDesc ShmemIndexHashDesc = {
.name = "ShmemIndex",
.init_size = SHMEM_INDEX_SIZE,
.max_size = SHMEM_INDEX_SIZE,
+ .ptr = &ShmemIndex
};
- /* primary index hashtable for shmem */
-#define ShmemIndex (ShmemIndexHashDesc.ptr)
-
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
@@ -205,7 +206,7 @@ ShmemInitRegistered(void)
result->allocated_size = allocated_size;
result->location = structPtr;
- registry[i]->ptr = structPtr;
+ *(registry[i]->ptr) = structPtr;
if (registry[i]->init_fn)
registry[i]->init_fn(registry[i]->init_fn_arg);
}
@@ -239,7 +240,7 @@ ShmemAttachRegistered(void)
registry[i]->name)));
}
- registry[i]->ptr = result->location;
+ *registry[i]->ptr = result->location;
if (registry[i]->attach_fn)
registry[i]->attach_fn(registry[i]->attach_fn_arg);
@@ -425,10 +426,11 @@ InitShmemIndex(void)
info.keysize = SHMEM_INDEX_KEYSIZE;
info.entrysize = sizeof(ShmemIndexEnt);
- ShmemIndex = ShmemInitHash("ShmemIndex",
+ *ShmemIndexHashDesc.ptr = ShmemInitHash("ShmemIndex",
SHMEM_INDEX_SIZE, SHMEM_INDEX_SIZE,
&info,
HASH_ELEM | HASH_STRINGS);
+ Assert(ShmemIndex != NULL && *ShmemIndexHashDesc.ptr == ShmemIndex);
}
/*
@@ -482,6 +484,12 @@ ShmemRegisterHash(ShmemHashDesc *desc, /* configuration */
desc->base_desc.init_fn_arg = desc;
desc->base_desc.attach_fn = shmem_hash_attach;
desc->base_desc.attach_fn_arg = desc;
+ /*
+ * We need a stable pointer to hold the pointer to the shared memory. Use
+ * the one passed in the descriptor now. It will be replaced with the hash
+ * table header by init or attach function.
+ */
+ desc->base_desc.ptr = (void **) desc->ptr;
desc->base_desc.extra_size = hash_estimate_size(desc->max_size, infoP->entrysize) - desc->base_desc.size;
@@ -499,10 +507,9 @@ shmem_hash_init(void *arg)
int hash_flags = desc->hash_flags;
/* Pass location of hashtable header to hash_create */
- desc->ptr = desc->base_desc.ptr;
- desc->infoP->hctl = (HASHHDR *) desc->ptr;
+ desc->infoP->hctl = (HASHHDR *) *desc->base_desc.ptr;
- desc->ptr = hash_create(desc->name, desc->init_size, desc->infoP, hash_flags);
+ *desc->ptr = hash_create(desc->name, desc->init_size, desc->infoP, hash_flags);
}
static void
@@ -518,9 +525,9 @@ shmem_hash_attach(void *arg)
hash_flags |= HASH_ATTACH;
/* Pass location of hashtable header to hash_create */
- desc->infoP->hctl = (HASHHDR *) desc->ptr;
+ desc->infoP->hctl = (HASHHDR *) *desc->base_desc.ptr;
- desc->ptr = hash_create(desc->name, desc->init_size, desc->infoP, hash_flags);
+ *desc->ptr = hash_create(desc->name, desc->init_size, desc->infoP, hash_flags);
}
/*
diff --git a/src/backend/storage/ipc/sinvaladt.c b/src/backend/storage/ipc/sinvaladt.c
index 0fe0f256971..8321bd9b52d 100644
--- a/src/backend/storage/ipc/sinvaladt.c
+++ b/src/backend/storage/ipc/sinvaladt.c
@@ -205,15 +205,16 @@ typedef struct SISeg
static void SharedInvalShmemInit(void *arg);
+/* pointer to the shared inval buffer */
+SISeg *shmInvalBuffer = NULL;
+
static ShmemStructDesc SharedInvalShmemDesc = {
.name = "shmInvalBuffer",
.size = 0, /* dynamic */
.init_fn = SharedInvalShmemInit,
+ .ptr = (void **) &shmInvalBuffer,
};
-/* pointer to the shared inval buffer */
-#define shmInvalBuffer ((SISeg *) SharedInvalShmemDesc.ptr)
-
static LocalTransactionId nextLocalTransactionId;
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 85375b5195e..a3d6557aa9d 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -77,27 +77,33 @@ PGPROC *MyProc = NULL;
static void ProcGlobalShmemInit(void *arg);
+/* Pointers to shared-memory structures */
+PROC_HDR *ProcGlobal = NULL;
+void *tmpAllProcs = NULL;
+void *tmpFastPathLockArray = NULL;
+NON_EXEC_STATIC PGPROC *AuxiliaryProcs = NULL;
+PGPROC *PreparedXactProcs = NULL;
+
+
static ShmemStructDesc ProcGlobalShmemDesc = {
.name = "Proc Header",
.size = sizeof(PROC_HDR),
.init_fn = ProcGlobalShmemInit,
+ .ptr = (void **) &ProcGlobal,
};
static ShmemStructDesc ProcGlobalAllProcsShmemDesc = {
.name = "PGPROC structures",
.size = 0, /* dynamic */
+ .ptr = (void **) &tmpAllProcs,
};
static ShmemStructDesc FastPathLockArrayShmemDesc = {
.name = "Fast-Path Lock Array",
.size = 0, /* dynamic */
+ .ptr = (void **) &tmpFastPathLockArray,
};
-/* Pointers to shared-memory structures */
-PROC_HDR *ProcGlobal = NULL;
-NON_EXEC_STATIC PGPROC *AuxiliaryProcs = NULL;
-PGPROC *PreparedXactProcs = NULL;
-
static uint32 TotalProcs;
static DeadLockState deadlock_state = DS_NOT_YET_CHECKED;
@@ -222,8 +228,7 @@ ProcGlobalShmemInit(void *arg)
Size fpLockBitsSize,
fpRelIdSize;
- ProcGlobal = ProcGlobalShmemDesc.ptr;
-
+ Assert(ProcGlobal);
ProcGlobal->spins_per_delay = DEFAULT_SPINS_PER_DELAY;
dlist_init(&ProcGlobal->freeProcs);
dlist_init(&ProcGlobal->autovacFreeProcs);
@@ -236,7 +241,8 @@ ProcGlobalShmemInit(void *arg)
pg_atomic_init_u32(&ProcGlobal->clogGroupFirst, INVALID_PROC_NUMBER);
SpinLockInit(ProcStructLock);
- ptr = ProcGlobalAllProcsShmemDesc.ptr;
+ Assert(tmpAllProcs);
+ ptr = tmpAllProcs;
requestSize = ProcGlobalAllProcsShmemDesc.size;
memset(ptr, 0, requestSize);
@@ -274,7 +280,8 @@ ProcGlobalShmemInit(void *arg)
fpLockBitsSize = MAXALIGN(FastPathLockGroupsPerBackend * sizeof(uint64));
fpRelIdSize = MAXALIGN(FastPathLockSlotsPerBackend() * sizeof(Oid));
- fpPtr = FastPathLockArrayShmemDesc.ptr;
+ Assert(tmpFastPathLockArray);
+ fpPtr = tmpFastPathLockArray;
requestSize = FastPathLockArrayShmemDesc.size;
memset(fpPtr, 0, requestSize);
diff --git a/src/include/access/transam.h b/src/include/access/transam.h
index 49d476e9d5c..6e5a546f411 100644
--- a/src/include/access/transam.h
+++ b/src/include/access/transam.h
@@ -334,7 +334,7 @@ extern bool TransactionStartedDuringRecovery(void);
/* in transam/varsup.c */
#ifndef FRONTEND
extern PGDLLIMPORT struct ShmemStructDesc TransamVariablesShmemDesc;
-#define TransamVariables ((TransamVariablesData *) TransamVariablesShmemDesc.ptr)
+extern PGDLLIMPORT TransamVariablesData *TransamVariables;
#endif
/*
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 40e2fc17056..cbd4ef8d03f 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -50,8 +50,8 @@ typedef struct ShmemStructDesc
*/
size_t extra_size;
- /* Pointer to the shared memory area, when it's allocated. */
- void *ptr;
+ /* Pointer to the variable to which pointer to this shared memory area is assigned after allocation. */
+ void **ptr;
} ShmemStructDesc;
/*
@@ -67,7 +67,7 @@ typedef struct ShmemHashDesc
size_t max_size; /* max number of entries */
HASHCTL *infoP;
- HTAB *ptr;
+ HTAB **ptr;
ShmemStructDesc base_desc;
} ShmemHashDesc;
--
2.34.1
[text/x-patch] 0001-wip-Introduce-a-new-way-of-registering-shar-20260210.patch (53.8K, ../../CAExHW5uVQGGn8p3Dq0uTSi8-+Tb83mfkhTzQLj_CSQJqdQ8p+g@mail.gmail.com/3-0001-wip-Introduce-a-new-way-of-registering-shar-20260210.patch)
download | inline diff:
From 49676c5ba088d13236f2c1c66800d7e7b1abbe5f Mon Sep 17 00:00:00 2001
From: Heikki Linnakangas <heikki.linnakangas@iki.fi>
Date: Mon, 9 Feb 2026 22:28:23 +0200
Subject: [PATCH 1/2] wip: Introduce a new way of registering shared memory
structs
---
.../pg_stat_statements/pg_stat_statements.c | 112 ++++-----
src/backend/access/transam/varsup.c | 32 +--
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/postmaster/launch_backend.c | 11 +-
src/backend/postmaster/postmaster.c | 2 +
src/backend/storage/ipc/dsm.c | 46 ++--
src/backend/storage/ipc/dsm_registry.c | 34 ++-
src/backend/storage/ipc/ipci.c | 51 ++--
src/backend/storage/ipc/pmsignal.c | 53 ++--
src/backend/storage/ipc/procarray.c | 127 +++++-----
src/backend/storage/ipc/procsignal.c | 63 +++--
src/backend/storage/ipc/shmem.c | 233 +++++++++++++++++-
src/backend/storage/ipc/sinvaladt.c | 39 +--
src/backend/storage/lmgr/proc.c | 156 ++++++------
src/backend/tcop/postgres.c | 2 +
src/include/access/transam.h | 12 +-
src/include/storage/dsm_registry.h | 3 +-
src/include/storage/ipc.h | 1 +
src/include/storage/pmsignal.h | 3 +-
src/include/storage/proc.h | 5 +-
src/include/storage/procarray.h | 3 +-
src/include/storage/procsignal.h | 3 +-
src/include/storage/shmem.h | 57 +++++
src/include/storage/sinvaladt.h | 3 +-
24 files changed, 665 insertions(+), 388 deletions(-)
diff --git a/contrib/pg_stat_statements/pg_stat_statements.c b/contrib/pg_stat_statements/pg_stat_statements.c
index 4a427533bd8..71debc8b47f 100644
--- a/contrib/pg_stat_statements/pg_stat_statements.c
+++ b/contrib/pg_stat_statements/pg_stat_statements.c
@@ -258,6 +258,25 @@ typedef struct pgssSharedState
pgssGlobalStats stats; /* global statistics for pgss */
} pgssSharedState;
+static void pgss_shmem_init(void *arg);
+
+static ShmemStructDesc pgssSharedStateShmemDesc = {
+ .name = "pg_stat_statements",
+ .size = sizeof(pgssSharedState),
+ .init_fn = pgss_shmem_init,
+};
+
+static ShmemHashDesc pgssSharedHashDesc = {
+ .name = "pg_stat_statements hash",
+ .init_size = 0, /* set from 'pgss_max' */
+ .max_size = 0, /* set from 'pgss_max' */
+};
+
+/* Links to shared memory state */
+#define pgss ((pgssSharedState *) pgssSharedStateShmemDesc.ptr)
+#define pgss_hash (pgssSharedHashDesc.ptr)
+
+
/*---- Local variables ----*/
/* Current nesting depth of planner/ExecutorRun/ProcessUtility calls */
@@ -274,10 +293,6 @@ static ExecutorFinish_hook_type prev_ExecutorFinish = NULL;
static ExecutorEnd_hook_type prev_ExecutorEnd = NULL;
static ProcessUtility_hook_type prev_ProcessUtility = NULL;
-/* Links to shared memory state */
-static pgssSharedState *pgss = NULL;
-static HTAB *pgss_hash = NULL;
-
/*---- GUC variables ----*/
typedef enum
@@ -365,7 +380,6 @@ static void pgss_store(const char *query, int64 queryId,
static void pg_stat_statements_internal(FunctionCallInfo fcinfo,
pgssVersion api_version,
bool showtext);
-static Size pgss_memsize(void);
static pgssEntry *entry_alloc(pgssHashKey *key, Size query_offset, int query_len,
int encoding, bool sticky);
static void entry_dealloc(void);
@@ -500,11 +514,39 @@ _PG_init(void)
static void
pgss_shmem_request(void)
{
+ HASHCTL info;
+
if (prev_shmem_request_hook)
prev_shmem_request_hook();
- RequestAddinShmemSpace(pgss_memsize());
RequestNamedLWLockTranche("pg_stat_statements", 1);
+
+ /*
+ * Register our shared memory state, including hash table
+ */
+ ShmemRegisterStruct(&pgssSharedStateShmemDesc);
+
+ info.keysize = sizeof(pgssHashKey);
+ info.entrysize = sizeof(pgssEntry);
+ pgssSharedHashDesc.init_size = pgss_max;
+ pgssSharedHashDesc.max_size = pgss_max;
+ ShmemRegisterHash(&pgssSharedHashDesc,
+ &info,
+ HASH_ELEM | HASH_BLOBS);
+}
+
+static void
+pgss_shmem_init(void *arg)
+{
+ pgss->lock = &(GetNamedLWLockTranche("pg_stat_statements"))->lock;
+ pgss->cur_median_usage = ASSUMED_MEDIAN_INIT;
+ pgss->mean_query_len = ASSUMED_LENGTH_INIT;
+ SpinLockInit(&pgss->mutex);
+ pgss->extent = 0;
+ pgss->n_writers = 0;
+ pgss->gc_count = 0;
+ pgss->stats.dealloc = 0;
+ pgss->stats.stats_reset = GetCurrentTimestamp();
}
/*
@@ -516,8 +558,6 @@ pgss_shmem_request(void)
static void
pgss_shmem_startup(void)
{
- bool found;
- HASHCTL info;
FILE *file = NULL;
FILE *qfile = NULL;
uint32 header;
@@ -530,42 +570,6 @@ pgss_shmem_startup(void)
if (prev_shmem_startup_hook)
prev_shmem_startup_hook();
- /* reset in case this is a restart within the postmaster */
- pgss = NULL;
- pgss_hash = NULL;
-
- /*
- * Create or attach to the shared memory state, including hash table
- */
- LWLockAcquire(AddinShmemInitLock, LW_EXCLUSIVE);
-
- pgss = ShmemInitStruct("pg_stat_statements",
- sizeof(pgssSharedState),
- &found);
-
- if (!found)
- {
- /* First time through ... */
- pgss->lock = &(GetNamedLWLockTranche("pg_stat_statements"))->lock;
- pgss->cur_median_usage = ASSUMED_MEDIAN_INIT;
- pgss->mean_query_len = ASSUMED_LENGTH_INIT;
- SpinLockInit(&pgss->mutex);
- pgss->extent = 0;
- pgss->n_writers = 0;
- pgss->gc_count = 0;
- pgss->stats.dealloc = 0;
- pgss->stats.stats_reset = GetCurrentTimestamp();
- }
-
- info.keysize = sizeof(pgssHashKey);
- info.entrysize = sizeof(pgssEntry);
- pgss_hash = ShmemInitHash("pg_stat_statements hash",
- pgss_max, pgss_max,
- &info,
- HASH_ELEM | HASH_BLOBS);
-
- LWLockRelease(AddinShmemInitLock);
-
/*
* If we're in the postmaster (or a standalone backend...), set up a shmem
* exit hook to dump the statistics to disk.
@@ -573,12 +577,6 @@ pgss_shmem_startup(void)
if (!IsUnderPostmaster)
on_shmem_exit(pgss_shmem_shutdown, (Datum) 0);
- /*
- * Done if some other process already completed our initialization.
- */
- if (found)
- return;
-
/*
* Note: we don't bother with locks here, because there should be no other
* processes running when this code is reached.
@@ -2082,20 +2080,6 @@ pg_stat_statements_info(PG_FUNCTION_ARGS)
PG_RETURN_DATUM(HeapTupleGetDatum(heap_form_tuple(tupdesc, values, nulls)));
}
-/*
- * Estimate shared memory space needed.
- */
-static Size
-pgss_memsize(void)
-{
- Size size;
-
- size = MAXALIGN(sizeof(pgssSharedState));
- size = add_size(size, hash_estimate_size(pgss_max, sizeof(pgssEntry)));
-
- return size;
-}
-
/*
* Allocate a new hashtable entry.
* caller must hold an exclusive lock on pgss->lock
diff --git a/src/backend/access/transam/varsup.c b/src/backend/access/transam/varsup.c
index 3e95d4cfd16..11ad90e7372 100644
--- a/src/backend/access/transam/varsup.c
+++ b/src/backend/access/transam/varsup.c
@@ -30,35 +30,27 @@
/* Number of OIDs to prefetch (preallocate) per XLOG write */
#define VAR_OID_PREFETCH 8192
-/* pointer to variables struct in shared memory */
-TransamVariablesData *TransamVariables = NULL;
+static void VarsupShmemInit(void *arg);
+ShmemStructDesc TransamVariablesShmemDesc = {
+ .name = "TransamVariables",
+ .size = sizeof(TransamVariablesData),
+ .init_fn = VarsupShmemInit,
+};
/*
* Initialization of shared memory for TransamVariables.
*/
-Size
-VarsupShmemSize(void)
+void
+VarsupShmemRegister(void)
{
- return sizeof(TransamVariablesData);
+ ShmemRegisterStruct(&TransamVariablesShmemDesc);
}
-void
-VarsupShmemInit(void)
+static void
+VarsupShmemInit(void *arg)
{
- bool found;
-
- /* Initialize our shared state struct */
- TransamVariables = ShmemInitStruct("TransamVariables",
- sizeof(TransamVariablesData),
- &found);
- if (!IsUnderPostmaster)
- {
- Assert(!found);
- memset(TransamVariables, 0, sizeof(TransamVariablesData));
- }
- else
- Assert(found);
+ memset(TransamVariables, 0, sizeof(TransamVariablesData));
}
/*
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index 7d32cd0e159..0ded7018e86 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -337,6 +337,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ RegisterShmemStructs();
+
CreateSharedMemoryAndSemaphores();
/*
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index 926fd6f2700..8f638118cdf 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -49,6 +49,7 @@
#include "replication/walreceiver.h"
#include "storage/dsm.h"
#include "storage/io_worker.h"
+#include "storage/ipc.h"
#include "storage/pg_shmem.h"
#include "tcop/backend_startup.h"
#include "utils/memutils.h"
@@ -104,12 +105,10 @@ typedef struct
char **LWLockTrancheNames;
int *LWLockCounter;
LWLockPadded *MainLWLockArray;
- slock_t *ProcStructLock;
PROC_HDR *ProcGlobal;
PGPROC *AuxiliaryProcs;
PGPROC *PreparedXactProcs;
volatile PMSignalData *PMSignalState;
- ProcSignalHeader *ProcSignal;
pid_t PostmasterPid;
TimestampTz PgStartTime;
TimestampTz PgReloadTime;
@@ -678,8 +677,12 @@ SubPostmasterMain(int argc, char *argv[])
/* Restore basic shared memory pointers */
if (UsedShmemSegAddr != NULL)
+ {
InitShmemAllocator(UsedShmemSegAddr);
+ RegisterShmemStructs();
+ }
+
/*
* Run the appropriate Main function
*/
@@ -735,12 +738,10 @@ save_backend_variables(BackendParameters *param,
param->LWLockTrancheNames = LWLockTrancheNames;
param->LWLockCounter = LWLockCounter;
param->MainLWLockArray = MainLWLockArray;
- param->ProcStructLock = ProcStructLock;
param->ProcGlobal = ProcGlobal;
param->AuxiliaryProcs = AuxiliaryProcs;
param->PreparedXactProcs = PreparedXactProcs;
param->PMSignalState = PMSignalState;
- param->ProcSignal = ProcSignal;
param->PostmasterPid = PostmasterPid;
param->PgStartTime = PgStartTime;
@@ -995,12 +996,10 @@ restore_backend_variables(BackendParameters *param)
LWLockTrancheNames = param->LWLockTrancheNames;
LWLockCounter = param->LWLockCounter;
MainLWLockArray = param->MainLWLockArray;
- ProcStructLock = param->ProcStructLock;
ProcGlobal = param->ProcGlobal;
AuxiliaryProcs = param->AuxiliaryProcs;
PreparedXactProcs = param->PreparedXactProcs;
PMSignalState = param->PMSignalState;
- ProcSignal = param->ProcSignal;
PostmasterPid = param->PostmasterPid;
PgStartTime = param->PgStartTime;
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index d6133bfebc6..f6d3369f917 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -968,6 +968,8 @@ PostmasterMain(int argc, char *argv[])
* shared memory, determine the value of any runtime-computed GUCs that
* depend on the amount of shared memory required.
*/
+ RegisterShmemStructs();
+
InitializeShmemGUCs();
/*
diff --git a/src/backend/storage/ipc/dsm.c b/src/backend/storage/ipc/dsm.c
index 6a5b16392f7..55f46c7687e 100644
--- a/src/backend/storage/ipc/dsm.c
+++ b/src/backend/storage/ipc/dsm.c
@@ -108,7 +108,15 @@ static inline bool is_main_region_dsm_handle(dsm_handle handle);
static bool dsm_init_done = false;
/* Preallocated DSM space in the main shared memory region. */
-static void *dsm_main_space_begin = NULL;
+static void dsm_main_space_init(void *);
+
+static ShmemStructDesc dsm_main_space_shmem_desc = {
+ .name = "Preallocated DSM",
+ .size = 0, /* dynamic */
+ .init_fn = dsm_main_space_init,
+};
+
+#define dsm_main_space_begin (dsm_main_space_shmem_desc.ptr)
/*
* List of dynamic shared memory segments used by this backend.
@@ -479,27 +487,29 @@ void
dsm_shmem_init(void)
{
size_t size = dsm_estimate_size();
- bool found;
if (size == 0)
return;
- dsm_main_space_begin = ShmemInitStruct("Preallocated DSM", size, &found);
- if (!found)
- {
- FreePageManager *fpm = (FreePageManager *) dsm_main_space_begin;
- size_t first_page = 0;
- size_t pages;
-
- /* Reserve space for the FreePageManager. */
- while (first_page * FPM_PAGE_SIZE < sizeof(FreePageManager))
- ++first_page;
-
- /* Initialize it and give it all the rest of the space. */
- FreePageManagerInitialize(fpm, dsm_main_space_begin);
- pages = (size / FPM_PAGE_SIZE) - first_page;
- FreePageManagerPut(fpm, first_page, pages);
- }
+ ShmemRegisterStruct(&dsm_main_space_shmem_desc);
+}
+
+static void
+dsm_main_space_init(void *arg)
+{
+ size_t size = dsm_main_space_shmem_desc.size;
+ FreePageManager *fpm = (FreePageManager *) dsm_main_space_begin;
+ size_t first_page = 0;
+ size_t pages;
+
+ /* Reserve space for the FreePageManager. */
+ while (first_page * FPM_PAGE_SIZE < sizeof(FreePageManager))
+ ++first_page;
+
+ /* Initialize it and give it all the rest of the space. */
+ FreePageManagerInitialize(fpm, dsm_main_space_begin);
+ pages = (size / FPM_PAGE_SIZE) - first_page;
+ FreePageManagerPut(fpm, first_page, pages);
}
/*
diff --git a/src/backend/storage/ipc/dsm_registry.c b/src/backend/storage/ipc/dsm_registry.c
index 068c1577b12..882af83b7b2 100644
--- a/src/backend/storage/ipc/dsm_registry.c
+++ b/src/backend/storage/ipc/dsm_registry.c
@@ -54,7 +54,15 @@ typedef struct DSMRegistryCtxStruct
dshash_table_handle dshh;
} DSMRegistryCtxStruct;
-static DSMRegistryCtxStruct *DSMRegistryCtx;
+static void DSMRegistryCtxShmemInit(void *arg);
+
+static ShmemStructDesc DSMRegistryCtxShmemDesc = {
+ .name = "DSM Registry Data",
+ .size = sizeof(DSMRegistryCtxStruct),
+ .init_fn = DSMRegistryCtxShmemInit,
+};
+
+#define DSMRegistryCtx ((DSMRegistryCtxStruct *) DSMRegistryCtxShmemDesc.ptr)
typedef struct NamedDSMState
{
@@ -113,27 +121,17 @@ static const dshash_parameters dsh_params = {
static dsa_area *dsm_registry_dsa;
static dshash_table *dsm_registry_table;
-Size
-DSMRegistryShmemSize(void)
+void
+DSMRegistryShmemRegister(void)
{
- return MAXALIGN(sizeof(DSMRegistryCtxStruct));
+ ShmemRegisterStruct(&DSMRegistryCtxShmemDesc);
}
-void
-DSMRegistryShmemInit(void)
+static void
+DSMRegistryCtxShmemInit(void *)
{
- bool found;
-
- DSMRegistryCtx = (DSMRegistryCtxStruct *)
- ShmemInitStruct("DSM Registry Data",
- DSMRegistryShmemSize(),
- &found);
-
- if (!found)
- {
- DSMRegistryCtx->dsah = DSA_HANDLE_INVALID;
- DSMRegistryCtx->dshh = DSHASH_HANDLE_INVALID;
- }
+ DSMRegistryCtx->dsah = DSA_HANDLE_INVALID;
+ DSMRegistryCtx->dshh = DSHASH_HANDLE_INVALID;
}
/*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 1f7e933d500..952988645d0 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -101,13 +101,14 @@ CalculateShmemSize(void)
size = add_size(size, hash_estimate_size(SHMEM_INDEX_SIZE,
sizeof(ShmemIndexEnt)));
size = add_size(size, dsm_estimate_size());
- size = add_size(size, DSMRegistryShmemSize());
+
+ size = add_size(size, ShmemRegisteredSize());
+
+ /* legacy subsystmes */
size = add_size(size, BufferManagerShmemSize());
size = add_size(size, LockManagerShmemSize());
size = add_size(size, PredicateLockShmemSize());
- size = add_size(size, ProcGlobalShmemSize());
size = add_size(size, XLogPrefetchShmemSize());
- size = add_size(size, VarsupShmemSize());
size = add_size(size, XLOGShmemSize());
size = add_size(size, XLogRecoveryShmemSize());
size = add_size(size, CLOGShmemSize());
@@ -117,11 +118,7 @@ CalculateShmemSize(void)
size = add_size(size, BackgroundWorkerShmemSize());
size = add_size(size, MultiXactShmemSize());
size = add_size(size, LWLockShmemSize());
- size = add_size(size, ProcArrayShmemSize());
size = add_size(size, BackendStatusShmemSize());
- size = add_size(size, SharedInvalShmemSize());
- size = add_size(size, PMSignalShmemSize());
- size = add_size(size, ProcSignalShmemSize());
size = add_size(size, CheckpointerShmemSize());
size = add_size(size, AutoVacuumShmemSize());
size = add_size(size, ReplicationSlotsShmemSize());
@@ -217,6 +214,10 @@ CreateSharedMemoryAndSemaphores(void)
*/
InitShmemAllocator(seghdr);
+ /* Reserve space for semaphores. */
+ if (!IsUnderPostmaster)
+ PGReserveSemaphores(ProcGlobalSemas());
+
/* Initialize subsystems */
CreateOrAttachShmemStructs();
@@ -230,6 +231,19 @@ CreateSharedMemoryAndSemaphores(void)
shmem_startup_hook();
}
+void
+RegisterShmemStructs(void)
+{
+ DSMRegistryShmemRegister();
+
+ ProcGlobalShmemRegister();
+ VarsupShmemRegister();
+ ProcArrayShmemRegister();
+ SharedInvalShmemRegister();
+ PMSignalShmemRegister();
+ ProcSignalShmemRegister();
+}
+
/*
* Initialize various subsystems, setting up their data structures in
* shared memory.
@@ -259,14 +273,23 @@ CreateOrAttachShmemStructs(void)
*/
InitShmemIndex();
+#ifdef EXEC_BACKEND
+ if (IsUnderPostmaster)
+ ShmemAttachRegistered();
+ else
+#endif
+ {
+ ShmemInitRegistered();
+ }
+
dsm_shmem_init();
- DSMRegistryShmemInit();
+ //DSMRegistryShmemInit();
/*
* Set up xlog, clog, and buffers
*/
- VarsupShmemInit();
XLOGShmemInit();
+
XLogPrefetchShmemInit();
XLogRecoveryShmemInit();
CLOGShmemInit();
@@ -288,23 +311,13 @@ CreateOrAttachShmemStructs(void)
/*
* Set up process table
*/
- if (!IsUnderPostmaster)
- InitProcGlobal();
- ProcArrayShmemInit();
BackendStatusShmemInit();
TwoPhaseShmemInit();
BackgroundWorkerShmemInit();
- /*
- * Set up shared-inval messaging
- */
- SharedInvalShmemInit();
-
/*
* Set up interprocess signaling mechanisms
*/
- PMSignalShmemInit();
- ProcSignalShmemInit();
CheckpointerShmemInit();
AutoVacuumShmemInit();
ReplicationSlotsShmemInit();
diff --git a/src/backend/storage/ipc/pmsignal.c b/src/backend/storage/ipc/pmsignal.c
index 4618820b337..23752500d16 100644
--- a/src/backend/storage/ipc/pmsignal.c
+++ b/src/backend/storage/ipc/pmsignal.c
@@ -80,9 +80,24 @@ struct PMSignalData
sig_atomic_t PMChildFlags[FLEXIBLE_ARRAY_MEMBER];
};
-/* PMSignalState pointer is valid in both postmaster and child processes */
+static void PMSignalShmemInit(void *);
+
+static ShmemStructDesc PMSignalShmemDesc = {
+ .name = "PMSignalState",
+ .size = 0, /* dynamic */
+ .init_fn = PMSignalShmemInit,
+};
+
+/*
+ * PMSignalState pointer is valid in both postmaster and child processes
+ *
+ * This is a stand-alone variable rather than just a #define over
+ * PMSignalShmemDesc.ptr because it is needed early at backend startup and
+ * passed as a backend parameter in EXEC_BACKEND mode
+ */
NON_EXEC_STATIC volatile PMSignalData *PMSignalState = NULL;
+
/*
* Local copy of PMSignalState->num_child_flags, only valid in the
* postmaster. Postmaster keeps a local copy so that it doesn't need to
@@ -123,39 +138,28 @@ postmaster_death_handler(SIGNAL_ARGS)
static void MarkPostmasterChildInactive(int code, Datum arg);
/*
- * PMSignalShmemSize
- * Compute space needed for pmsignal.c's shared memory
+ * PMSignalShmemRegister - Register our shared memory
*/
-Size
-PMSignalShmemSize(void)
+void
+PMSignalShmemRegister(void)
{
Size size;
size = offsetof(PMSignalData, PMChildFlags);
size = add_size(size, mul_size(MaxLivePostmasterChildren(),
sizeof(sig_atomic_t)));
-
- return size;
+ PMSignalShmemDesc.size = size;
+ ShmemRegisterStruct(&PMSignalShmemDesc);
}
-/*
- * PMSignalShmemInit - initialize during shared-memory creation
- */
-void
-PMSignalShmemInit(void)
+static void
+PMSignalShmemInit(void *arg)
{
- bool found;
-
- PMSignalState = (PMSignalData *)
- ShmemInitStruct("PMSignalState", PMSignalShmemSize(), &found);
-
- if (!found)
- {
- /* initialize all flags to zeroes */
- MemSet(unvolatize(PMSignalData *, PMSignalState), 0, PMSignalShmemSize());
- num_child_flags = MaxLivePostmasterChildren();
- PMSignalState->num_child_flags = num_child_flags;
- }
+ /* initialize all flags to zeroes */
+ PMSignalState = PMSignalShmemDesc.ptr;
+ MemSet(unvolatize(PMSignalData *, PMSignalState), 0, PMSignalShmemDesc.size);
+ num_child_flags = MaxLivePostmasterChildren();
+ PMSignalState->num_child_flags = num_child_flags;
}
/*
@@ -291,6 +295,7 @@ RegisterPostmasterChildActive(void)
{
int slot = MyPMChildSlot;
+ Assert(PMSignalState);
Assert(slot > 0 && slot <= PMSignalState->num_child_flags);
slot--;
Assert(PMSignalState->PMChildFlags[slot] == PM_CHILD_ASSIGNED);
diff --git a/src/backend/storage/ipc/procarray.c b/src/backend/storage/ipc/procarray.c
index 301f54fb5a8..08c63bcb2a7 100644
--- a/src/backend/storage/ipc/procarray.c
+++ b/src/backend/storage/ipc/procarray.c
@@ -101,6 +101,18 @@ typedef struct ProcArrayStruct
int pgprocnos[FLEXIBLE_ARRAY_MEMBER];
} ProcArrayStruct;
+static void ProcArrayShmemInit(void *arg);
+static void ProcArrayShmemAttach(void *arg);
+
+static ShmemStructDesc ProcArrayShmemDesc = {
+ .name = "Proc Array",
+ .size = 0, /* dynamic */
+ .init_fn = ProcArrayShmemInit,
+ .attach_fn = ProcArrayShmemAttach,
+};
+
+#define procArray ((ProcArrayStruct *) ProcArrayShmemDesc.ptr)
+
/*
* State for the GlobalVisTest* family of functions. Those functions can
* e.g. be used to decide if a deleted row can be removed without violating
@@ -267,9 +279,6 @@ typedef enum KAXCompressReason
KAX_STARTUP_PROCESS_IDLE, /* startup process is about to sleep */
} KAXCompressReason;
-
-static ProcArrayStruct *procArray;
-
static PGPROC *allProcs;
/*
@@ -280,8 +289,23 @@ static TransactionId cachedXidIsNotInProgress = InvalidTransactionId;
/*
* Bookkeeping for tracking emulated transactions in recovery
*/
-static TransactionId *KnownAssignedXids;
-static bool *KnownAssignedXidsValid;
+
+static ShmemStructDesc KnownAssignedXidsShmemDesc = {
+ .name = "KnownAssignedXids",
+ .size = 0, /* dynamic */
+ .init_fn = NULL,
+};
+
+#define KnownAssignedXids ((TransactionId *) KnownAssignedXidsShmemDesc.ptr)
+
+static ShmemStructDesc KnownAssignedXidsValidShmemDesc = {
+ .name = "KnownAssignedXidsValid",
+ .size = 0, /* dynamic */
+ .init_fn = NULL,
+};
+
+#define KnownAssignedXidsValid ((bool *) KnownAssignedXidsValidShmemDesc.ptr)
+
static TransactionId latestObservedXid = InvalidTransactionId;
/*
@@ -372,18 +396,19 @@ static inline FullTransactionId FullXidRelativeTo(FullTransactionId rel,
static void GlobalVisUpdateApply(ComputeXidHorizonsResult *horizons);
/*
- * Report shared-memory space needed by ProcArrayShmemInit
+ * Register the shared PGPROC array during postmaster startup.
*/
-Size
-ProcArrayShmemSize(void)
+void
+ProcArrayShmemRegister(void)
{
- Size size;
-
- /* Size of the ProcArray structure itself */
#define PROCARRAY_MAXPROCS (MaxBackends + max_prepared_xacts)
- size = offsetof(ProcArrayStruct, pgprocnos);
- size = add_size(size, mul_size(sizeof(int), PROCARRAY_MAXPROCS));
+ /* Create or attach to the ProcArray shared structure */
+ ProcArrayShmemDesc.size =
+ add_size(offsetof(ProcArrayStruct, pgprocnos),
+ mul_size(sizeof(int),
+ PROCARRAY_MAXPROCS));
+ ShmemRegisterStruct(&ProcArrayShmemDesc);
/*
* During Hot Standby processing we have a data structure called
@@ -403,64 +428,38 @@ ProcArrayShmemSize(void)
if (EnableHotStandby)
{
- size = add_size(size,
- mul_size(sizeof(TransactionId),
- TOTAL_MAX_CACHED_SUBXIDS));
- size = add_size(size,
- mul_size(sizeof(bool), TOTAL_MAX_CACHED_SUBXIDS));
+ KnownAssignedXidsShmemDesc.size =
+ mul_size(sizeof(TransactionId),
+ TOTAL_MAX_CACHED_SUBXIDS);
+ ShmemRegisterStruct(&KnownAssignedXidsShmemDesc);
+
+ KnownAssignedXidsValidShmemDesc.size =
+ mul_size(sizeof(bool), TOTAL_MAX_CACHED_SUBXIDS);
+ ShmemRegisterStruct(&KnownAssignedXidsValidShmemDesc);
}
-
- return size;
}
-/*
- * Initialize the shared PGPROC array during postmaster startup.
- */
-void
-ProcArrayShmemInit(void)
+static void
+ProcArrayShmemInit(void *arg)
{
- bool found;
-
- /* Create or attach to the ProcArray shared structure */
- procArray = (ProcArrayStruct *)
- ShmemInitStruct("Proc Array",
- add_size(offsetof(ProcArrayStruct, pgprocnos),
- mul_size(sizeof(int),
- PROCARRAY_MAXPROCS)),
- &found);
-
- if (!found)
- {
- /*
- * We're the first - initialize.
- */
- procArray->numProcs = 0;
- procArray->maxProcs = PROCARRAY_MAXPROCS;
- procArray->maxKnownAssignedXids = TOTAL_MAX_CACHED_SUBXIDS;
- procArray->numKnownAssignedXids = 0;
- procArray->tailKnownAssignedXids = 0;
- procArray->headKnownAssignedXids = 0;
- procArray->lastOverflowedXid = InvalidTransactionId;
- procArray->replication_slot_xmin = InvalidTransactionId;
- procArray->replication_slot_catalog_xmin = InvalidTransactionId;
- TransamVariables->xactCompletionCount = 1;
- }
+ procArray->numProcs = 0;
+ procArray->maxProcs = PROCARRAY_MAXPROCS;
+ procArray->maxKnownAssignedXids = TOTAL_MAX_CACHED_SUBXIDS;
+ procArray->numKnownAssignedXids = 0;
+ procArray->tailKnownAssignedXids = 0;
+ procArray->headKnownAssignedXids = 0;
+ procArray->lastOverflowedXid = InvalidTransactionId;
+ procArray->replication_slot_xmin = InvalidTransactionId;
+ procArray->replication_slot_catalog_xmin = InvalidTransactionId;
+ TransamVariables->xactCompletionCount = 1;
allProcs = ProcGlobal->allProcs;
+}
- /* Create or attach to the KnownAssignedXids arrays too, if needed */
- if (EnableHotStandby)
- {
- KnownAssignedXids = (TransactionId *)
- ShmemInitStruct("KnownAssignedXids",
- mul_size(sizeof(TransactionId),
- TOTAL_MAX_CACHED_SUBXIDS),
- &found);
- KnownAssignedXidsValid = (bool *)
- ShmemInitStruct("KnownAssignedXidsValid",
- mul_size(sizeof(bool), TOTAL_MAX_CACHED_SUBXIDS),
- &found);
- }
+static void
+ProcArrayShmemAttach(void *arg)
+{
+ allProcs = ProcGlobal->allProcs;
}
/*
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 8e56922dcea..5743f088324 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -102,7 +102,16 @@ struct ProcSignalHeader
#define BARRIER_CLEAR_BIT(flags, type) \
((flags) &= ~(((uint32) 1) << (uint32) (type)))
-NON_EXEC_STATIC ProcSignalHeader *ProcSignal = NULL;
+static void ProcSignalShmemInit(void *arg);
+
+static ShmemStructDesc ProcSignalShmemDesc = {
+ .name = "ProcSignal",
+ .size = 0, /* dynamic */
+ .init_fn = ProcSignalShmemInit,
+};
+
+#define ProcSignal ((ProcSignalHeader *) ProcSignalShmemDesc.ptr)
+
static ProcSignalSlot *MyProcSignalSlot = NULL;
static bool CheckProcSignal(ProcSignalReason reason);
@@ -110,51 +119,37 @@ static void CleanupProcSignalState(int status, Datum arg);
static void ResetProcSignalBarrierBits(uint32 flags);
/*
- * ProcSignalShmemSize
- * Compute space needed for ProcSignal's shared memory
+ * ProcSignalShmemRegister
+ * Register ProcSignal's shared memory needs at postmaster startup
*/
-Size
-ProcSignalShmemSize(void)
+void
+ProcSignalShmemRegister(void)
{
Size size;
size = mul_size(NumProcSignalSlots, sizeof(ProcSignalSlot));
size = add_size(size, offsetof(ProcSignalHeader, psh_slot));
- return size;
+
+ ProcSignalShmemDesc.size = size;
+ ShmemRegisterStruct(&ProcSignalShmemDesc);
}
-/*
- * ProcSignalShmemInit
- * Allocate and initialize ProcSignal's shared memory
- */
-void
-ProcSignalShmemInit(void)
+static void
+ProcSignalShmemInit(void *arg)
{
- Size size = ProcSignalShmemSize();
- bool found;
+ pg_atomic_init_u64(&ProcSignal->psh_barrierGeneration, 0);
- ProcSignal = (ProcSignalHeader *)
- ShmemInitStruct("ProcSignal", size, &found);
-
- /* If we're first, initialize. */
- if (!found)
+ for (int i = 0; i < NumProcSignalSlots; ++i)
{
- int i;
-
- pg_atomic_init_u64(&ProcSignal->psh_barrierGeneration, 0);
+ ProcSignalSlot *slot = &ProcSignal->psh_slot[i];
- for (i = 0; i < NumProcSignalSlots; ++i)
- {
- ProcSignalSlot *slot = &ProcSignal->psh_slot[i];
-
- SpinLockInit(&slot->pss_mutex);
- pg_atomic_init_u32(&slot->pss_pid, 0);
- slot->pss_cancel_key_len = 0;
- MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
- pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
- pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
- ConditionVariableInit(&slot->pss_barrierCV);
- }
+ SpinLockInit(&slot->pss_mutex);
+ pg_atomic_init_u32(&slot->pss_pid, 0);
+ slot->pss_cancel_key_len = 0;
+ MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags));
+ pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX);
+ pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0);
+ ConditionVariableInit(&slot->pss_barrierCV);
}
}
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 9f362ce8641..faa0fcbd21e 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -19,6 +19,8 @@
* methods). The routines in this file are used for allocating and
* binding to shared memory data structures.
*
+ * FIXME: NOTES below are outdated
+ *
* NOTES:
* (a) There are three kinds of shared memory data structures
* available to POSTGRES: fixed-size structures, queues and hash
@@ -76,6 +78,16 @@
#include "storage/spin.h"
#include "utils/builtins.h"
+/* size constants for the shmem index table */
+ /* max size of data structure string name */
+#define SHMEM_INDEX_KEYSIZE (48)
+ /* estimated size of the shmem index table (not a hard limit) */
+#define SHMEM_INDEX_SIZE (64)
+
+/* these are in postmaster private memory */
+static ShmemStructDesc *registry[SHMEM_INDEX_SIZE];
+static int num_registrations = 0;
+
/*
* This is the first data structure stored in the shared memory segment, at
* the offset that PGShmemHeader->content_offset points to. Allocations by
@@ -95,6 +107,9 @@ typedef struct ShmemAllocatorData
static void *ShmemAllocRaw(Size size, Size *allocated_size);
+static void shmem_hash_init(void *arg);
+static void shmem_hash_attach(void *arg);
+
/* shared memory global variables */
static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */
@@ -103,13 +118,137 @@ static void *ShmemEnd; /* end+1 address of shared memory */
static ShmemAllocatorData *ShmemAllocator;
slock_t *ShmemLock; /* points to ShmemAllocator->shmem_lock */
-static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */
+
+
+static ShmemHashDesc ShmemIndexHashDesc = {
+ .name = "ShmemIndex",
+ .init_size = SHMEM_INDEX_SIZE,
+ .max_size = SHMEM_INDEX_SIZE,
+};
+
+ /* primary index hashtable for shmem */
+#define ShmemIndex (ShmemIndexHashDesc.ptr)
+
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
static bool firstNumaTouch = true;
Datum pg_numa_available(PG_FUNCTION_ARGS);
+
+void
+ShmemRegisterStruct(ShmemStructDesc *desc)
+{
+ elog(DEBUG2, "REGISTER: %s with size %zd", desc->name, desc->size);
+
+ registry[num_registrations++] = desc;
+}
+
+size_t
+ShmemRegisteredSize(void)
+{
+ size_t size;
+
+ size = 0;
+ for (int i = 0; i < num_registrations; i++)
+ {
+ size = add_size(size, registry[i]->size);
+ size = add_size(size, registry[i]->extra_size);
+ }
+
+ elog(DEBUG2, "SIZE: total %zd", size);
+
+ return size;
+}
+
+void
+ShmemInitRegistered(void)
+{
+ /* Should be called only by the postmaster or a standalone backend. */
+ Assert(!IsUnderPostmaster);
+
+ for (int i = 0; i < num_registrations; i++)
+ {
+ size_t allocated_size;
+ void *structPtr;
+ bool found;
+ ShmemIndexEnt *result;
+
+ elog(DEBUG2, "INIT [%d/%d]: %s", i, num_registrations, registry[i]->name);
+
+ /* look it up in the shmem index */
+ result = (ShmemIndexEnt *)
+ hash_search(ShmemIndex, registry[i]->name, HASH_ENTER_NULL, &found);
+ if (!result)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("could not create ShmemIndex entry for data structure \"%s\"",
+ registry[i]->name)));
+ }
+ if (found)
+ elog(ERROR, "shmem struct \"%s\" is already initialized", registry[i]->name);
+
+ /* allocate and initialize it */
+ structPtr = ShmemAllocRaw(registry[i]->size, &allocated_size);
+ if (structPtr == NULL)
+ {
+ /* out of memory; remove the failed ShmemIndex entry */
+ hash_search(ShmemIndex, registry[i]->name, HASH_REMOVE, NULL);
+ ereport(ERROR,
+ (errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("not enough shared memory for data structure"
+ " \"%s\" (%zu bytes requested)",
+ registry[i]->name, registry[i]->size)));
+ }
+ result->size = registry[i]->size;
+ result->allocated_size = allocated_size;
+ result->location = structPtr;
+
+ registry[i]->ptr = structPtr;
+ if (registry[i]->init_fn)
+ registry[i]->init_fn(registry[i]->init_fn_arg);
+ }
+}
+
+#ifdef EXEC_BACKEND
+void
+ShmemAttachRegistered(void)
+{
+ /* Must be initializing a (non-standalone) backend */
+ Assert(IsUnderPostmaster);
+ Assert(ShmemAllocator->index != NULL);
+
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+
+ for (int i = 0; i < num_registrations; i++)
+ {
+ bool found;
+ ShmemIndexEnt *result;
+
+ elog(LOG, "ATTACH [%d/%d]: %s", i, num_registrations, registry[i]->name);
+
+ /* look it up in the shmem index */
+ result = (ShmemIndexEnt *)
+ hash_search(ShmemIndex, registry[i]->name, HASH_FIND, &found);
+ if (!found)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("could not find ShmemIndex entry for data structure \"%s\"",
+ registry[i]->name)));
+ }
+
+ registry[i]->ptr = result->location;
+
+ if (registry[i]->attach_fn)
+ registry[i]->attach_fn(registry[i]->attach_fn_arg);
+ }
+
+ LWLockRelease(ShmemIndexLock);
+}
+#endif
+
/*
* InitShmemAllocator() --- set up basic pointers to shared memory.
*
@@ -292,6 +431,98 @@ InitShmemIndex(void)
HASH_ELEM | HASH_STRINGS);
}
+/*
+ * ShmemInitHash -- Create and initialize, or attach to, a
+ * shared memory hash table.
+ *
+ * We assume caller is doing some kind of synchronization
+ * so that two processes don't try to create/initialize the same
+ * table at once. (In practice, all creations are done in the postmaster
+ * process; child processes should always be attaching to existing tables.)
+ *
+ * max_size is the estimated maximum number of hashtable entries. This is
+ * not a hard limit, but the access efficiency will degrade if it is
+ * exceeded substantially (since it's used to compute directory size and
+ * the hash table buckets will get overfull).
+ *
+ * init_size is the number of hashtable entries to preallocate. For a table
+ * whose maximum size is certain, this should be equal to max_size; that
+ * ensures that no run-time out-of-shared-memory failures can occur.
+ *
+ * *infoP and hash_flags must specify at least the entry sizes and key
+ * comparison semantics (see hash_create()). Flag bits and values specific
+ * to shared-memory hash tables are added here, except that callers may
+ * choose to specify HASH_PARTITION and/or HASH_FIXED_SIZE.
+ *
+ * Note: before Postgres 9.0, this function returned NULL for some failure
+ * cases. Now, it always throws error instead, so callers need not check
+ * for NULL.
+ */
+void
+ShmemRegisterHash(ShmemHashDesc *desc, /* configuration */
+ HASHCTL *infoP, /* info about key and bucket size */
+ int hash_flags) /* info about infoP */
+{
+ /*
+ * Hash tables allocated in shared memory have a fixed directory; it can't
+ * grow or other backends wouldn't be able to find it. So, make sure we
+ * make it big enough to start with.
+ *
+ * The shared memory allocator must be specified too.
+ */
+ infoP->dsize = infoP->max_dsize = hash_select_dirsize(desc->max_size);
+ infoP->alloc = ShmemAllocNoError;
+ hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE;
+
+ /* look it up in the shmem index */
+ memset(&desc->base_desc, 0, sizeof(desc->base_desc));
+ desc->base_desc.name = desc->name;
+ desc->base_desc.size = hash_get_shared_size(infoP, hash_flags);
+ desc->base_desc.init_fn = shmem_hash_init;
+ desc->base_desc.init_fn_arg = desc;
+ desc->base_desc.attach_fn = shmem_hash_attach;
+ desc->base_desc.attach_fn_arg = desc;
+
+ desc->base_desc.extra_size = hash_estimate_size(desc->max_size, infoP->entrysize) - desc->base_desc.size;
+
+ desc->hash_flags = hash_flags;
+ desc->infoP = MemoryContextAlloc(TopMemoryContext, sizeof(HASHCTL));
+ memcpy(desc->infoP, infoP, sizeof(HASHCTL));
+
+ ShmemRegisterStruct(&desc->base_desc);
+}
+
+static void
+shmem_hash_init(void *arg)
+{
+ ShmemHashDesc *desc = (ShmemHashDesc *) arg;
+ int hash_flags = desc->hash_flags;
+
+ /* Pass location of hashtable header to hash_create */
+ desc->ptr = desc->base_desc.ptr;
+ desc->infoP->hctl = (HASHHDR *) desc->ptr;
+
+ desc->ptr = hash_create(desc->name, desc->init_size, desc->infoP, hash_flags);
+}
+
+static void
+shmem_hash_attach(void *arg)
+{
+ ShmemHashDesc *desc = (ShmemHashDesc *) arg;
+ int hash_flags = desc->hash_flags;
+
+ /*
+ * if it already exists, attach to it rather than allocate and initialize
+ * new space
+ */
+ hash_flags |= HASH_ATTACH;
+
+ /* Pass location of hashtable header to hash_create */
+ desc->infoP->hctl = (HASHHDR *) desc->ptr;
+
+ desc->ptr = hash_create(desc->name, desc->init_size, desc->infoP, hash_flags);
+}
+
/*
* ShmemInitHash -- Create and initialize, or attach to, a
* shared memory hash table.
diff --git a/src/backend/storage/ipc/sinvaladt.c b/src/backend/storage/ipc/sinvaladt.c
index a7a7cc4f0a9..0fe0f256971 100644
--- a/src/backend/storage/ipc/sinvaladt.c
+++ b/src/backend/storage/ipc/sinvaladt.c
@@ -203,7 +203,16 @@ typedef struct SISeg
*/
#define NumProcStateSlots (MaxBackends + NUM_AUXILIARY_PROCS)
-static SISeg *shmInvalBuffer; /* pointer to the shared inval buffer */
+static void SharedInvalShmemInit(void *arg);
+
+static ShmemStructDesc SharedInvalShmemDesc = {
+ .name = "shmInvalBuffer",
+ .size = 0, /* dynamic */
+ .init_fn = SharedInvalShmemInit,
+};
+
+/* pointer to the shared inval buffer */
+#define shmInvalBuffer ((SISeg *) SharedInvalShmemDesc.ptr)
static LocalTransactionId nextLocalTransactionId;
@@ -212,10 +221,11 @@ static void CleanupInvalidationState(int status, Datum arg);
/*
- * SharedInvalShmemSize --- return shared-memory space needed
+ * SharedInvalShmemRegister
+ * Register shared memory needs for the SI message buffer
*/
-Size
-SharedInvalShmemSize(void)
+void
+SharedInvalShmemRegister(void)
{
Size size;
@@ -223,26 +233,17 @@ SharedInvalShmemSize(void)
size = add_size(size, mul_size(sizeof(ProcState), NumProcStateSlots)); /* procState */
size = add_size(size, mul_size(sizeof(int), NumProcStateSlots)); /* pgprocnos */
- return size;
+ /* Allocate space in shared memory */
+ SharedInvalShmemDesc.size = size;
+ ShmemRegisterStruct(&SharedInvalShmemDesc);
}
-/*
- * SharedInvalShmemInit
- * Create and initialize the SI message buffer
- */
-void
-SharedInvalShmemInit(void)
+static void
+SharedInvalShmemInit(void *arg)
{
int i;
- bool found;
-
- /* Allocate space in shared memory */
- shmInvalBuffer = (SISeg *)
- ShmemInitStruct("shmInvalBuffer", SharedInvalShmemSize(), &found);
- if (found)
- return;
- /* Clear message counters, save size of procState array, init spinlock */
+ /* Clear message counters, save size of procState array FIXME, init spinlock */
shmInvalBuffer->minMsgNum = 0;
shmInvalBuffer->maxMsgNum = 0;
shmInvalBuffer->nextThreshold = CLEANUP_MIN;
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index c7a001b3b79..85375b5195e 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -73,13 +73,33 @@ PGPROC *MyProc = NULL;
* relatively infrequently (only at backend startup or shutdown) and not for
* very long, so a spinlock is okay.
*/
-NON_EXEC_STATIC slock_t *ProcStructLock = NULL;
+#define ProcStructLock (&ProcGlobal->freeProcsLock)
+
+static void ProcGlobalShmemInit(void *arg);
+
+static ShmemStructDesc ProcGlobalShmemDesc = {
+ .name = "Proc Header",
+ .size = sizeof(PROC_HDR),
+ .init_fn = ProcGlobalShmemInit,
+};
+
+static ShmemStructDesc ProcGlobalAllProcsShmemDesc = {
+ .name = "PGPROC structures",
+ .size = 0, /* dynamic */
+};
+
+static ShmemStructDesc FastPathLockArrayShmemDesc = {
+ .name = "Fast-Path Lock Array",
+ .size = 0, /* dynamic */
+};
/* Pointers to shared-memory structures */
PROC_HDR *ProcGlobal = NULL;
NON_EXEC_STATIC PGPROC *AuxiliaryProcs = NULL;
PGPROC *PreparedXactProcs = NULL;
+static uint32 TotalProcs;
+
static DeadLockState deadlock_state = DS_NOT_YET_CHECKED;
/* Is a deadlock check pending? */
@@ -91,24 +111,6 @@ static void AuxiliaryProcKill(int code, Datum arg);
static void CheckDeadLock(void);
-/*
- * Report shared-memory space needed by PGPROC.
- */
-static Size
-PGProcShmemSize(void)
-{
- Size size = 0;
- Size TotalProcs =
- add_size(MaxBackends, add_size(NUM_AUXILIARY_PROCS, max_prepared_xacts));
-
- size = add_size(size, mul_size(TotalProcs, sizeof(PGPROC)));
- size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->xids)));
- size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->subxidStates)));
- size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->statusFlags)));
-
- return size;
-}
-
/*
* Report shared-memory space needed by Fast-Path locks.
*/
@@ -116,8 +118,6 @@ static Size
FastPathLockShmemSize(void)
{
Size size = 0;
- Size TotalProcs =
- add_size(MaxBackends, add_size(NUM_AUXILIARY_PROCS, max_prepared_xacts));
Size fpLockBitsSize,
fpRelIdSize;
@@ -133,25 +133,6 @@ FastPathLockShmemSize(void)
return size;
}
-/*
- * Report shared-memory space needed by InitProcGlobal.
- */
-Size
-ProcGlobalShmemSize(void)
-{
- Size size = 0;
-
- /* ProcGlobal */
- size = add_size(size, sizeof(PROC_HDR));
- size = add_size(size, sizeof(slock_t));
-
- size = add_size(size, PGSemaphoreShmemSize(ProcGlobalSemas()));
- size = add_size(size, PGProcShmemSize());
- size = add_size(size, FastPathLockShmemSize());
-
- return size;
-}
-
/*
* Report number of semaphores needed by InitProcGlobal.
*/
@@ -186,35 +167,63 @@ ProcGlobalSemas(void)
* implementation typically requires us to create semaphores in the
* postmaster, not in backends.
*
- * Note: this is NOT called by individual backends under a postmaster,
+ * Note: this is NOT called by individual backends under a postmaster, XXX
* not even in the EXEC_BACKEND case. The ProcGlobal and AuxiliaryProcs
* pointers must be propagated specially for EXEC_BACKEND operation.
*/
void
-InitProcGlobal(void)
+ProcGlobalShmemRegister(void)
+{
+ Size size = 0;
+
+ /*
+ * Reserve all the PGPROC structures we'll need. There are
+ * six separate consumers: (1) normal backends, (2) autovacuum workers and
+ * special workers, (3) background workers, (4) walsenders, (5) auxiliary
+ * processes, and (6) prepared transactions. (For largely-historical
+ * reasons, we combine autovacuum and special workers into one category
+ * with a single freelist.) Each PGPROC structure is dedicated to exactly
+ * one of these purposes, and they do not move between groups.
+ */
+ TotalProcs =
+ add_size(MaxBackends, add_size(NUM_AUXILIARY_PROCS, max_prepared_xacts));
+
+ size = add_size(size, mul_size(TotalProcs, sizeof(PGPROC)));
+
+ /* FIXME: the sizeofs look dangerous because ProcGlobal is not initialized yet */
+ size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->xids)));
+ size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->subxidStates)));
+ size = add_size(size, mul_size(TotalProcs, sizeof(*ProcGlobal->statusFlags)));
+
+ ProcGlobalAllProcsShmemDesc.size = size;
+ ShmemRegisterStruct(&ProcGlobalAllProcsShmemDesc);
+
+ FastPathLockArrayShmemDesc.size = FastPathLockShmemSize();
+ ShmemRegisterStruct(&FastPathLockArrayShmemDesc);
+
+ /*
+ * Create the ProcGlobal shared structure last. Its init callback
+ * initializes the others too.
+ */
+ ShmemRegisterStruct(&ProcGlobalShmemDesc);
+}
+
+static void
+ProcGlobalShmemInit(void *arg)
{
+ char *ptr;
+ size_t requestSize;
PGPROC *procs;
int i,
j;
- bool found;
- uint32 TotalProcs = MaxBackends + NUM_AUXILIARY_PROCS + max_prepared_xacts;
-
/* Used for setup of per-backend fast-path slots. */
char *fpPtr,
*fpEndPtr PG_USED_FOR_ASSERTS_ONLY;
Size fpLockBitsSize,
fpRelIdSize;
- Size requestSize;
- char *ptr;
- /* Create the ProcGlobal shared structure */
- ProcGlobal = (PROC_HDR *)
- ShmemInitStruct("Proc Header", sizeof(PROC_HDR), &found);
- Assert(!found);
+ ProcGlobal = ProcGlobalShmemDesc.ptr;
- /*
- * Initialize the data structures.
- */
ProcGlobal->spins_per_delay = DEFAULT_SPINS_PER_DELAY;
dlist_init(&ProcGlobal->freeProcs);
dlist_init(&ProcGlobal->autovacFreeProcs);
@@ -225,23 +234,11 @@ InitProcGlobal(void)
ProcGlobal->checkpointerProc = INVALID_PROC_NUMBER;
pg_atomic_init_u32(&ProcGlobal->procArrayGroupFirst, INVALID_PROC_NUMBER);
pg_atomic_init_u32(&ProcGlobal->clogGroupFirst, INVALID_PROC_NUMBER);
+ SpinLockInit(ProcStructLock);
- /*
- * Create and initialize all the PGPROC structures we'll need. There are
- * six separate consumers: (1) normal backends, (2) autovacuum workers and
- * special workers, (3) background workers, (4) walsenders, (5) auxiliary
- * processes, and (6) prepared transactions. (For largely-historical
- * reasons, we combine autovacuum and special workers into one category
- * with a single freelist.) Each PGPROC structure is dedicated to exactly
- * one of these purposes, and they do not move between groups.
- */
- requestSize = PGProcShmemSize();
-
- ptr = ShmemInitStruct("PGPROC structures",
- requestSize,
- &found);
-
- MemSet(ptr, 0, requestSize);
+ ptr = ProcGlobalAllProcsShmemDesc.ptr;
+ requestSize = ProcGlobalAllProcsShmemDesc.size;
+ memset(ptr, 0, requestSize);
procs = (PGPROC *) ptr;
ptr = ptr + TotalProcs * sizeof(PGPROC);
@@ -277,20 +274,13 @@ InitProcGlobal(void)
fpLockBitsSize = MAXALIGN(FastPathLockGroupsPerBackend * sizeof(uint64));
fpRelIdSize = MAXALIGN(FastPathLockSlotsPerBackend() * sizeof(Oid));
- requestSize = FastPathLockShmemSize();
-
- fpPtr = ShmemInitStruct("Fast-Path Lock Array",
- requestSize,
- &found);
-
- MemSet(fpPtr, 0, requestSize);
+ fpPtr = FastPathLockArrayShmemDesc.ptr;
+ requestSize = FastPathLockArrayShmemDesc.size;
+ memset(fpPtr, 0, requestSize);
/* For asserts checking we did not overflow. */
fpEndPtr = fpPtr + requestSize;
- /* Reserve space for semaphores. */
- PGReserveSemaphores(ProcGlobalSemas());
-
for (i = 0; i < TotalProcs; i++)
{
PGPROC *proc = &procs[i];
@@ -380,12 +370,6 @@ InitProcGlobal(void)
*/
AuxiliaryProcs = &procs[MaxBackends];
PreparedXactProcs = &procs[MaxBackends + NUM_AUXILIARY_PROCS];
-
- /* Create ProcStructLock spinlock, too */
- ProcStructLock = (slock_t *) ShmemInitStruct("ProcStructLock spinlock",
- sizeof(slock_t),
- &found);
- SpinLockInit(ProcStructLock);
}
/*
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 02e9aaa6bca..eed188416ee 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4117,6 +4117,8 @@ PostgresSingleUserMain(int argc, char *argv[],
* shared memory, determine the value of any runtime-computed GUCs that
* depend on the amount of shared memory required.
*/
+ RegisterShmemStructs();
+
InitializeShmemGUCs();
/*
diff --git a/src/include/access/transam.h b/src/include/access/transam.h
index 6fa91bfcdc0..49d476e9d5c 100644
--- a/src/include/access/transam.h
+++ b/src/include/access/transam.h
@@ -15,7 +15,9 @@
#define TRANSAM_H
#include "access/xlogdefs.h"
-
+#ifndef FRONTEND
+#include "storage/shmem.h"
+#endif
/* ----------------
* Special transaction ID values
@@ -330,7 +332,10 @@ TransactionIdFollowsOrEquals(TransactionId id1, TransactionId id2)
extern bool TransactionStartedDuringRecovery(void);
/* in transam/varsup.c */
-extern PGDLLIMPORT TransamVariablesData *TransamVariables;
+#ifndef FRONTEND
+extern PGDLLIMPORT struct ShmemStructDesc TransamVariablesShmemDesc;
+#define TransamVariables ((TransamVariablesData *) TransamVariablesShmemDesc.ptr)
+#endif
/*
* prototypes for functions in transam/transam.c
@@ -345,8 +350,7 @@ extern TransactionId TransactionIdLatest(TransactionId mainxid,
extern XLogRecPtr TransactionIdGetCommitLSN(TransactionId xid);
/* in transam/varsup.c */
-extern Size VarsupShmemSize(void);
-extern void VarsupShmemInit(void);
+extern void VarsupShmemRegister(void);
extern FullTransactionId GetNewTransactionId(bool isSubXact);
extern void AdvanceNextFullTransactionIdPastXid(TransactionId xid);
extern FullTransactionId ReadNextFullTransactionId(void);
diff --git a/src/include/storage/dsm_registry.h b/src/include/storage/dsm_registry.h
index 506fae2c9ca..9a1b4d982af 100644
--- a/src/include/storage/dsm_registry.h
+++ b/src/include/storage/dsm_registry.h
@@ -22,7 +22,6 @@ extern dsa_area *GetNamedDSA(const char *name, bool *found);
extern dshash_table *GetNamedDSHash(const char *name,
const dshash_parameters *params,
bool *found);
-extern Size DSMRegistryShmemSize(void);
-extern void DSMRegistryShmemInit(void);
+extern void DSMRegistryShmemRegister(void);
#endif /* DSM_REGISTRY_H */
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index da32787ab51..8a3b71ad5d3 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -77,6 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
/* ipci.c */
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
+extern void RegisterShmemStructs(void);
extern Size CalculateShmemSize(void);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h
index 206fb78f8a5..7cdc4852334 100644
--- a/src/include/storage/pmsignal.h
+++ b/src/include/storage/pmsignal.h
@@ -66,8 +66,7 @@ extern PGDLLIMPORT volatile PMSignalData *PMSignalState;
/*
* prototypes for functions in pmsignal.c
*/
-extern Size PMSignalShmemSize(void);
-extern void PMSignalShmemInit(void);
+extern void PMSignalShmemRegister(void);
extern void SendPostmasterSignal(PMSignalReason reason);
extern bool CheckPostmasterSignal(PMSignalReason reason);
extern void SetQuitSignalReason(QuitSignalReason reason);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 679f0624f92..37023e1a93f 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -418,6 +418,9 @@ typedef struct PROC_HDR
dlist_head bgworkerFreeProcs;
/* Head of list of walsender free PGPROC structures */
dlist_head walsenderFreeProcs;
+
+ slock_t freeProcsLock;
+
/* First pgproc waiting for group XID clear */
pg_atomic_uint32 procArrayGroupFirst;
/* First pgproc waiting for group transaction status update */
@@ -488,7 +491,7 @@ extern PGDLLIMPORT PGPROC *AuxiliaryProcs;
* Function Prototypes
*/
extern int ProcGlobalSemas(void);
-extern Size ProcGlobalShmemSize(void);
+extern void ProcGlobalShmemRegister(void);
extern void InitProcGlobal(void);
extern void InitProcess(void);
extern void InitProcessPhase2(void);
diff --git a/src/include/storage/procarray.h b/src/include/storage/procarray.h
index 3a8593f87ba..41753c3a630 100644
--- a/src/include/storage/procarray.h
+++ b/src/include/storage/procarray.h
@@ -20,8 +20,7 @@
#include "utils/snapshot.h"
-extern Size ProcArrayShmemSize(void);
-extern void ProcArrayShmemInit(void);
+extern void ProcArrayShmemRegister(void);
extern void ProcArrayAdd(PGPROC *proc);
extern void ProcArrayRemove(PGPROC *proc, TransactionId latestXid);
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index e52b8eb7697..f2df1f30c5f 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -71,8 +71,7 @@ typedef enum
/*
* prototypes for functions in procsignal.c
*/
-extern Size ProcSignalShmemSize(void);
-extern void ProcSignalShmemInit(void);
+extern void ProcSignalShmemRegister(void);
extern void ProcSignalInit(const uint8 *cancel_key, int cancel_key_len);
extern int SendProcSignal(pid_t pid, ProcSignalReason reason,
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 89d45287c17..40e2fc17056 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -24,6 +24,53 @@
#include "storage/spin.h"
#include "utils/hsearch.h"
+typedef void (*ShmemInitCallback) (void *arg);
+typedef void (*ShmemAttachCallback) (void *arg);
+
+/*
+ * Descriptor for a named area or struct in shared memory
+ */
+typedef struct ShmemStructDesc
+{
+ /* Name of the shared memory area. Must be unique across the system */
+ const char *name;
+
+ size_t size;
+
+ size_t alignment;
+ ShmemInitCallback init_fn;
+ ShmemInitCallback attach_fn;
+ void *init_fn_arg;
+ void *attach_fn_arg;
+
+ /*
+ * Extra space to allocated in the shared memory segment, but it's not
+ * part of the struct itself. This is used for shared memory hash tables
+ * that can grow beyond the initial size when more buckets are allocated.
+ */
+ size_t extra_size;
+
+ /* Pointer to the shared memory area, when it's allocated. */
+ void *ptr;
+} ShmemStructDesc;
+
+/*
+ * Descriptor for shared memory hash table
+ */
+typedef struct ShmemHashDesc
+{
+ const char *name;
+
+ int hash_flags;
+
+ size_t init_size; /* initial number of entries */
+ size_t max_size; /* max number of entries */
+ HASHCTL *infoP;
+
+ HTAB *ptr;
+
+ ShmemStructDesc base_desc;
+} ShmemHashDesc;
/* shmem.c */
extern PGDLLIMPORT slock_t *ShmemLock;
@@ -34,9 +81,19 @@ extern void *ShmemAlloc(Size size);
extern void *ShmemAllocNoError(Size size);
extern bool ShmemAddrIsValid(const void *addr);
extern void InitShmemIndex(void);
+
+extern void ShmemRegisterHash(ShmemHashDesc *desc, HASHCTL *infoP, int hash_flags);
+extern void ShmemRegisterStruct(ShmemStructDesc *desc);
+
+/* Legacy functions */
extern HTAB *ShmemInitHash(const char *name, int64 init_size, int64 max_size,
HASHCTL *infoP, int hash_flags);
extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr);
+
+extern size_t ShmemRegisteredSize(void);
+extern void ShmemInitRegistered(void);
+extern void ShmemAttachRegistered(void);
+
extern Size add_size(Size s1, Size s2);
extern Size mul_size(Size s1, Size s2);
diff --git a/src/include/storage/sinvaladt.h b/src/include/storage/sinvaladt.h
index a1694500a85..4edba2936e6 100644
--- a/src/include/storage/sinvaladt.h
+++ b/src/include/storage/sinvaladt.h
@@ -28,8 +28,7 @@
/*
* prototypes for functions in sinvaladt.c
*/
-extern Size SharedInvalShmemSize(void);
-extern void SharedInvalShmemInit(void);
+extern void SharedInvalShmemRegister(void);
extern void SharedInvalBackendInit(bool sendOnly);
extern void SIInsertDataEntries(const SharedInvalidationMessage *data, int n);
base-commit: c67bef3f3252a3a38bf347f9f119944176a796ce
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:14 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-10 15:23 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-10 15:39 ` Heikki Linnakangas <hlinnaka@iki.fi>
0 siblings, 0 replies; 167+ messages in thread
From: Heikki Linnakangas @ 2026-02-10 15:39 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On 10/02/2026 17:23, Ashutosh Bapat wrote:
> I didn't look too deeply but what's broken in the EXEC_BACKEND case?
Dunno, I didn't look too deeply either :-D. But I think it's just some
silly bug, i.e. the design should work with EXEC_BACKEND.
>> --------------
>>
>> /* This struct lives in shared memory */
>> typedef struct
>> {
>> int field;
>> } FoobarSharedCtlData;
>>
>> static void FoobarShmemInit(void *arg);
>>
>> /* Descriptor for the shared memory area */
>> ShmemStructDesc FoobarShmemDesc = {
>> .name = "Foobar subsystem",
>> .size = sizeof(FoobarSharedCtlData),
>> .init_fn = FoobarShmemInit,
>> };
>>
>> /* Pointer to the shared memory struct */
>> #define FoobarCtl ((FoobarSharedCtlData *) FoobarShmemDesc.ptr)
>>
>
> I don't like this much since it limits the ability to debug. A macro
> is not present in the symbol table.
True, although FoobarShmemDesc would still be present in the symbol table.
> How about something like attached?
> Once we do that we can make the ShmemStructDesc local to
> ShmemRegisterStruct() calls and construct them on the fly. There's no
> other use for them.
Works for me.
> LGTM.
>
> One more comment.
> /* estimated size of the shmem index table (not a hard limit) */
> #define SHMEM_INDEX_SIZE (64)
>
> What do you mean by "not a hard limit"?
That was just copy-pasted from existing code. (Don't look too deeply ;-) )
> We will be limited by the number of entries in the array; a limit that
> doesn't exist in the earlier implementation. I mean one can easily
> have 100s of extensions loaded, each claiming a couple shared memory
> structures. How large do we expect the array to be? Maybe we should
> create a linked list which is converted to an array at the end of
> registration. The array and its size can be easily passed to the child
> through the launch backend as compared to the list. Looking at
> launch_backend changes, even that doesn't seem to be required since
> SubPostmasterMain() calls RegisterShmemStructs().
Or you can also just repalloc() a larger array if it fills up.
> Also, do we want to discuss this in a thread of its own?
Makes sense.
- Heikki
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:14 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
@ 2026-02-12 20:42 ` Andres Freund <andres@anarazel.de>
1 sibling, 0 replies; 167+ messages in thread
From: Andres Freund @ 2026-02-12 20:42 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
Hi,
On 2026-02-10 00:14:23 +0200, Heikki Linnakangas wrote:
> Putting this patch aside for a moment, I don't much like our current
> interface for defining shared memory structs anyway. The
> [SubSystem]ShmemSize() functions feel too detached from the
> ShmemInitStruct() calls.
Very much agreed.
Besides the other weaknesses, it's particularly annoying to use in extensions,
because extensions need to chain the shmem init hook.
I also really dislike that we support unattributed shared memory
allocations. For one they end up interspersing memory from different
subsystems, for another the lack of attributability is bad; it sure is weird
that one can't see the amount of memory used by the buffer mapping table or
the lock table.
> I feel that it'd be good to have a single definition of each shmem struct,
> and derive all the other things from there.
>
> Attached is a proof-of-concept of what I have in mind. Don't look too
> closely at how it's implemented, it's very hacky and EXEC_BACKEND mode is
> slightly broken, for example. The point is to demonstrate what the callers
> would look like. I converted only a few subsystems to use the new API, the
> rest still use ShmemInitStruct() and ShmemInitHash().
>
> With this, initialization of a subsystem that defines a shared memory area
> looks like this:
>
> --------------
>
> /* This struct lives in shared memory */
> typedef struct
> {
> int field;
> } FoobarSharedCtlData;
>
> static void FoobarShmemInit(void *arg);
>
> /* Descriptor for the shared memory area */
> ShmemStructDesc FoobarShmemDesc = {
> .name = "Foobar subsystem",
> .size = sizeof(FoobarSharedCtlData),
> .init_fn = FoobarShmemInit,
> };
I wonder if the size determination should be a callback too. That way
extensions wouldn't need to bother with shmem_request_hook, as introduced in
commit 4f2400cb3f1
Author: Robert Haas <rhaas@postgresql.org>
Date: 2022-05-13 09:31:06 -0400
Add a new shmem_request_hook hook.
Currently, preloaded libraries are expected to request additional
shared memory and LWLocks in _PG_init(). However, it is not unusal
for such requests to depend on MaxBackends, which won't be
initialized at that time. Such requests could also depend on GUCs
that other modules might change. This introduces a new hook where
modules can safely use MaxBackends and GUCs to request additional
shared memory and LWLocks.
...
And it'd allow us to make the ShmemStructDesc variables const.
Eventually I'd really like to not have multiple lists of core subsystems
(e.g. currently in CalculateShmemSize(), BaseInit() & InitPostgres(), each
top-level sigsetjmp() block, ...) but drive it all through one table. This
seems like a nice step towards that.
> The ShmemStructDesc provides room for extending the facility in the future.
> For example, you could specify alignment there, or an additional "attach"
> callback when you need to do more per-backend initialization in EXEC_BACKEND
> mode. And with the resizeable shared memory, a max size.
And perhaps information about how to deal with NUMAn (e.g. interleave the proc
array, but use more complicated behaviour for buffer blocks) and whether to
use huge pages (it's not clear it's a win for e.g. pgproc).
> Thoughts?
There's stuff to quibble with in the details (being a POC), but I think as a
whole this would be quite an improvement.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-09 22:37 ` Andres Freund <andres@anarazel.de>
2026-02-10 05:32 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Andres Freund @ 2026-02-09 22:37 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
Hi,
On 2026-02-09 20:45:28 +0530, Ashutosh Bapat wrote:
> 2. Address space reservation for shared memory
> ============================================
>
> Currently the shared memory layout is designed to pack everything tight
> together, leaving no space between mappings for resizing. Here is how it
> looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
> anonymous shared memory we talk about:
>
> 00400000-00490000 /path/bin/postgres
> ...
> 012d9000-0133e000 [heap]
> 7f443a800000-7f470a800000 /dev/zero (deleted)
> 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
> 7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
> ...
>
> Make the layout more dynamic via splitting every shared memory segment
> into two parts:
>
> * An anonymous file, which actually contains shared memory content.
> Such an anonymous file is created via memfd_create, it lives in
> memory, behaves like a regular file and semantically equivalent to an
> anonymous memory allocated via mmap with MAP_ANONYMOUS.
>
> * A reservation mapping, which size is much larger than required shared
> segment size. This mapping is created with flag MAP_NORESERVE (to not
> count the reserved space against memory limits). The anonymous file is
> mapped into this reservation mapping.
>
> If we have to change the address maps while resizing the shared buffer
> pool, it is needed to be done in Postmaster too, so that the new
> backends will inherit the resized address space from the Postmaster.
> However, Postmaster is not invovled in ProcSignalBarrier mechanism and
> we don't want it to spend time in things other than its core
> functionality. To achive that, maximum required address space maps are
> setup upfront with read and write access when starting the server. When
> resizing the buffer pool only the backing file object is resized from
> the coordinator. This also makes the ProcSignalBarrier handling code
> light for backends other than the coordinator.
>
> The resulting layout looks like this:
>
> 00400000-00490000 /path/bin/postgres
> ...
> 3f526000-3f590000 rw-p [heap]
> 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
> 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
> 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
> 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
>
> To resize a shared memory segment in this layout it's possible to use
> ftruncate on the memory mapped file.
>
> This approach also do not impact the actual memory usage as reported by
> the kernel.
I still don't see what the point of having multiple mappings and using memfd
is. We need to reserve the address space for the maximum sized allocation in
postmaster, otherwise there's absolutely no guarantee that it's available at
those addresses in all the children - which you do as you explain
here. Therefore, the maximum size of each "suballocation" needs to be reserved
ahead of time. At which point I don't see the point of having multiple
mmaps. It just makes things more complicated and expensive (each mmap makes
fork & exit slower).
Even if we decide to use memfd, because we consider MADV_DONTNEED to not be
suitable for some reason, what's the point of having more than one mapping
using memfd?
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:37 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2026-02-10 05:32 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:43 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-12 20:12 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
0 siblings, 2 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2026-02-10 05:32 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
On Tue, Feb 10, 2026 at 4:07 AM Andres Freund <andres@anarazel.de> wrote:
>
> Hi,
>
> On 2026-02-09 20:45:28 +0530, Ashutosh Bapat wrote:
> > 2. Address space reservation for shared memory
> > ============================================
> >
> > Currently the shared memory layout is designed to pack everything tight
> > together, leaving no space between mappings for resizing. Here is how it
> > looks like for one mapping in /proc/$PID/maps, /dev/zero represents the
> > anonymous shared memory we talk about:
> >
> > 00400000-00490000 /path/bin/postgres
> > ...
> > 012d9000-0133e000 [heap]
> > 7f443a800000-7f470a800000 /dev/zero (deleted)
> > 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive
> > 7f4718400000-7f4718401000 /usr/lib64/libstdc++.so.6.0.34
> > ...
> >
> > Make the layout more dynamic via splitting every shared memory segment
> > into two parts:
> >
> > * An anonymous file, which actually contains shared memory content.
> > Such an anonymous file is created via memfd_create, it lives in
> > memory, behaves like a regular file and semantically equivalent to an
> > anonymous memory allocated via mmap with MAP_ANONYMOUS.
> >
> > * A reservation mapping, which size is much larger than required shared
> > segment size. This mapping is created with flag MAP_NORESERVE (to not
> > count the reserved space against memory limits). The anonymous file is
> > mapped into this reservation mapping.
> >
> > If we have to change the address maps while resizing the shared buffer
> > pool, it is needed to be done in Postmaster too, so that the new
> > backends will inherit the resized address space from the Postmaster.
> > However, Postmaster is not invovled in ProcSignalBarrier mechanism and
> > we don't want it to spend time in things other than its core
> > functionality. To achive that, maximum required address space maps are
> > setup upfront with read and write access when starting the server. When
> > resizing the buffer pool only the backing file object is resized from
> > the coordinator. This also makes the ProcSignalBarrier handling code
> > light for backends other than the coordinator.
> >
> > The resulting layout looks like this:
> >
> > 00400000-00490000 /path/bin/postgres
> > ...
> > 3f526000-3f590000 rw-p [heap]
> > 7fbd827fe000-7fbd8bdde000 rw-s /memfd:main (deleted) -- anon file
> > 7fbd8bdde000-7fbe82800000 ---s /memfd:main (deleted) -- reservation
> > 7fbe82800000-7fbe90670000 r--p /usr/lib/locale/locale-archive
> > 7fbe90800000-7fbe90941000 r-xp /usr/lib64/libstdc++.so.6.0.34
I had revised this commit message to reflect the current state, but it
seems this still leaked from the previous commit message. There is
only one mapping for main and not two as seen above. I have just
removed the layout from commit message now. Sorry for the misleading
writeup.
> >
> > To resize a shared memory segment in this layout it's possible to use
> > ftruncate on the memory mapped file.
> >
> > This approach also do not impact the actual memory usage as reported by
> > the kernel.
>
> I still don't see what the point of having multiple mappings and using memfd
> is. We need to reserve the address space for the maximum sized allocation in
> postmaster, otherwise there's absolutely no guarantee that it's available at
> those addresses in all the children - which you do as you explain
> here. Therefore, the maximum size of each "suballocation" needs to be reserved
> ahead of time. At which point I don't see the point of having multiple
> mmaps. It just makes things more complicated and expensive (each mmap makes
> fork & exit slower).
>
> Even if we decide to use memfd, because we consider MADV_DONTNEED to not be
> suitable for some reason, what's the point of having more than one mapping
> using memfd?
There are just two mappings now compared to 6 earlier. If I am reading
Jakub's benchmarking correctly, even 6 segments didn't show much
regression in his benchmarks. Having just two should not see much
regression. If we use multiple mappings we could control the
properties of each segment separately - e.g. use huge pages for some
(buffer blocks) and not for others. In Windows it seems it is easy to
create multiple segments than punching holes in an existing segments.
When we port the feature to Windows or other platforms, being able to
treat all the segments in the same way would be an advantage.
Said that I am not discarding the idea of using a single fd and then
punching holes using fallocate() altogether; we will use it if
multiple mappings do not bring any advantages. Let's also see how the
on-demand shared memory segment feature being discussed in this thread
with Heikki gets shaped.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:37 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2026-02-10 05:32 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-10 14:43 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
1 sibling, 0 replies; 167+ messages in thread
From: Jakub Wartak @ 2026-02-10 14:43 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Heikki Linnakangas <hlinnaka@iki.fi>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
On Tue, Feb 10, 2026 at 6:32 AM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
[..]
> > > To resize a shared memory segment in this layout it's possible to use
> > > ftruncate on the memory mapped file.
> > >
> > > This approach also do not impact the actual memory usage as reported by
> > > the kernel.
> >
> > I still don't see what the point of having multiple mappings and using memfd
> > is. We need to reserve the address space for the maximum sized allocation in
> > postmaster, otherwise there's absolutely no guarantee that it's available at
> > those addresses in all the children - which you do as you explain
> > here. Therefore, the maximum size of each "suballocation" needs to be reserved
> > ahead of time. At which point I don't see the point of having multiple
> > mmaps. It just makes things more complicated and expensive (each mmap makes
> > fork & exit slower).
> >
> > Even if we decide to use memfd, because we consider MADV_DONTNEED to not be
> > suitable for some reason, what's the point of having more than one mapping
> > using memfd?
>
> There are just two mappings now compared to 6 earlier. If I am reading
> Jakub's benchmarking correctly, even 6 segments didn't show much
> regression in his benchmarks. Having just two should not see much
> regression. If we use multiple mappings we could control the
> properties of each segment separately - e.g. use huge pages for some
> (buffer blocks) and not for others.
FWIW, Tomas ended up technically using multiple mmap segments too due to NUMA
and there appears to be no other way (2*NUMA nodes, for at least
Buffer Blocks and
PGPROC as I remember, or maybe it was even 3*nodes??). I hope we
attack that problem
again one day and we can measure the impact again there if needed.
-J.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:37 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2026-02-10 05:32 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-12 20:12 ` Andres Freund <andres@anarazel.de>
2026-02-13 11:49 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Andres Freund @ 2026-02-12 20:12 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
Hi,
On 2026-02-10 11:02:13 +0530, Ashutosh Bapat wrote:
> > I still don't see what the point of having multiple mappings and using memfd
> > is. We need to reserve the address space for the maximum sized allocation in
> > postmaster, otherwise there's absolutely no guarantee that it's available at
> > those addresses in all the children - which you do as you explain
> > here. Therefore, the maximum size of each "suballocation" needs to be reserved
> > ahead of time. At which point I don't see the point of having multiple
> > mmaps. It just makes things more complicated and expensive (each mmap makes
> > fork & exit slower).
> >
> > Even if we decide to use memfd, because we consider MADV_DONTNEED to not be
> > suitable for some reason, what's the point of having more than one mapping
> > using memfd?
(this should reference MADV_REMOVE, not MADV_DONTNEED)
> There are just two mappings now compared to 6 earlier. If I am reading
> Jakub's benchmarking correctly, even 6 segments didn't show much
> regression in his benchmarks. Having just two should not see much
> regression. If we use multiple mappings we could control the
> properties of each segment separately - e.g. use huge pages for some
> (buffer blocks) and not for others. In Windows it seems it is easy to
> create multiple segments than punching holes in an existing segments.
> When we port the feature to Windows or other platforms, being able to
> treat all the segments in the same way would be an advantage.
> Said that I am not discarding the idea of using a single fd and then
> punching holes using fallocate() altogether; we will use it if
> multiple mappings do not bring any advantages. Let's also see how the
> on-demand shared memory segment feature being discussed in this thread
> with Heikki gets shaped.
I think the multiple memory mappings approach is just too restrictive. If we
e.g. eventually want to make some of the other major allocations that depend
on NBuffers react to resizing shared buffers, it's very easy to do if all it
requires is calling
madvise(TYPEALIGN(start, page_size), MADV_REMOVE, TYPEALIGN_DOWN(end, page_size));
There are several cases that are pretty easy to handle that way:
- Buffer Blocks
- Buffer Descriptors
- Sync request queue (part of the "Checkpointer Data" allocation)
- Checkpoint BufferIds (for sorting the to-be-checkpointed data)
- Buffer IO Condition Variables
But if you want to support making these resizable with the separate mappings
approach, it gets considerably more complicated and the number of mappings
increases more substantially.
We also don't need a lot less infrastructure in shmem.c that way. We could
e.g. make ShmemInitStruct() reservere the entire requested size (to avoid OOM
killer issues) and have a ShmemInitStructExt() that allows the caller choose
whether to reserve. No different segment IDs etc are needed.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Re: Changing shared_buffers without restart Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:37 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2026-02-10 05:32 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 20:12 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2026-02-13 11:49 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2026-02-13 11:49 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
On Fri, Feb 13, 2026 at 1:42 AM Andres Freund <andres@anarazel.de> wrote:
>
> Hi,
>
> On 2026-02-10 11:02:13 +0530, Ashutosh Bapat wrote:
> > > I still don't see what the point of having multiple mappings and using memfd
> > > is. We need to reserve the address space for the maximum sized allocation in
> > > postmaster, otherwise there's absolutely no guarantee that it's available at
> > > those addresses in all the children - which you do as you explain
> > > here. Therefore, the maximum size of each "suballocation" needs to be reserved
> > > ahead of time. At which point I don't see the point of having multiple
> > > mmaps. It just makes things more complicated and expensive (each mmap makes
> > > fork & exit slower).
> > >
> > > Even if we decide to use memfd, because we consider MADV_DONTNEED to not be
> > > suitable for some reason, what's the point of having more than one mapping
> > > using memfd?
>
> (this should reference MADV_REMOVE, not MADV_DONTNEED)
>
>
> > There are just two mappings now compared to 6 earlier. If I am reading
> > Jakub's benchmarking correctly, even 6 segments didn't show much
> > regression in his benchmarks. Having just two should not see much
> > regression. If we use multiple mappings we could control the
> > properties of each segment separately - e.g. use huge pages for some
> > (buffer blocks) and not for others. In Windows it seems it is easy to
> > create multiple segments than punching holes in an existing segments.
> > When we port the feature to Windows or other platforms, being able to
> > treat all the segments in the same way would be an advantage.
>
> > Said that I am not discarding the idea of using a single fd and then
> > punching holes using fallocate() altogether; we will use it if
> > multiple mappings do not bring any advantages. Let's also see how the
> > on-demand shared memory segment feature being discussed in this thread
> > with Heikki gets shaped.
>
> I think the multiple memory mappings approach is just too restrictive. If we
> e.g. eventually want to make some of the other major allocations that depend
> on NBuffers react to resizing shared buffers, it's very easy to do if all it
> requires is calling
> madvise(TYPEALIGN(start, page_size), MADV_REMOVE, TYPEALIGN_DOWN(end, page_size));
>
> There are several cases that are pretty easy to handle that way:
> - Buffer Blocks
> - Buffer Descriptors
> - Sync request queue (part of the "Checkpointer Data" allocation)
> - Checkpoint BufferIds (for sorting the to-be-checkpointed data)
> - Buffer IO Condition Variables
>
> But if you want to support making these resizable with the separate mappings
> approach, it gets considerably more complicated and the number of mappings
> increases more substantially.
>
> We also don't need a lot less infrastructure in shmem.c that way. We could
> e.g. make ShmemInitStruct() reservere the entire requested size (to avoid OOM
> killer issues) and have a ShmemInitStructExt() that allows the caller choose
> whether to reserve. No different segment IDs etc are needed.
I have started a new thread to discuss resizable shared memory
structures [1]. Copied this discussion over there. We will come back
to this thread once the discussion there is settled and discuss
specifically resizable buffer pool.
[1] https://www.postgresql.org/message-id/CAExHW5vM1bneLYfg0wGeAa=52UiJ3z4vKd3AJ72X8Fw6k3KKrg@mail.gmail...
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-09 13:41 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-09 14:29 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2 siblings, 2 replies; 167+ messages in thread
From: Jakub Wartak @ 2026-02-09 13:41 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On Wed, Jan 28, 2026 at 2:19 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>v 20260128*.patch
Short intro: I've started trying out these patches for slightly another reason
than the online buffers resize. There's was recent post [1] that was brought to
attention by Alvaro. That article is complaining about postmaster being
unscalable and more or less saturating @ 2-3k new connections / second and
postmaster becoming a CPU hog (one could argue that's too much and not sensible
setup).
I've thought that the potential main reason of the hit would be slow fork(),
so I had an idea why we fork() with majority of memory being shared_buffers
(BufferBlocks) that is not really used inside postmaster itself
(I mean it does not use it, only backends do use it). I've thought it could
be cool if we could just init the memory, leave just the fd from memfd_create
for s_b around (that is unmap() BufferBlocks from the postmaster thus lowering
its RSS/smaps footprint) and then on fork() the fork() would NOT have to copy
that big kernel VMA for shared_buffers. Instead (in theory - only the fd that
is the reference - thereby we could increase the scalability of the postmaster
(kernel would need to perform less work during fork()). Later on, the classic
backends on their side would mmap() the region back from the fd created earlier
(in postmaster) using memfd_create(2), but that would happen as part of many
backends (so workload would be spread across many CPUs). The critical
assumption here is that although on Linux there seems to be huge PMD sharing for
MAP_SHARED | MAP_HUGETLB, I was still wondering if we couldn't accelerate it
further by simply not having at all this memory before calling fork().
Initially I've created simple PoC bench on 64GB even with hugepages showed some
potential:
Scenario 1 (mmap inherited): 20001 total forks, 0.302ms per fork
Scenario 2 (MADV_DONTFORK): 20001 total forks, 0.292ms per fork
Scenario 3 (memfd_create): 20002 total forks, 0.145ms per fork
Quite unexpectedly that's how I discovered Your's and Dimitry's patch
as it already
had separation of memory segments (rather than one big mmap() blob) and
memfd_create(2) used too, so I just gave it a try. So I've tried to benchmark
Your's patchset when it comes to establishing new connections:
1s4c 32GB RAM, 6.14.x kernel, 16GB shared_buffers
benchmark: /usr/pgsql19/bin/pgbench -n --connect -j 4 -c 100
-f <(echo "SELECT 1;") postgres -P 1 -T 30
# master
latency average = 358.681 ms
latency stddev = 225.813 ms
average connection time = 2.989 ms
tps = 1329.733460 (including reconnection times)
# memfd/thispatchset
latency average = 363.584 ms
latency stddev = 230.529 ms
average connection time = 3.022 ms
tps = 1315.810761 (including reconnection times)
# memfd+mytrick, showed some promise in low stddev, but not in TPS
latency average = 34.229 ms
latency stddev = 22.059 ms
average connection time = 2.908 ms
tps = 1369.785773 (including reconnection times)
Another box, 4s32c64, 128GB RAM, 6.14.x kernel,
64GB shared_buffers (4 NUMA nodes)
benchmark: /usr/pgsql19/bin/pgbench -n --connect -j 128 -c 1000
-f <(echo "SELECT 1;") postgres -P 1 -T 30
#master
latency average = 240.179 ms
latency stddev = 119.379 ms
average connection time = 62.049 ms
tps = 2058.434343 (including reconnection times)
#memfd
latency average = 268.384 ms
latency stddev = 133.501 ms
average connection time = 69.081 ms
tps = 1847.422995 (including reconnection times)
#memfd+mytrick
latency average = 261.726 ms
latency stddev = 130.161 ms
average connection time = 67.579 ms
tps = 1889.988400 (including reconnection times)
So:
a) yes, my idea fizzled - still no crystal clear idea why - but at least
I've tried Your's patch :) We are still in the ballpark of ~1800..3000
new connections per second.
and here proper review against patchset follows:
b) the patch changes the behavior on startup and it appears that now
the patch tries to touch all the memory during startup which takes
much more time (I'm thinking of HA failover/promote scenarios where
long startup on could mean trouble e.g. after pg_rewind). E.g. without
patch it takes 1-2s and with the patch it takes 49s, no HugePages with
64GB s_b on slow machine). It happens due to that new fallocate() from
shmem_fallocate(). If it is supposed to stay like that IMHO log should
elog() what it is doing ("allocating memory...", otherwise users can
be left confused. It almost behaves like MAP_POPULATE would be
used.
c) as per above measurements, on NUMA it appears that there's seems be
like 1847/2058=~89% of baseline regression, when it comes to the
establishing new connections and you are operating on sysv_shmem.c
(so affecting all users). Possibly this would have to be re-tested
on some more modern hardware (I don't see it on single socket, but I
see on multiple sockets)
d) MADV_HUGEPAGES is Linux 4.14+ and although released nearly 10
years ago the buildfarm probably has some animals (Ubuntu 16?) that
still use such
old kernels (??))
e) so maybe because of b+c+d we should consider putting it under some new
shared_memory_type in the long run?
e) With huge_pages=on and no asserts it seemed to never work for me due to:
FATAL: segment[main]: could not truncate anonymous file to
size 313483264: Invalid argument
and please see this (this is with both(!)
max_shared_buffers=shared_buffers=1GB),
for some reason ftruncate() ended up calling ~ 2x more.
[pid 1252287] memfd_create("main", MFD_HUGETLB) = 4
[pid 1252287] mmap(NULL, 157286400, PROT_NONE, MAP_SHARED|MAP_NORESE..
[pid 1252287] mprotect(0x7f2a1a400000, 157286400, PROT_READ|PROT_WRI..
[pid 1252287] ftruncate(4, 313483264) = -1 EINVAL (Invalid argument)
it appears that I'm getting this due to bug in
round_off_mapping_sizes_for_hugepages() as before it I'm getting:
shmem_reserved=156196864, shmem_req_size=156196864
and after it it's called it returning:
shmem_reserved=157286400, shmem_req_size=313483264
Maybe TYPE ALIGN() would be a better fit for this there.
-J.
[1] - https://www.recall.ai/blog/postgres-postmaster-does-not-scale
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 0399265c4dd..b80fcd58931 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -758,6 +758,8 @@ round_off_mapping_sizes_for_hugepages(MemoryMappingSizes *mapping, int hugepages
if (hugepagesize == 0)
return;
+ elog(WARNING, "shmem_reserved=%ld, shmem_req_size=%ld", mapping->shmem_reserved, mapping->shmem_req_size);
+
if (mapping->shmem_req_size % hugepagesize != 0)
mapping->shmem_req_size += add_size(mapping->shmem_req_size,
hugepagesize - (mapping->shmem_req_size % hugepagesize));
@@ -839,7 +841,7 @@ CreateAnonymousSegment(int segment_id, MemoryMappingSizes *mapping)
(errmsg("segment[%s]: could not create anonymous shared memory file: %m",
segname)));
- elog(DEBUG1, "segment[%s]: mmap(%zu)", segname, mapping->shmem_req_size);
+ elog(WARNING, "segment[%s]: mmap(%zu)", segname, mapping->shmem_req_size);
/*
* Reserve maximum required address space for future expansion of this
@@ -894,7 +896,6 @@ CreateAnonymousSegment(int segment_id, MemoryMappingSizes *mapping)
anonseg->addr = ptr;
anonseg->size = mapping->shmem_reserved;
}
-
/*
* PrepareHugePages
*
@@ -1418,6 +1419,29 @@ PGSharedMemoryDetach(void)
}
}
+void
+JWPGSharedMemoryBuffersDetachTrick(void)
+{
+ AnonShmemSegment *sbseg = &AnonShmemSegs[BUFFERS_SHMEM_SEGMENT];
+ elog(WARNING, "unmapping s_b, but not closing fd(%d), from postmaster to accelerate fork()", sbseg->fd);
+ if(munmap(sbseg->addr, sbseg->size) != 0) {
+ ereport(FATAL, (errmsg("sb segment: could not unmap anonymous shared memory: %m")));
+ }
+}
+
+void
+JWPGSharedMemoryBuffersReattachTrick(void)
+{
+ AnonShmemSegment *sbseg = &AnonShmemSegs[BUFFERS_SHMEM_SEGMENT];
+
+ if(sbseg->fd == -1)
+ elog(PANIC, "children got wrong memfd fd");
+
+ sbseg->addr = mmap(NULL, sbseg->size, PROT_READ | PROT_WRITE, MAP_SHARED, sbseg->fd, 0);
+ if (sbseg->addr == MAP_FAILED)
+ elog(PANIC, "children failed mmap: %m");
+}
+
void
ShmemControlInit(void)
{
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index 80e3088fc7e..b59c82f6c1f 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -93,6 +93,7 @@ BackgroundWriterMain(const void *startup_data, size_t startup_data_len)
WritebackContext wb_context;
Assert(startup_data_len == 0);
+ elog(WARNING, "bwriter starting");
MyBackendType = B_BG_WRITER;
AuxiliaryProcessMainCommon();
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index 85da8ac381a..43db2570a95 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -227,6 +227,9 @@ postmaster_child_launch(BackendType child_type, int child_slot,
conn_timing.fork_end = GetCurrentTimestamp();
}
+ //elog(WARNING, "starting %d", pid);
+ JWPGSharedMemoryBuffersReattachTrick();
+
/* Close the postmaster's sockets */
ClosePostmasterPorts(child_type == B_LOGGER);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 1e92b0bcc5e..a18cb6227c5 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -257,7 +257,7 @@ CreateSharedMemoryAndSemaphores(void)
inhseg->UsedShmemSegID = i;
/* Compute the size of the shared-memory block */
- elog(DEBUG3, "invoking IpcMemoryCreate(segment %s, size=%zu, reserved address space=%zu)",
+ elog(WARNING, "invoking IpcMemoryCreate(segment %s, size=%zu, reserved address space=%zu)",
MappingName(i), mapping->shmem_req_size, mapping->shmem_reserved);
/*
@@ -289,6 +289,10 @@ CreateSharedMemoryAndSemaphores(void)
*/
if (shmem_startup_hook)
shmem_startup_hook();
+
+ /* JW HACK */
+ JWPGSharedMemoryBuffersDetachTrick();
+
}
/*
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index ac679259787..2160af8de1a 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -219,6 +219,8 @@ extern PGShmemHeader *PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *m
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
+extern void JWPGSharedMemoryBuffersDetachTrick(void);
+extern void JWPGSharedMemoryBuffersReattachTrick(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
extern bool PGSharedMemoryResize(int segment_id, MemoryMappingSizes *mapping_sizes);
Attachments:
[text/plain] jw_no_sb_in_postmaster__noeffect.txt (4.5K, ../../CAKZiRmwxVqEbp7JgOed=BCT6cq8RNuHk3N0vuwro65Tsw9E8NA@mail.gmail.com/2-jw_no_sb_in_postmaster__noeffect.txt)
download | inline diff:
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 0399265c4dd..b80fcd58931 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -758,6 +758,8 @@ round_off_mapping_sizes_for_hugepages(MemoryMappingSizes *mapping, int hugepages
if (hugepagesize == 0)
return;
+ elog(WARNING, "shmem_reserved=%ld, shmem_req_size=%ld", mapping->shmem_reserved, mapping->shmem_req_size);
+
if (mapping->shmem_req_size % hugepagesize != 0)
mapping->shmem_req_size += add_size(mapping->shmem_req_size,
hugepagesize - (mapping->shmem_req_size % hugepagesize));
@@ -839,7 +841,7 @@ CreateAnonymousSegment(int segment_id, MemoryMappingSizes *mapping)
(errmsg("segment[%s]: could not create anonymous shared memory file: %m",
segname)));
- elog(DEBUG1, "segment[%s]: mmap(%zu)", segname, mapping->shmem_req_size);
+ elog(WARNING, "segment[%s]: mmap(%zu)", segname, mapping->shmem_req_size);
/*
* Reserve maximum required address space for future expansion of this
@@ -894,7 +896,6 @@ CreateAnonymousSegment(int segment_id, MemoryMappingSizes *mapping)
anonseg->addr = ptr;
anonseg->size = mapping->shmem_reserved;
}
-
/*
* PrepareHugePages
*
@@ -1418,6 +1419,29 @@ PGSharedMemoryDetach(void)
}
}
+void
+JWPGSharedMemoryBuffersDetachTrick(void)
+{
+ AnonShmemSegment *sbseg = &AnonShmemSegs[BUFFERS_SHMEM_SEGMENT];
+ elog(WARNING, "unmapping s_b, but not closing fd(%d), from postmaster to accelerate fork()", sbseg->fd);
+ if(munmap(sbseg->addr, sbseg->size) != 0) {
+ ereport(FATAL, (errmsg("sb segment: could not unmap anonymous shared memory: %m")));
+ }
+}
+
+void
+JWPGSharedMemoryBuffersReattachTrick(void)
+{
+ AnonShmemSegment *sbseg = &AnonShmemSegs[BUFFERS_SHMEM_SEGMENT];
+
+ if(sbseg->fd == -1)
+ elog(PANIC, "children got wrong memfd fd");
+
+ sbseg->addr = mmap(NULL, sbseg->size, PROT_READ | PROT_WRITE, MAP_SHARED, sbseg->fd, 0);
+ if (sbseg->addr == MAP_FAILED)
+ elog(PANIC, "children failed mmap: %m");
+}
+
void
ShmemControlInit(void)
{
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index 80e3088fc7e..b59c82f6c1f 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -93,6 +93,7 @@ BackgroundWriterMain(const void *startup_data, size_t startup_data_len)
WritebackContext wb_context;
Assert(startup_data_len == 0);
+ elog(WARNING, "bwriter starting");
MyBackendType = B_BG_WRITER;
AuxiliaryProcessMainCommon();
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index 85da8ac381a..43db2570a95 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -227,6 +227,9 @@ postmaster_child_launch(BackendType child_type, int child_slot,
conn_timing.fork_end = GetCurrentTimestamp();
}
+ //elog(WARNING, "starting %d", pid);
+ JWPGSharedMemoryBuffersReattachTrick();
+
/* Close the postmaster's sockets */
ClosePostmasterPorts(child_type == B_LOGGER);
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 1e92b0bcc5e..a18cb6227c5 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -257,7 +257,7 @@ CreateSharedMemoryAndSemaphores(void)
inhseg->UsedShmemSegID = i;
/* Compute the size of the shared-memory block */
- elog(DEBUG3, "invoking IpcMemoryCreate(segment %s, size=%zu, reserved address space=%zu)",
+ elog(WARNING, "invoking IpcMemoryCreate(segment %s, size=%zu, reserved address space=%zu)",
MappingName(i), mapping->shmem_req_size, mapping->shmem_reserved);
/*
@@ -289,6 +289,10 @@ CreateSharedMemoryAndSemaphores(void)
*/
if (shmem_startup_hook)
shmem_startup_hook();
+
+ /* JW HACK */
+ JWPGSharedMemoryBuffersDetachTrick();
+
}
/*
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index ac679259787..2160af8de1a 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -219,6 +219,8 @@ extern PGShmemHeader *PGSharedMemoryCreate(int segment_id, MemoryMappingSizes *m
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
+extern void JWPGSharedMemoryBuffersDetachTrick(void);
+extern void JWPGSharedMemoryBuffersReattachTrick(void);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags,
int *memfd_flags);
extern bool PGSharedMemoryResize(int segment_id, MemoryMappingSizes *mapping_sizes);
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
@ 2026-02-09 14:29 ` Andres Freund <andres@anarazel.de>
2026-02-10 12:50 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
1 sibling, 1 reply; 167+ messages in thread
From: Andres Freund @ 2026-02-09 14:29 UTC (permalink / raw)
To: Jakub Wartak <jakub.wartak@enterprisedb.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
Hi,
On 2026-02-09 14:41:12 +0100, Jakub Wartak wrote:
> I've thought that the potential main reason of the hit would be slow fork(),
> so I had an idea why we fork() with majority of memory being shared_buffers
> (BufferBlocks) that is not really used inside postmaster itself
> (I mean it does not use it, only backends do use it). I've thought it could
> be cool if we could just init the memory, leave just the fd from memfd_create
> for s_b around (that is unmap() BufferBlocks from the postmaster thus lowering
> its RSS/smaps footprint) and then on fork() the fork() would NOT have to copy
> that big kernel VMA for shared_buffers. Instead (in theory - only the fd that
> is the reference - thereby we could increase the scalability of the postmaster
> (kernel would need to perform less work during fork()). Later on, the classic
> backends on their side would mmap() the region back from the fd created earlier
> (in postmaster) using memfd_create(2), but that would happen as part of many
> backends (so workload would be spread across many CPUs).
FWIW, when looking at this in the past there were two noteworthy things:
1) The main driver of slowness was *NOT* shared buffers, but all the libraries
we link to. Particularly openssl makes things a *lot* slower, due to all
the small mappings it creates. If you compare the fork speed of a postgres
with minimal dependencies and one with all the dependencies, you'll see a
huge difference.
The reason that openssl is so bad is that it modifies data in all the
copy-on-write mappings during process exit processing. See [1].
2) A lot of the slowness isn't actually from the fork overhead itself, but
from fork competing with the processing during process exit, as both taking
conflicting locks.
I seriously doubt it's a good idea to delay the mmapping until after the fork,
that'll just lead to more different mappings to exist that then all need to be
tracked separately by the kernel.
Greetings,
Andres Freund
[1] https://postgr.es/m/hgs2vs74tzxigf2xqosez7rpf3ia5e7izalg5gz3lv3nqfptxx%40thanmprbpl4e
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-09 14:29 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
@ 2026-02-10 12:50 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
0 siblings, 0 replies; 167+ messages in thread
From: Jakub Wartak @ 2026-02-10 12:50 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com
On Mon, Feb 9, 2026 at 3:29 PM Andres Freund <andres@anarazel.de> wrote:
>
> Hi,
>
> On 2026-02-09 14:41:12 +0100, Jakub Wartak wrote:
> > I've thought that the potential main reason of the hit would be slow fork(),
> > so I had an idea why we fork() with majority of memory being shared_buffers
> > (BufferBlocks) that is not really used inside postmaster itself
> > (I mean it does not use it, only backends do use it). I've thought it could
> > be cool if we could just init the memory, leave just the fd from memfd_create
> > for s_b around (that is unmap() BufferBlocks from the postmaster thus lowering
> > its RSS/smaps footprint) and then on fork() the fork() would NOT have to copy
> > that big kernel VMA for shared_buffers. Instead (in theory - only the fd that
> > is the reference - thereby we could increase the scalability of the postmaster
> > (kernel would need to perform less work during fork()). Later on, the classic
> > backends on their side would mmap() the region back from the fd created earlier
> > (in postmaster) using memfd_create(2), but that would happen as part of many
> > backends (so workload would be spread across many CPUs).
>
> FWIW, when looking at this in the past there were two noteworthy things:
>
> 1) The main driver of slowness was *NOT* shared buffers, but all the libraries
> we link to. Particularly openssl makes things a *lot* slower, due to all
> the small mappings it creates. If you compare the fork speed of a postgres
> with minimal dependencies and one with all the dependencies, you'll see a
> huge difference.
>
> The reason that openssl is so bad is that it modifies data in all the
> copy-on-write mappings during process exit processing. See [1].
>
>
> 2) A lot of the slowness isn't actually from the fork overhead itself, but
> from fork competing with the processing during process exit, as both taking
> conflicting locks.
Interesting, thanks for sharing this. I've studied fork() itself a
little bit more
(the fork() vs various factors without crazy exit() handlers). See attached
results from 2 machines or just run fork_bench C proggie. My conclusions on
on 6.14.x are following (those are mostly notes for myself while
studying those, but I
think I'll share, maybe just one variable is missing here: how fork() ends up
being affected by NUMA - future TODO for me ):
MAP_SHARED (findings for this $thread)
--------------------------------------
a) In "mmap-MAP_SHARED" cases, the max number of fork()/s drops but very
slightly as the number of (still only MAP_SHARED!) segments increase. This
applies to both with huge pages and without them. Memfd_normal seems to
behave almost in identical way, so at least from that angle the patch seems
to be ok (assuming it has just two segments today, yesterday it had 6 for
me ;))
b) My wild trick/assumption - not related to $thread - under "memfd_unmap"
that I've posted earlier - assuming it will double postmaster scalability -
is double fizzled right now, as you say the overhead of unmapping segments
Before fork()ing and keeping just mem fd to restore that mmap MAP_SHARED
segment from child for some reason degrades performance compared to just
letting them persist or using MADV_DONTNEED. Probably it's page faulting
as you say, I haven't measured that. RIP idea.
MAP_PRIVATE (this can be ignored for the purposes of this $thread)
------------------------------------------------------------------
Nevertheless quite interesting to see how those two modes compare and it
touches aspect of openssl and e.g. io_uring using to create many VMAs too
[1]
c) MAP_PRIVATE seems to be way slower because fork() must copy PTEs and
mark them as CoW. Performance drops as the total memory (number of pages)
increases. We should not have big MAP_ANONYMOUS|MAP_PRIVATE segments (or
even just many segments [1]) in the postmaster if we want fast fork().
But even still having a lot of MAP_PRIVATE (in some edge case? large
heap?), really benefits from huge pages there.
My takeaway from this is - and it's unrelated to this $thread, but still
interesting finding for future: once we'll have multithreading, we
might be not able to fork() efficiently from there (or it will be big
huge impact for MAP_PRIVATE/big heap for all threads). It will clearly
depend on the architecture: but if postmaster will be removed and one
a giant PID will have multiple TIDs and somebody does want to run COPY
TO/FROM PROGRAM often from there, we are screwed unless those segments will
be MADV_DONOTFORK.
> I seriously doubt it's a good idea to delay the mmapping until after the fork,
> that'll just lead to more different mappings to exist that then all need to be
> tracked separately by the kernel.
Right, the raw numbers are not showing this as a good idea.
-J.
[1] - https://www.postgresql.org/message-id/7bduf2aqh6ygz7qugmb65ohczozeed36oscviebhjcvussjqt4%405fcoh7427...
Starting Sweep (HugePages: ENABLED, Max Mem: 16GB, Max Time/Test: 2.0s)
Mode | Segs | SizeGB | Forks/sec
----------------------------------------------------------------------
mmap-MAP_SHARED | 1 | 1 | 11147.96
mmap-MAP_SHARED | 2 | 1 | 11159.49
mmap-MAP_SHARED | 4 | 1 | 11352.20
mmap-MAP_SHARED | 8 | 1 | 10733.16
mmap-MAP_SHARED | 16 | 1 | 10264.63
mmap-MAP_SHARED | 1 | 2 | 11745.78
mmap-MAP_SHARED | 2 | 2 | 11873.96
mmap-MAP_SHARED | 4 | 2 | 11821.10
mmap-MAP_SHARED | 8 | 2 | 10799.39
mmap-MAP_SHARED | 1 | 4 | 11943.26
mmap-MAP_SHARED | 2 | 4 | 11842.58
mmap-MAP_SHARED | 4 | 4 | 10946.43
mmap-MAP_SHARED | 1 | 8 | 11402.92
mmap-MAP_SHARED | 2 | 8 | 11701.11
mmap-MAP_SHARED | 1 | 16 | 11924.66
----------------------------------------------------------------------
mmap-MAP_PRIVATE | 1 | 1 | 6691.29
mmap-MAP_PRIVATE | 2 | 1 | 2640.46
mmap-MAP_PRIVATE | 4 | 1 | 1094.15
mmap-MAP_PRIVATE | 8 | 1 | 583.11
mmap-MAP_PRIVATE | 16 | 1 | 315.90
mmap-MAP_PRIVATE | 1 | 2 | 2611.27
mmap-MAP_PRIVATE | 2 | 2 | 1171.92
mmap-MAP_PRIVATE | 4 | 2 | 623.90
mmap-MAP_PRIVATE | 8 | 2 | 315.12
mmap-MAP_PRIVATE | 1 | 4 | 1136.03
mmap-MAP_PRIVATE | 2 | 4 | 601.53
mmap-MAP_PRIVATE | 4 | 4 | 318.12
mmap-MAP_PRIVATE | 1 | 8 | 626.62
mmap-MAP_PRIVATE | 2 | 8 | 313.79
mmap-MAP_PRIVATE | 1 | 16 | 318.76
----------------------------------------------------------------------
memfd_normal | 1 | 1 | 12030.31
memfd_normal | 2 | 1 | 11910.85
memfd_normal | 4 | 1 | 11290.21
memfd_normal | 8 | 1 | 10934.57
memfd_normal | 16 | 1 | 10211.38
memfd_normal | 1 | 2 | 11841.26
memfd_normal | 2 | 2 | 11679.65
memfd_normal | 4 | 2 | 11603.79
memfd_normal | 8 | 2 | 11159.31
memfd_normal | 1 | 4 | 12076.96
memfd_normal | 2 | 4 | 12020.41
memfd_normal | 4 | 4 | 11157.79
memfd_normal | 1 | 8 | 11907.32
memfd_normal | 2 | 8 | 11647.34
memfd_normal | 1 | 16 | 11596.94
----------------------------------------------------------------------
memfd_MADV_DONTNEED | 1 | 1 | 11883.35
memfd_MADV_DONTNEED | 2 | 1 | 11558.44
memfd_MADV_DONTNEED | 4 | 1 | 11272.62
memfd_MADV_DONTNEED | 8 | 1 | 11406.86
memfd_MADV_DONTNEED | 16 | 1 | 10397.43
memfd_MADV_DONTNEED | 1 | 2 | 12038.87
memfd_MADV_DONTNEED | 2 | 2 | 11832.76
memfd_MADV_DONTNEED | 4 | 2 | 11595.20
memfd_MADV_DONTNEED | 8 | 2 | 11045.98
memfd_MADV_DONTNEED | 1 | 4 | 12084.83
memfd_MADV_DONTNEED | 2 | 4 | 11696.95
memfd_MADV_DONTNEED | 4 | 4 | 11359.32
memfd_MADV_DONTNEED | 1 | 8 | 12103.40
memfd_MADV_DONTNEED | 2 | 8 | 12009.92
memfd_MADV_DONTNEED | 1 | 16 | 11995.43
----------------------------------------------------------------------
memfd_unmap | 1 | 1 | 11489.67
memfd_unmap | 2 | 1 | 10878.10
memfd_unmap | 4 | 1 | 10294.98
memfd_unmap | 8 | 1 | 9788.87
memfd_unmap | 16 | 1 | 8231.27
memfd_unmap | 1 | 2 | 11091.21
memfd_unmap | 2 | 2 | 10794.88
memfd_unmap | 4 | 2 | 10565.36
memfd_unmap | 8 | 2 | 9575.96
memfd_unmap | 1 | 4 | 11341.85
memfd_unmap | 2 | 4 | 11077.75
memfd_unmap | 4 | 4 | 10298.29
memfd_unmap | 1 | 8 | 11251.97
memfd_unmap | 2 | 8 | 11038.77
memfd_unmap | 1 | 16 | 10997.57
----------------------------------------------------------------------
Starting Sweep (HugePages: ENABLED, Max Mem: 16GB, Max Time/Test: 2.0s)
Mode | Segs | SizeGB | Forks/sec
----------------------------------------------------------------------
mmap-MAP_SHARED | 1 | 1 | 11908.55
mmap-MAP_SHARED | 2 | 1 | 11832.81
mmap-MAP_SHARED | 4 | 1 | 11613.15
mmap-MAP_SHARED | 8 | 1 | 11246.94
mmap-MAP_SHARED | 16 | 1 | 10601.63
mmap-MAP_SHARED | 1 | 2 | 11903.21
mmap-MAP_SHARED | 2 | 2 | 11802.68
mmap-MAP_SHARED | 4 | 2 | 11574.65
mmap-MAP_SHARED | 8 | 2 | 11223.44
mmap-MAP_SHARED | 1 | 4 | 11907.58
mmap-MAP_SHARED | 2 | 4 | 11794.99
mmap-MAP_SHARED | 4 | 4 | 11588.97
mmap-MAP_SHARED | 1 | 8 | 11909.40
mmap-MAP_SHARED | 2 | 8 | 11791.38
mmap-MAP_SHARED | 1 | 16 | 11873.59
----------------------------------------------------------------------
mmap-MAP_PRIVATE | 1 | 1 | 4570.23
mmap-MAP_PRIVATE | 2 | 1 | 2839.87
mmap-MAP_PRIVATE | 4 | 1 | 1606.92
mmap-MAP_PRIVATE | 8 | 1 | 861.98
mmap-MAP_PRIVATE | 16 | 1 | 447.01
mmap-MAP_PRIVATE | 1 | 2 | 2823.46
mmap-MAP_PRIVATE | 2 | 2 | 1608.86
mmap-MAP_PRIVATE | 4 | 2 | 862.26
mmap-MAP_PRIVATE | 8 | 2 | 448.19
mmap-MAP_PRIVATE | 1 | 4 | 1617.32
mmap-MAP_PRIVATE | 2 | 4 | 865.24
mmap-MAP_PRIVATE | 4 | 4 | 449.12
mmap-MAP_PRIVATE | 1 | 8 | 864.51
mmap-MAP_PRIVATE | 2 | 8 | 449.74
mmap-MAP_PRIVATE | 1 | 16 | 449.38
----------------------------------------------------------------------
memfd_normal | 1 | 1 | 11923.72
memfd_normal | 2 | 1 | 11801.36
memfd_normal | 4 | 1 | 11627.44
memfd_normal | 8 | 1 | 11222.97
memfd_normal | 16 | 1 | 10520.60
memfd_normal | 1 | 2 | 11912.25
memfd_normal | 2 | 2 | 11829.10
memfd_normal | 4 | 2 | 11582.64
memfd_normal | 8 | 2 | 11195.80
memfd_normal | 1 | 4 | 11889.64
memfd_normal | 2 | 4 | 11811.02
memfd_normal | 4 | 4 | 11588.37
memfd_normal | 1 | 8 | 11864.42
memfd_normal | 2 | 8 | 11753.83
memfd_normal | 1 | 16 | 11923.17
----------------------------------------------------------------------
memfd_MADV_DONTNEED | 1 | 1 | 11948.33
memfd_MADV_DONTNEED | 2 | 1 | 11801.81
memfd_MADV_DONTNEED | 4 | 1 | 11588.46
memfd_MADV_DONTNEED | 8 | 1 | 11190.21
memfd_MADV_DONTNEED | 16 | 1 | 10510.68
memfd_MADV_DONTNEED | 1 | 2 | 11963.67
memfd_MADV_DONTNEED | 2 | 2 | 11818.99
memfd_MADV_DONTNEED | 4 | 2 | 11579.91
memfd_MADV_DONTNEED | 8 | 2 | 11237.64
memfd_MADV_DONTNEED | 1 | 4 | 11936.34
memfd_MADV_DONTNEED | 2 | 4 | 11797.59
memfd_MADV_DONTNEED | 4 | 4 | 11604.26
memfd_MADV_DONTNEED | 1 | 8 | 11905.17
memfd_MADV_DONTNEED | 2 | 8 | 11779.81
memfd_MADV_DONTNEED | 1 | 16 | 11917.30
----------------------------------------------------------------------
memfd_unmap | 1 | 1 | 11180.23
memfd_unmap | 2 | 1 | 10989.48
memfd_unmap | 4 | 1 | 10652.56
memfd_unmap | 8 | 1 | 9901.21
memfd_unmap | 16 | 1 | 8851.35
memfd_unmap | 1 | 2 | 11193.39
memfd_unmap | 2 | 2 | 10970.04
memfd_unmap | 4 | 2 | 10627.30
memfd_unmap | 8 | 2 | 9870.38
memfd_unmap | 1 | 4 | 11204.30
memfd_unmap | 2 | 4 | 10998.67
memfd_unmap | 4 | 4 | 10628.48
memfd_unmap | 1 | 8 | 11190.34
memfd_unmap | 2 | 8 | 11004.03
memfd_unmap | 1 | 16 | 11134.26
----------------------------------------------------------------------
Starting Sweep (HugePages: DISABLED, Max Mem: 16GB, Max Time/Test: 2.0s)
Mode | Segs | SizeGB | Forks/sec
----------------------------------------------------------------------
mmap-MAP_SHARED | 1 | 1 | 11567.86
mmap-MAP_SHARED | 2 | 1 | 11508.67
mmap-MAP_SHARED | 4 | 1 | 11391.57
mmap-MAP_SHARED | 8 | 1 | 11107.72
mmap-MAP_SHARED | 16 | 1 | 10611.98
mmap-MAP_SHARED | 1 | 2 | 11586.96
mmap-MAP_SHARED | 2 | 2 | 11546.36
mmap-MAP_SHARED | 4 | 2 | 11404.84
mmap-MAP_SHARED | 8 | 2 | 11115.03
mmap-MAP_SHARED | 1 | 4 | 11595.14
mmap-MAP_SHARED | 2 | 4 | 11461.75
mmap-MAP_SHARED | 4 | 4 | 11353.02
mmap-MAP_SHARED | 1 | 8 | 11582.73
mmap-MAP_SHARED | 2 | 8 | 11540.00
mmap-MAP_SHARED | 1 | 16 | 11604.37
----------------------------------------------------------------------
mmap-MAP_PRIVATE | 1 | 1 | 48.95
mmap-MAP_PRIVATE | 2 | 1 | 28.03
mmap-MAP_PRIVATE | 4 | 1 | 15.89
mmap-MAP_PRIVATE | 8 | 1 | 8.36
mmap-MAP_PRIVATE | 16 | 1 | 4.28
mmap-MAP_PRIVATE | 1 | 2 | 28.43
mmap-MAP_PRIVATE | 2 | 2 | 15.69
mmap-MAP_PRIVATE | 4 | 2 | 8.37
mmap-MAP_PRIVATE | 8 | 2 | 4.29
mmap-MAP_PRIVATE | 1 | 4 | 15.80
mmap-MAP_PRIVATE | 2 | 4 | 8.37
mmap-MAP_PRIVATE | 4 | 4 | 4.28
mmap-MAP_PRIVATE | 1 | 8 | 8.37
mmap-MAP_PRIVATE | 2 | 8 | 4.28
mmap-MAP_PRIVATE | 1 | 16 | 4.28
----------------------------------------------------------------------
memfd_normal | 1 | 1 | 11611.36
memfd_normal | 2 | 1 | 11537.67
memfd_normal | 4 | 1 | 11371.99
memfd_normal | 8 | 1 | 11109.29
memfd_normal | 16 | 1 | 10559.06
memfd_normal | 1 | 2 | 11586.08
memfd_normal | 2 | 2 | 11529.02
memfd_normal | 4 | 2 | 11387.26
memfd_normal | 8 | 2 | 11111.67
memfd_normal | 1 | 4 | 11605.29
memfd_normal | 2 | 4 | 11545.69
memfd_normal | 4 | 4 | 11391.74
memfd_normal | 1 | 8 | 11600.11
memfd_normal | 2 | 8 | 11551.25
memfd_normal | 1 | 16 | 11631.39
----------------------------------------------------------------------
memfd_MADV_DONTNEED | 1 | 1 | 11579.61
memfd_MADV_DONTNEED | 2 | 1 | 11506.17
memfd_MADV_DONTNEED | 4 | 1 | 11371.25
memfd_MADV_DONTNEED | 8 | 1 | 11072.90
memfd_MADV_DONTNEED | 16 | 1 | 10534.44
memfd_MADV_DONTNEED | 1 | 2 | 11604.92
memfd_MADV_DONTNEED | 2 | 2 | 11499.87
memfd_MADV_DONTNEED | 4 | 2 | 11335.11
memfd_MADV_DONTNEED | 8 | 2 | 11074.28
memfd_MADV_DONTNEED | 1 | 4 | 11591.28
memfd_MADV_DONTNEED | 2 | 4 | 11517.90
memfd_MADV_DONTNEED | 4 | 4 | 11359.59
memfd_MADV_DONTNEED | 1 | 8 | 11609.35
memfd_MADV_DONTNEED | 2 | 8 | 11501.70
memfd_MADV_DONTNEED | 1 | 16 | 11554.79
----------------------------------------------------------------------
memfd_unmap | 1 | 1 | 10947.47
memfd_unmap | 2 | 1 | 10753.28
memfd_unmap | 4 | 1 | 10460.15
memfd_unmap | 8 | 1 | 9794.75
memfd_unmap | 16 | 1 | 8898.24
memfd_unmap | 1 | 2 | 10941.07
memfd_unmap | 2 | 2 | 10773.05
memfd_unmap | 4 | 2 | 10450.26
memfd_unmap | 8 | 2 | 9812.40
memfd_unmap | 1 | 4 | 10900.59
memfd_unmap | 2 | 4 | 10745.32
memfd_unmap | 4 | 4 | 10483.82
memfd_unmap | 1 | 8 | 10931.35
memfd_unmap | 2 | 8 | 10743.44
memfd_unmap | 1 | 16 | 10864.12
----------------------------------------------------------------------
Starting Sweep (HugePages: DISABLED, Max Mem: 16GB, Max Time/Test: 2.0s)
Mode | Segs | SizeGB | Forks/sec
----------------------------------------------------------------------
mmap-MAP_SHARED | 1 | 1 | 12244.18
mmap-MAP_SHARED | 2 | 1 | 11631.35
mmap-MAP_SHARED | 4 | 1 | 12374.73
mmap-MAP_SHARED | 8 | 1 | 11916.43
mmap-MAP_SHARED | 16 | 1 | 10885.38
mmap-MAP_SHARED | 1 | 2 | 12701.33
mmap-MAP_SHARED | 2 | 2 | 12488.93
mmap-MAP_SHARED | 4 | 2 | 12483.24
mmap-MAP_SHARED | 8 | 2 | 11703.18
mmap-MAP_SHARED | 1 | 4 | 12659.76
mmap-MAP_SHARED | 2 | 4 | 12417.44
mmap-MAP_SHARED | 4 | 4 | 12642.99
mmap-MAP_SHARED | 1 | 8 | 12463.28
mmap-MAP_SHARED | 2 | 8 | 12486.95
mmap-MAP_SHARED | 1 | 16 | 11921.65
----------------------------------------------------------------------
mmap-MAP_PRIVATE | 1 | 1 | 26.91
mmap-MAP_PRIVATE | 2 | 1 | 13.85
mmap-MAP_PRIVATE | 4 | 1 | 8.65
mmap-MAP_PRIVATE | 8 | 1 | 6.53
mmap-MAP_PRIVATE | 16 | 1 | 4.14
mmap-MAP_PRIVATE | 1 | 2 | 13.70
mmap-MAP_PRIVATE | 2 | 2 | 8.57
mmap-MAP_PRIVATE | 4 | 2 | 6.21
mmap-MAP_PRIVATE | 8 | 2 | 4.20
mmap-MAP_PRIVATE | 1 | 4 | 8.73
mmap-MAP_PRIVATE | 2 | 4 | 6.34
mmap-MAP_PRIVATE | 4 | 4 | 4.09
mmap-MAP_PRIVATE | 1 | 8 | 6.31
mmap-MAP_PRIVATE | 2 | 8 | 4.03
mmap-MAP_PRIVATE | 1 | 16 | 4.13
----------------------------------------------------------------------
memfd_normal | 1 | 1 | 12169.04
memfd_normal | 2 | 1 | 12204.76
memfd_normal | 4 | 1 | 12002.28
memfd_normal | 8 | 1 | 11272.83
memfd_normal | 16 | 1 | 10703.16
memfd_normal | 1 | 2 | 12351.89
memfd_normal | 2 | 2 | 11459.10
memfd_normal | 4 | 2 | 12159.86
memfd_normal | 8 | 2 | 11333.70
memfd_normal | 1 | 4 | 12522.09
memfd_normal | 2 | 4 | 11972.86
memfd_normal | 4 | 4 | 11639.81
memfd_normal | 1 | 8 | 11995.48
memfd_normal | 2 | 8 | 11801.37
memfd_normal | 1 | 16 | 12156.54
----------------------------------------------------------------------
memfd_MADV_DONTNEED | 1 | 1 | 11040.77
memfd_MADV_DONTNEED | 2 | 1 | 9978.94
memfd_MADV_DONTNEED | 4 | 1 | 10240.39
memfd_MADV_DONTNEED | 8 | 1 | 10383.98
memfd_MADV_DONTNEED | 16 | 1 | 10137.84
memfd_MADV_DONTNEED | 1 | 2 | 11614.28
memfd_MADV_DONTNEED | 2 | 2 | 10961.29
memfd_MADV_DONTNEED | 4 | 2 | 10311.64
memfd_MADV_DONTNEED | 8 | 2 | 10350.50
memfd_MADV_DONTNEED | 1 | 4 | 11184.21
memfd_MADV_DONTNEED | 2 | 4 | 11274.44
memfd_MADV_DONTNEED | 4 | 4 | 10870.33
memfd_MADV_DONTNEED | 1 | 8 | 11501.86
memfd_MADV_DONTNEED | 2 | 8 | 11055.71
memfd_MADV_DONTNEED | 1 | 16 | 11550.38
----------------------------------------------------------------------
memfd_unmap | 1 | 1 | 10653.19
memfd_unmap | 2 | 1 | 10162.33
memfd_unmap | 4 | 1 | 9651.69
memfd_unmap | 8 | 1 | 8701.46
memfd_unmap | 16 | 1 | 7492.18
memfd_unmap | 1 | 2 | 10435.16
memfd_unmap | 2 | 2 | 10020.18
memfd_unmap | 4 | 2 | 10151.69
memfd_unmap | 8 | 2 | 8746.93
memfd_unmap | 1 | 4 | 10220.18
memfd_unmap | 2 | 4 | 9965.44
memfd_unmap | 4 | 4 | 9342.17
memfd_unmap | 1 | 8 | 10609.20
memfd_unmap | 2 | 8 | 10210.61
memfd_unmap | 1 | 16 | 10569.24
----------------------------------------------------------------------
Attachments:
[text/plain] laptop_hugepages.txt (4.4K, ../../CAKZiRmx-ycn+TT3_n97K40aNf4Ug0V5ywi3wu9p7fFwkWO+udg@mail.gmail.com/2-laptop_hugepages.txt)
download | inline:
Starting Sweep (HugePages: ENABLED, Max Mem: 16GB, Max Time/Test: 2.0s)
Mode | Segs | SizeGB | Forks/sec
----------------------------------------------------------------------
mmap-MAP_SHARED | 1 | 1 | 11147.96
mmap-MAP_SHARED | 2 | 1 | 11159.49
mmap-MAP_SHARED | 4 | 1 | 11352.20
mmap-MAP_SHARED | 8 | 1 | 10733.16
mmap-MAP_SHARED | 16 | 1 | 10264.63
mmap-MAP_SHARED | 1 | 2 | 11745.78
mmap-MAP_SHARED | 2 | 2 | 11873.96
mmap-MAP_SHARED | 4 | 2 | 11821.10
mmap-MAP_SHARED | 8 | 2 | 10799.39
mmap-MAP_SHARED | 1 | 4 | 11943.26
mmap-MAP_SHARED | 2 | 4 | 11842.58
mmap-MAP_SHARED | 4 | 4 | 10946.43
mmap-MAP_SHARED | 1 | 8 | 11402.92
mmap-MAP_SHARED | 2 | 8 | 11701.11
mmap-MAP_SHARED | 1 | 16 | 11924.66
----------------------------------------------------------------------
mmap-MAP_PRIVATE | 1 | 1 | 6691.29
mmap-MAP_PRIVATE | 2 | 1 | 2640.46
mmap-MAP_PRIVATE | 4 | 1 | 1094.15
mmap-MAP_PRIVATE | 8 | 1 | 583.11
mmap-MAP_PRIVATE | 16 | 1 | 315.90
mmap-MAP_PRIVATE | 1 | 2 | 2611.27
mmap-MAP_PRIVATE | 2 | 2 | 1171.92
mmap-MAP_PRIVATE | 4 | 2 | 623.90
mmap-MAP_PRIVATE | 8 | 2 | 315.12
mmap-MAP_PRIVATE | 1 | 4 | 1136.03
mmap-MAP_PRIVATE | 2 | 4 | 601.53
mmap-MAP_PRIVATE | 4 | 4 | 318.12
mmap-MAP_PRIVATE | 1 | 8 | 626.62
mmap-MAP_PRIVATE | 2 | 8 | 313.79
mmap-MAP_PRIVATE | 1 | 16 | 318.76
----------------------------------------------------------------------
memfd_normal | 1 | 1 | 12030.31
memfd_normal | 2 | 1 | 11910.85
memfd_normal | 4 | 1 | 11290.21
memfd_normal | 8 | 1 | 10934.57
memfd_normal | 16 | 1 | 10211.38
memfd_normal | 1 | 2 | 11841.26
memfd_normal | 2 | 2 | 11679.65
memfd_normal | 4 | 2 | 11603.79
memfd_normal | 8 | 2 | 11159.31
memfd_normal | 1 | 4 | 12076.96
memfd_normal | 2 | 4 | 12020.41
memfd_normal | 4 | 4 | 11157.79
memfd_normal | 1 | 8 | 11907.32
memfd_normal | 2 | 8 | 11647.34
memfd_normal | 1 | 16 | 11596.94
----------------------------------------------------------------------
memfd_MADV_DONTNEED | 1 | 1 | 11883.35
memfd_MADV_DONTNEED | 2 | 1 | 11558.44
memfd_MADV_DONTNEED | 4 | 1 | 11272.62
memfd_MADV_DONTNEED | 8 | 1 | 11406.86
memfd_MADV_DONTNEED | 16 | 1 | 10397.43
memfd_MADV_DONTNEED | 1 | 2 | 12038.87
memfd_MADV_DONTNEED | 2 | 2 | 11832.76
memfd_MADV_DONTNEED | 4 | 2 | 11595.20
memfd_MADV_DONTNEED | 8 | 2 | 11045.98
memfd_MADV_DONTNEED | 1 | 4 | 12084.83
memfd_MADV_DONTNEED | 2 | 4 | 11696.95
memfd_MADV_DONTNEED | 4 | 4 | 11359.32
memfd_MADV_DONTNEED | 1 | 8 | 12103.40
memfd_MADV_DONTNEED | 2 | 8 | 12009.92
memfd_MADV_DONTNEED | 1 | 16 | 11995.43
----------------------------------------------------------------------
memfd_unmap | 1 | 1 | 11489.67
memfd_unmap | 2 | 1 | 10878.10
memfd_unmap | 4 | 1 | 10294.98
memfd_unmap | 8 | 1 | 9788.87
memfd_unmap | 16 | 1 | 8231.27
memfd_unmap | 1 | 2 | 11091.21
memfd_unmap | 2 | 2 | 10794.88
memfd_unmap | 4 | 2 | 10565.36
memfd_unmap | 8 | 2 | 9575.96
memfd_unmap | 1 | 4 | 11341.85
memfd_unmap | 2 | 4 | 11077.75
memfd_unmap | 4 | 4 | 10298.29
memfd_unmap | 1 | 8 | 11251.97
memfd_unmap | 2 | 8 | 11038.77
memfd_unmap | 1 | 16 | 10997.57
----------------------------------------------------------------------
[text/plain] 1s4c4t__hugepages.txt (4.4K, ../../CAKZiRmx-ycn+TT3_n97K40aNf4Ug0V5ywi3wu9p7fFwkWO+udg@mail.gmail.com/3-1s4c4t__hugepages.txt)
download | inline:
Starting Sweep (HugePages: ENABLED, Max Mem: 16GB, Max Time/Test: 2.0s)
Mode | Segs | SizeGB | Forks/sec
----------------------------------------------------------------------
mmap-MAP_SHARED | 1 | 1 | 11908.55
mmap-MAP_SHARED | 2 | 1 | 11832.81
mmap-MAP_SHARED | 4 | 1 | 11613.15
mmap-MAP_SHARED | 8 | 1 | 11246.94
mmap-MAP_SHARED | 16 | 1 | 10601.63
mmap-MAP_SHARED | 1 | 2 | 11903.21
mmap-MAP_SHARED | 2 | 2 | 11802.68
mmap-MAP_SHARED | 4 | 2 | 11574.65
mmap-MAP_SHARED | 8 | 2 | 11223.44
mmap-MAP_SHARED | 1 | 4 | 11907.58
mmap-MAP_SHARED | 2 | 4 | 11794.99
mmap-MAP_SHARED | 4 | 4 | 11588.97
mmap-MAP_SHARED | 1 | 8 | 11909.40
mmap-MAP_SHARED | 2 | 8 | 11791.38
mmap-MAP_SHARED | 1 | 16 | 11873.59
----------------------------------------------------------------------
mmap-MAP_PRIVATE | 1 | 1 | 4570.23
mmap-MAP_PRIVATE | 2 | 1 | 2839.87
mmap-MAP_PRIVATE | 4 | 1 | 1606.92
mmap-MAP_PRIVATE | 8 | 1 | 861.98
mmap-MAP_PRIVATE | 16 | 1 | 447.01
mmap-MAP_PRIVATE | 1 | 2 | 2823.46
mmap-MAP_PRIVATE | 2 | 2 | 1608.86
mmap-MAP_PRIVATE | 4 | 2 | 862.26
mmap-MAP_PRIVATE | 8 | 2 | 448.19
mmap-MAP_PRIVATE | 1 | 4 | 1617.32
mmap-MAP_PRIVATE | 2 | 4 | 865.24
mmap-MAP_PRIVATE | 4 | 4 | 449.12
mmap-MAP_PRIVATE | 1 | 8 | 864.51
mmap-MAP_PRIVATE | 2 | 8 | 449.74
mmap-MAP_PRIVATE | 1 | 16 | 449.38
----------------------------------------------------------------------
memfd_normal | 1 | 1 | 11923.72
memfd_normal | 2 | 1 | 11801.36
memfd_normal | 4 | 1 | 11627.44
memfd_normal | 8 | 1 | 11222.97
memfd_normal | 16 | 1 | 10520.60
memfd_normal | 1 | 2 | 11912.25
memfd_normal | 2 | 2 | 11829.10
memfd_normal | 4 | 2 | 11582.64
memfd_normal | 8 | 2 | 11195.80
memfd_normal | 1 | 4 | 11889.64
memfd_normal | 2 | 4 | 11811.02
memfd_normal | 4 | 4 | 11588.37
memfd_normal | 1 | 8 | 11864.42
memfd_normal | 2 | 8 | 11753.83
memfd_normal | 1 | 16 | 11923.17
----------------------------------------------------------------------
memfd_MADV_DONTNEED | 1 | 1 | 11948.33
memfd_MADV_DONTNEED | 2 | 1 | 11801.81
memfd_MADV_DONTNEED | 4 | 1 | 11588.46
memfd_MADV_DONTNEED | 8 | 1 | 11190.21
memfd_MADV_DONTNEED | 16 | 1 | 10510.68
memfd_MADV_DONTNEED | 1 | 2 | 11963.67
memfd_MADV_DONTNEED | 2 | 2 | 11818.99
memfd_MADV_DONTNEED | 4 | 2 | 11579.91
memfd_MADV_DONTNEED | 8 | 2 | 11237.64
memfd_MADV_DONTNEED | 1 | 4 | 11936.34
memfd_MADV_DONTNEED | 2 | 4 | 11797.59
memfd_MADV_DONTNEED | 4 | 4 | 11604.26
memfd_MADV_DONTNEED | 1 | 8 | 11905.17
memfd_MADV_DONTNEED | 2 | 8 | 11779.81
memfd_MADV_DONTNEED | 1 | 16 | 11917.30
----------------------------------------------------------------------
memfd_unmap | 1 | 1 | 11180.23
memfd_unmap | 2 | 1 | 10989.48
memfd_unmap | 4 | 1 | 10652.56
memfd_unmap | 8 | 1 | 9901.21
memfd_unmap | 16 | 1 | 8851.35
memfd_unmap | 1 | 2 | 11193.39
memfd_unmap | 2 | 2 | 10970.04
memfd_unmap | 4 | 2 | 10627.30
memfd_unmap | 8 | 2 | 9870.38
memfd_unmap | 1 | 4 | 11204.30
memfd_unmap | 2 | 4 | 10998.67
memfd_unmap | 4 | 4 | 10628.48
memfd_unmap | 1 | 8 | 11190.34
memfd_unmap | 2 | 8 | 11004.03
memfd_unmap | 1 | 16 | 11134.26
----------------------------------------------------------------------
[text/plain] 1s4c4t__nohugepages.txt (4.4K, ../../CAKZiRmx-ycn+TT3_n97K40aNf4Ug0V5ywi3wu9p7fFwkWO+udg@mail.gmail.com/4-1s4c4t__nohugepages.txt)
download | inline:
Starting Sweep (HugePages: DISABLED, Max Mem: 16GB, Max Time/Test: 2.0s)
Mode | Segs | SizeGB | Forks/sec
----------------------------------------------------------------------
mmap-MAP_SHARED | 1 | 1 | 11567.86
mmap-MAP_SHARED | 2 | 1 | 11508.67
mmap-MAP_SHARED | 4 | 1 | 11391.57
mmap-MAP_SHARED | 8 | 1 | 11107.72
mmap-MAP_SHARED | 16 | 1 | 10611.98
mmap-MAP_SHARED | 1 | 2 | 11586.96
mmap-MAP_SHARED | 2 | 2 | 11546.36
mmap-MAP_SHARED | 4 | 2 | 11404.84
mmap-MAP_SHARED | 8 | 2 | 11115.03
mmap-MAP_SHARED | 1 | 4 | 11595.14
mmap-MAP_SHARED | 2 | 4 | 11461.75
mmap-MAP_SHARED | 4 | 4 | 11353.02
mmap-MAP_SHARED | 1 | 8 | 11582.73
mmap-MAP_SHARED | 2 | 8 | 11540.00
mmap-MAP_SHARED | 1 | 16 | 11604.37
----------------------------------------------------------------------
mmap-MAP_PRIVATE | 1 | 1 | 48.95
mmap-MAP_PRIVATE | 2 | 1 | 28.03
mmap-MAP_PRIVATE | 4 | 1 | 15.89
mmap-MAP_PRIVATE | 8 | 1 | 8.36
mmap-MAP_PRIVATE | 16 | 1 | 4.28
mmap-MAP_PRIVATE | 1 | 2 | 28.43
mmap-MAP_PRIVATE | 2 | 2 | 15.69
mmap-MAP_PRIVATE | 4 | 2 | 8.37
mmap-MAP_PRIVATE | 8 | 2 | 4.29
mmap-MAP_PRIVATE | 1 | 4 | 15.80
mmap-MAP_PRIVATE | 2 | 4 | 8.37
mmap-MAP_PRIVATE | 4 | 4 | 4.28
mmap-MAP_PRIVATE | 1 | 8 | 8.37
mmap-MAP_PRIVATE | 2 | 8 | 4.28
mmap-MAP_PRIVATE | 1 | 16 | 4.28
----------------------------------------------------------------------
memfd_normal | 1 | 1 | 11611.36
memfd_normal | 2 | 1 | 11537.67
memfd_normal | 4 | 1 | 11371.99
memfd_normal | 8 | 1 | 11109.29
memfd_normal | 16 | 1 | 10559.06
memfd_normal | 1 | 2 | 11586.08
memfd_normal | 2 | 2 | 11529.02
memfd_normal | 4 | 2 | 11387.26
memfd_normal | 8 | 2 | 11111.67
memfd_normal | 1 | 4 | 11605.29
memfd_normal | 2 | 4 | 11545.69
memfd_normal | 4 | 4 | 11391.74
memfd_normal | 1 | 8 | 11600.11
memfd_normal | 2 | 8 | 11551.25
memfd_normal | 1 | 16 | 11631.39
----------------------------------------------------------------------
memfd_MADV_DONTNEED | 1 | 1 | 11579.61
memfd_MADV_DONTNEED | 2 | 1 | 11506.17
memfd_MADV_DONTNEED | 4 | 1 | 11371.25
memfd_MADV_DONTNEED | 8 | 1 | 11072.90
memfd_MADV_DONTNEED | 16 | 1 | 10534.44
memfd_MADV_DONTNEED | 1 | 2 | 11604.92
memfd_MADV_DONTNEED | 2 | 2 | 11499.87
memfd_MADV_DONTNEED | 4 | 2 | 11335.11
memfd_MADV_DONTNEED | 8 | 2 | 11074.28
memfd_MADV_DONTNEED | 1 | 4 | 11591.28
memfd_MADV_DONTNEED | 2 | 4 | 11517.90
memfd_MADV_DONTNEED | 4 | 4 | 11359.59
memfd_MADV_DONTNEED | 1 | 8 | 11609.35
memfd_MADV_DONTNEED | 2 | 8 | 11501.70
memfd_MADV_DONTNEED | 1 | 16 | 11554.79
----------------------------------------------------------------------
memfd_unmap | 1 | 1 | 10947.47
memfd_unmap | 2 | 1 | 10753.28
memfd_unmap | 4 | 1 | 10460.15
memfd_unmap | 8 | 1 | 9794.75
memfd_unmap | 16 | 1 | 8898.24
memfd_unmap | 1 | 2 | 10941.07
memfd_unmap | 2 | 2 | 10773.05
memfd_unmap | 4 | 2 | 10450.26
memfd_unmap | 8 | 2 | 9812.40
memfd_unmap | 1 | 4 | 10900.59
memfd_unmap | 2 | 4 | 10745.32
memfd_unmap | 4 | 4 | 10483.82
memfd_unmap | 1 | 8 | 10931.35
memfd_unmap | 2 | 8 | 10743.44
memfd_unmap | 1 | 16 | 10864.12
----------------------------------------------------------------------
[text/plain] laptop_nohugepages.txt (4.4K, ../../CAKZiRmx-ycn+TT3_n97K40aNf4Ug0V5ywi3wu9p7fFwkWO+udg@mail.gmail.com/5-laptop_nohugepages.txt)
download | inline:
Starting Sweep (HugePages: DISABLED, Max Mem: 16GB, Max Time/Test: 2.0s)
Mode | Segs | SizeGB | Forks/sec
----------------------------------------------------------------------
mmap-MAP_SHARED | 1 | 1 | 12244.18
mmap-MAP_SHARED | 2 | 1 | 11631.35
mmap-MAP_SHARED | 4 | 1 | 12374.73
mmap-MAP_SHARED | 8 | 1 | 11916.43
mmap-MAP_SHARED | 16 | 1 | 10885.38
mmap-MAP_SHARED | 1 | 2 | 12701.33
mmap-MAP_SHARED | 2 | 2 | 12488.93
mmap-MAP_SHARED | 4 | 2 | 12483.24
mmap-MAP_SHARED | 8 | 2 | 11703.18
mmap-MAP_SHARED | 1 | 4 | 12659.76
mmap-MAP_SHARED | 2 | 4 | 12417.44
mmap-MAP_SHARED | 4 | 4 | 12642.99
mmap-MAP_SHARED | 1 | 8 | 12463.28
mmap-MAP_SHARED | 2 | 8 | 12486.95
mmap-MAP_SHARED | 1 | 16 | 11921.65
----------------------------------------------------------------------
mmap-MAP_PRIVATE | 1 | 1 | 26.91
mmap-MAP_PRIVATE | 2 | 1 | 13.85
mmap-MAP_PRIVATE | 4 | 1 | 8.65
mmap-MAP_PRIVATE | 8 | 1 | 6.53
mmap-MAP_PRIVATE | 16 | 1 | 4.14
mmap-MAP_PRIVATE | 1 | 2 | 13.70
mmap-MAP_PRIVATE | 2 | 2 | 8.57
mmap-MAP_PRIVATE | 4 | 2 | 6.21
mmap-MAP_PRIVATE | 8 | 2 | 4.20
mmap-MAP_PRIVATE | 1 | 4 | 8.73
mmap-MAP_PRIVATE | 2 | 4 | 6.34
mmap-MAP_PRIVATE | 4 | 4 | 4.09
mmap-MAP_PRIVATE | 1 | 8 | 6.31
mmap-MAP_PRIVATE | 2 | 8 | 4.03
mmap-MAP_PRIVATE | 1 | 16 | 4.13
----------------------------------------------------------------------
memfd_normal | 1 | 1 | 12169.04
memfd_normal | 2 | 1 | 12204.76
memfd_normal | 4 | 1 | 12002.28
memfd_normal | 8 | 1 | 11272.83
memfd_normal | 16 | 1 | 10703.16
memfd_normal | 1 | 2 | 12351.89
memfd_normal | 2 | 2 | 11459.10
memfd_normal | 4 | 2 | 12159.86
memfd_normal | 8 | 2 | 11333.70
memfd_normal | 1 | 4 | 12522.09
memfd_normal | 2 | 4 | 11972.86
memfd_normal | 4 | 4 | 11639.81
memfd_normal | 1 | 8 | 11995.48
memfd_normal | 2 | 8 | 11801.37
memfd_normal | 1 | 16 | 12156.54
----------------------------------------------------------------------
memfd_MADV_DONTNEED | 1 | 1 | 11040.77
memfd_MADV_DONTNEED | 2 | 1 | 9978.94
memfd_MADV_DONTNEED | 4 | 1 | 10240.39
memfd_MADV_DONTNEED | 8 | 1 | 10383.98
memfd_MADV_DONTNEED | 16 | 1 | 10137.84
memfd_MADV_DONTNEED | 1 | 2 | 11614.28
memfd_MADV_DONTNEED | 2 | 2 | 10961.29
memfd_MADV_DONTNEED | 4 | 2 | 10311.64
memfd_MADV_DONTNEED | 8 | 2 | 10350.50
memfd_MADV_DONTNEED | 1 | 4 | 11184.21
memfd_MADV_DONTNEED | 2 | 4 | 11274.44
memfd_MADV_DONTNEED | 4 | 4 | 10870.33
memfd_MADV_DONTNEED | 1 | 8 | 11501.86
memfd_MADV_DONTNEED | 2 | 8 | 11055.71
memfd_MADV_DONTNEED | 1 | 16 | 11550.38
----------------------------------------------------------------------
memfd_unmap | 1 | 1 | 10653.19
memfd_unmap | 2 | 1 | 10162.33
memfd_unmap | 4 | 1 | 9651.69
memfd_unmap | 8 | 1 | 8701.46
memfd_unmap | 16 | 1 | 7492.18
memfd_unmap | 1 | 2 | 10435.16
memfd_unmap | 2 | 2 | 10020.18
memfd_unmap | 4 | 2 | 10151.69
memfd_unmap | 8 | 2 | 8746.93
memfd_unmap | 1 | 4 | 10220.18
memfd_unmap | 2 | 4 | 9965.44
memfd_unmap | 4 | 4 | 9342.17
memfd_unmap | 1 | 8 | 10609.20
memfd_unmap | 2 | 8 | 10210.61
memfd_unmap | 1 | 16 | 10569.24
----------------------------------------------------------------------
[text/x-csrc] fork_bench6.c (4.7K, ../../CAKZiRmx-ycn+TT3_n97K40aNf4Ug0V5ywi3wu9p7fFwkWO+udg@mail.gmail.com/6-fork_bench6.c)
download | inline:
#define _GNU_SOURCE
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/wait.h>
#include <sys/time.h>
#include <fcntl.h>
#include <getopt.h>
#define TEST_DURATION_SEC 2.0
#define MAX_TOTAL_GB 16
typedef enum {
MODE_MMAP_SHARED,
MODE_MMAP_PRIVATE,
MODE_MEMFD_NORMAL,
MODE_MEMFD_DONTNEED,
MODE_MEMFD_UNMAP
} bench_mode_t;
const char* get_mode_name(bench_mode_t mode) {
switch(mode) {
case MODE_MMAP_SHARED: return "mmap-MAP_SHARED";
case MODE_MMAP_PRIVATE: return "mmap-MAP_PRIVATE";
case MODE_MEMFD_NORMAL: return "memfd_normal";
case MODE_MEMFD_DONTNEED: return "memfd_MADV_DONTNEED";
case MODE_MEMFD_UNMAP: return "memfd_unmap";
default: return "unknown";
}
}
double get_now() {
struct timeval tv;
gettimeofday(&tv, NULL);
return (double)tv.tv_sec + (double)tv.tv_usec / 1000000.0;
}
double run_benchmark(bench_mode_t mode, int num_segs, int size_gb, int use_huge) {
size_t segment_size = (size_t)size_gb * 1024 * 1024 * 1024ULL;
int *fds = malloc(num_segs * sizeof(int));
void **ptrs = malloc(num_segs * sizeof(void *));
for (int i = 0; i < num_segs; i++) {
int shared_flag = (mode == MODE_MMAP_PRIVATE) ? MAP_PRIVATE : MAP_SHARED;
if (mode >= MODE_MEMFD_NORMAL) {
char name[32];
sprintf(name, "bench_%d", i);
int mfd_flags = MFD_CLOEXEC | (use_huge ? MFD_HUGETLB : 0);
fds[i] = memfd_create(name, mfd_flags);
if (fds[i] < 0) return -1.0;
if (ftruncate(fds[i], segment_size) < 0) return -1.0;
ptrs[i] = mmap(NULL, segment_size, PROT_READ | PROT_WRITE, shared_flag, fds[i], 0);
} else {
int flags = shared_flag | MAP_ANONYMOUS | (use_huge ? MAP_HUGETLB : 0);
ptrs[i] = mmap(NULL, segment_size, PROT_READ | PROT_WRITE, flags, -1, 0);
}
if (ptrs[i] == MAP_FAILED) return -1.0;
memset(ptrs[i], 'A', segment_size);
if (mode == MODE_MEMFD_DONTNEED) madvise(ptrs[i], segment_size, MADV_DONTNEED);
if (mode == MODE_MEMFD_UNMAP) munmap(ptrs[i], segment_size);
}
long fork_count = 0;
double start_time = get_now();
while ((get_now() - start_time) < TEST_DURATION_SEC) {
pid_t pid = fork();
if (pid < 0) break;
if (pid == 0) {
if (mode == MODE_MEMFD_UNMAP) {
for (int j = 0; j < num_segs; j++)
mmap(NULL, segment_size, PROT_READ | PROT_WRITE, MAP_SHARED, fds[j], 0);
}
_exit(0);
} else {
waitpid(pid, NULL, 0);
fork_count++;
}
}
double total_time = get_now() - start_time;
for (int i = 0; i < num_segs; i++) {
if (mode != MODE_MEMFD_UNMAP) munmap(ptrs[i], segment_size);
if (mode >= MODE_MEMFD_NORMAL) close(fds[i]);
}
free(fds); free(ptrs);
return (double)fork_count / total_time;
}
void print_usage(char* prog) {
printf("Usage: %s [--huge-pages | --no-huge-pages]\n", prog);
exit(1);
}
int main(int argc, char **argv) {
int huge_flag = -1;
static struct option long_options[] = {
{"huge-pages", no_argument, 0, 'h'},
{"no-huge-pages", no_argument, 0, 'n'},
{0, 0, 0, 0}
};
int opt;
while ((opt = getopt_long(argc, argv, "hn", long_options, NULL)) != -1) {
switch (opt) {
case 'h': huge_flag = 1; break;
case 'n': huge_flag = 0; break;
default: print_usage(argv[0]);
}
}
if (huge_flag == -1) print_usage(argv[0]);
printf("Starting Sweep (HugePages: %s, Max Mem: %dGB, Max Time/Test: %.1fs)\n",
huge_flag ? "ENABLED" : "DISABLED", MAX_TOTAL_GB, TEST_DURATION_SEC);
printf("%-20s | %-5s | %-6s | %-12s\n", "Mode", "Segs", "SizeGB", "Forks/sec");
printf("----------------------------------------------------------------------\n");
bench_mode_t modes[] = {MODE_MMAP_SHARED, MODE_MMAP_PRIVATE, MODE_MEMFD_NORMAL, MODE_MEMFD_DONTNEED, MODE_MEMFD_UNMAP};
for (int m = 0; m < 5; m++) {
for (int s = 1; s <= 16; s *= 2) {
for (int n = 1; n <= 16; n *= 2) {
if (n * s > MAX_TOTAL_GB) continue;
double rate = run_benchmark(modes[m], n, s, huge_flag);
if (rate < 0) {
printf("%-20s | %-5d | %-6d | [ERR: OOM/POOL]\n", get_mode_name(modes[m]), n, s);
} else {
printf("%-20s | %-5d | %-6d | %-12.2f\n", get_mode_name(modes[m]), n, s, rate);
}
}
}
printf("----------------------------------------------------------------------\n");
}
return 0;
}
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
@ 2026-02-10 06:17 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
1 sibling, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2026-02-10 06:17 UTC (permalink / raw)
To: Jakub Wartak <jakub.wartak@enterprisedb.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On Mon, Feb 9, 2026 at 7:11 PM Jakub Wartak
<jakub.wartak@enterprisedb.com> wrote:
>
> On Wed, Jan 28, 2026 at 2:19 PM Ashutosh Bapat
> <ashutosh.bapat.oss@gmail.com> wrote:
>
> >v 20260128*.patch
>
> Short intro: I've started trying out these patches for slightly another reason
> than the online buffers resize. There's was recent post [1] that was brought to
> attention by Alvaro. That article is complaining about postmaster being
> unscalable and more or less saturating @ 2-3k new connections / second and
> postmaster becoming a CPU hog (one could argue that's too much and not sensible
> setup).
>
> I've thought that the potential main reason of the hit would be slow fork(),
> so I had an idea why we fork() with majority of memory being shared_buffers
> (BufferBlocks) that is not really used inside postmaster itself
> (I mean it does not use it, only backends do use it). I've thought it could
> be cool if we could just init the memory, leave just the fd from memfd_create
> for s_b around (that is unmap() BufferBlocks from the postmaster thus lowering
> its RSS/smaps footprint) and then on fork() the fork() would NOT have to copy
> that big kernel VMA for shared_buffers. Instead (in theory - only the fd that
> is the reference - thereby we could increase the scalability of the postmaster
> (kernel would need to perform less work during fork()). Later on, the classic
> backends on their side would mmap() the region back from the fd created earlier
> (in postmaster) using memfd_create(2), but that would happen as part of many
> backends (so workload would be spread across many CPUs). The critical
> assumption here is that although on Linux there seems to be huge PMD sharing for
> MAP_SHARED | MAP_HUGETLB, I was still wondering if we couldn't accelerate it
> further by simply not having at all this memory before calling fork().
> Initially I've created simple PoC bench on 64GB even with hugepages showed some
> potential:
> Scenario 1 (mmap inherited): 20001 total forks, 0.302ms per fork
> Scenario 2 (MADV_DONTFORK): 20001 total forks, 0.292ms per fork
> Scenario 3 (memfd_create): 20002 total forks, 0.145ms per fork
>
> Quite unexpectedly that's how I discovered Your's and Dimitry's patch
> as it already
> had separation of memory segments (rather than one big mmap() blob) and
> memfd_create(2) used too, so I just gave it a try. So I've tried to benchmark
> Your's patchset when it comes to establishing new connections:
>
> 1s4c 32GB RAM, 6.14.x kernel, 16GB shared_buffers
> benchmark: /usr/pgsql19/bin/pgbench -n --connect -j 4 -c 100
> -f <(echo "SELECT 1;") postgres -P 1 -T 30
>
> # master
> latency average = 358.681 ms
> latency stddev = 225.813 ms
> average connection time = 2.989 ms
> tps = 1329.733460 (including reconnection times)
>
> # memfd/thispatchset
> latency average = 363.584 ms
> latency stddev = 230.529 ms
> average connection time = 3.022 ms
> tps = 1315.810761 (including reconnection times)
>
> # memfd+mytrick, showed some promise in low stddev, but not in TPS
> latency average = 34.229 ms
> latency stddev = 22.059 ms
> average connection time = 2.908 ms
> tps = 1369.785773 (including reconnection times)
>
> Another box, 4s32c64, 128GB RAM, 6.14.x kernel,
> 64GB shared_buffers (4 NUMA nodes)
>
> benchmark: /usr/pgsql19/bin/pgbench -n --connect -j 128 -c 1000
> -f <(echo "SELECT 1;") postgres -P 1 -T 30
>
> #master
> latency average = 240.179 ms
> latency stddev = 119.379 ms
> average connection time = 62.049 ms
> tps = 2058.434343 (including reconnection times)
>
> #memfd
> latency average = 268.384 ms
> latency stddev = 133.501 ms
> average connection time = 69.081 ms
> tps = 1847.422995 (including reconnection times)
>
> #memfd+mytrick
> latency average = 261.726 ms
> latency stddev = 130.161 ms
> average connection time = 67.579 ms
> tps = 1889.988400 (including reconnection times)
>
Thanks for the benchmarks. I can see
1. There's isn't much impact of having multiple segments on new connection time.
2. fallocate seems to be behind the regression on machine with 4 NUMA nodes.
Am I reading it correctly?
The latest patches 20260209 use only two segments. Please check if
that improves the situation further.
> So:
> a) yes, my idea fizzled - still no crystal clear idea why - but at least
> I've tried Your's patch :) We are still in the ballpark of ~1800..3000
> new connections per second.
>
> and here proper review against patchset follows:
> b) the patch changes the behavior on startup and it appears that now
> the patch tries to touch all the memory during startup which takes
> much more time (I'm thinking of HA failover/promote scenarios where
> long startup on could mean trouble e.g. after pg_rewind). E.g. without
> patch it takes 1-2s and with the patch it takes 49s, no HugePages with
> 64GB s_b on slow machine). It happens due to that new fallocate() from
> shmem_fallocate(). If it is supposed to stay like that IMHO log should
> elog() what it is doing ("allocating memory...", otherwise users can
> be left confused. It almost behaves like MAP_POPULATE would be
> used.
>
> c) as per above measurements, on NUMA it appears that there's seems be
> like 1847/2058=~89% of baseline regression, when it comes to the
> establishing new connections and you are operating on sysv_shmem.c
> (so affecting all users). Possibly this would have to be re-tested
> on some more modern hardware (I don't see it on single socket, but I
> see on multiple sockets)
I have added a TODO in the code to investigate this case later as we
fine tune the code.
>
> d) MADV_HUGEPAGES is Linux 4.14+ and although released nearly 10
> years ago the buildfarm probably has some animals (Ubuntu 16?) that
> still use such
> old kernels (??))
>
> e) so maybe because of b+c+d we should consider putting it under some new
> shared_memory_type in the long run?
That may be a good idea so as to avoid hitting segfault at run time
because of lack of memory to back the shared memory.
>
> e) With huge_pages=on and no asserts it seemed to never work for me due to:
> FATAL: segment[main]: could not truncate anonymous file to
> size 313483264: Invalid argument
> and please see this (this is with both(!)
> max_shared_buffers=shared_buffers=1GB),
> for some reason ftruncate() ended up calling ~ 2x more.
> [pid 1252287] memfd_create("main", MFD_HUGETLB) = 4
> [pid 1252287] mmap(NULL, 157286400, PROT_NONE, MAP_SHARED|MAP_NORESE..
> [pid 1252287] mprotect(0x7f2a1a400000, 157286400, PROT_READ|PROT_WRI..
> [pid 1252287] ftruncate(4, 313483264) = -1 EINVAL (Invalid argument)
> it appears that I'm getting this due to bug in
> round_off_mapping_sizes_for_hugepages() as before it I'm getting:
> shmem_reserved=156196864, shmem_req_size=156196864
> and after it it's called it returning:
> shmem_reserved=157286400, shmem_req_size=313483264
> Maybe TYPE ALIGN() would be a better fit for this there.
>
I see the bug. Fixed in the attached diff. Please apply it on top of
20260209 and let me know if it fixes the issue for you. I will include
it in the next set of patches.
--
Best Wishes,
Ashutosh Bapat
Attachments:
[application/octet-stream] huge_page_fix.diff.no_ci (3.5K, ../../CAExHW5vEfDQuqgV0Z_8=5htZTt186VioD+d2YtszywegAag5=Q@mail.gmail.com/2-huge_page_fix.diff.no_ci)
download
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-10 14:37 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Jakub Wartak @ 2026-02-10 14:37 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On Tue, Feb 10, 2026 at 7:17 AM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> On Mon, Feb 9, 2026 at 7:11 PM Jakub Wartak
> <jakub.wartak@enterprisedb.com> wrote:
> > 1s4c 32GB RAM, 6.14.x kernel, 16GB shared_buffers
> > benchmark: /usr/pgsql19/bin/pgbench -n --connect -j 4 -c 100
> > -f <(echo "SELECT 1;") postgres -P 1 -T 30
> >
> > # master
> > latency average = 358.681 ms
> > latency stddev = 225.813 ms
> > average connection time = 2.989 ms
> > tps = 1329.733460 (including reconnection times)
> >
> > # memfd/thispatchset
> > latency average = 363.584 ms
> > latency stddev = 230.529 ms
> > average connection time = 3.022 ms
> > tps = 1315.810761 (including reconnection times)
> >
> > Another box, 4s32c64, 128GB RAM, 6.14.x kernel,
> > 64GB shared_buffers (4 NUMA nodes)
> >
> > benchmark: /usr/pgsql19/bin/pgbench -n --connect -j 128 -c 1000
> > -f <(echo "SELECT 1;") postgres -P 1 -T 30
> >
> > #master
> > latency average = 240.179 ms
> > latency stddev = 119.379 ms
> > average connection time = 62.049 ms
> > tps = 2058.434343 (including reconnection times)
> >
> > #memfd
> > latency average = 268.384 ms
> > latency stddev = 133.501 ms
> > average connection time = 69.081 ms
> > tps = 1847.422995 (including reconnection times)
Hi Ashutosh!
>
> Thanks for the benchmarks. I can see
> 1. There's isn't much impact of having multiple segments on new connection time.
> The latest patches 20260209 use only two segments. Please check if
> that improves the situation further.
Well there was regression 2058 -> 1847 conns/s on previous patchset on
this legacy
NUMA box, but now it's appears to be gone thanks to probably just two
regions in v20260209
as I'm getting better results:
latency average = 244.292 ms
latency stddev = 121.505 ms
average connection time = 62.973 ms
tps = 2027.553831 (including reconnection times)
That box is legacy, slow, deprecated, Andres hates it, but it's
sometimes easier to spot such
things with the naked eye.
> 2. fallocate seems to be behind the regression on machine with 4 NUMA nodes.
Yes, that fallocate() on startup can take a lot of time (without HPs
here, 32GB):
[pid 948989] 15:16:49 ftruncate(5, 34359746560) = 0
[pid 948989] 15:16:49 fallocate(5, 0, 0, 34359746560.....................) = 0
[pid 948989] 15:17:10 fallocate(6, 0, 0, 209776) = 0
I'm kind of wondering and worried about this behavior e.g. on modern >=1TB RAM
machines (so 256GB for s_b) . If you would shield that as non-default
shared_memory_type then probably not all folks will be impacted, but only
those who prefer to have online resize of buffers.
> > and here proper review against patchset follows:
> > b) the patch changes the behavior on startup and it appears that now
> > the patch tries to touch all the memory during startup which takes
> > much more time (I'm thinking of HA failover/promote scenarios where
> > long startup on could mean trouble e.g. after pg_rewind). E.g. without
> > patch it takes 1-2s and with the patch it takes 49s, no HugePages with
> > 64GB s_b on slow machine). It happens due to that new fallocate() from
> > shmem_fallocate(). If it is supposed to stay like that IMHO log should
> > elog() what it is doing ("allocating memory...", otherwise users can
> > be left confused. It almost behaves like MAP_POPULATE would be
> > used.
> >
> > c) as per above measurements, on NUMA it appears that there's seems be
> > like 1847/2058=~89% of baseline regression, when it comes to the
> > establishing new connections and you are operating on sysv_shmem.c
> > (so affecting all users). Possibly this would have to be re-tested
> > on some more modern hardware (I don't see it on single socket, but I
> > see on multiple sockets)
>
> I have added a TODO in the code to investigate this case later as we
> fine tune the code.
Well with the new patch version I think you can remove it. It think it should be
solved as per above pgbench number (with v20260209/just 2 segments)
and the numbers from fork-microbenchmark in parallel reply to Andres.
> > d) MADV_HUGEPAGES is Linux 4.14+ and although released nearly 10
> > years ago the buildfarm probably has some animals (Ubuntu 16?) that
> > still use such
> > old kernels (??))
> >
> > e) so maybe because of b+c+d we should consider putting it under some new
> > shared_memory_type in the long run?
>
> That may be a good idea so as to avoid hitting segfault at run time
> because of lack of memory to back the shared memory.
Not sure what you mean about that segfault, but right now it's just the
fallocate() that might cause a long startup time.
> > e) With huge_pages=on and no asserts it seemed to never work for me due to:
> > FATAL: segment[main]: could not truncate anonymous file to
> > size 313483264: Invalid argument
> > and please see this (this is with both(!)
> > max_shared_buffers=shared_buffers=1GB),
> > for some reason ftruncate() ended up calling ~ 2x more.
> > [pid 1252287] memfd_create("main", MFD_HUGETLB) = 4
> > [pid 1252287] mmap(NULL, 157286400, PROT_NONE, MAP_SHARED|MAP_NORESE..
> > [pid 1252287] mprotect(0x7f2a1a400000, 157286400, PROT_READ|PROT_WRI..
> > [pid 1252287] ftruncate(4, 313483264) = -1 EINVAL (Invalid argument)
> > it appears that I'm getting this due to bug in
> > round_off_mapping_sizes_for_hugepages() as before it I'm getting:
> > shmem_reserved=156196864, shmem_req_size=156196864
> > and after it it's called it returning:
> > shmem_reserved=157286400, shmem_req_size=313483264
> > Maybe TYPE ALIGN() would be a better fit for this there.
> >
>
> I see the bug. Fixed in the attached diff. Please apply it on top of
> 20260209 and let me know if it fixes the issue for you. I will include
> it in the next set of patches.
Yes, it fixes that "bug1", given
shared_buffers = '32 GB'
max_shared_buffers = '32 GB'
max_connections = 1000
huge_pages = 'on'
without it , it was:
mmap(NULL, 35399925760, PROT_NONE,
MAP_SHARED|MAP_ANONYMOUS|MAP_HUGETLB, -1, 0) = 0x7fad34c00000
mmap(NULL, 1038090240, PROT_NONE,
MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 4, 0) = 0x7facf6e00000
ftruncate(4, 2074263552) = -1 EINVAL (Invalid argument)
and with it:
mmap(NULL, 1038090240, PROT_NONE,
MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 4, 0) = 0x7fbf49a00000
ftruncate(4, 1038090240) = 0
mmap(NULL, 34361835520, PROT_NONE,
MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 5, 0) = 0x7fb749800000
ftruncate(5, 34361835520) = 0
So bug1 should be fixed. However there's something odd afterwards:
postgres=# show huge_pages;
huge_pages
------------
on
(1 row)
postgres=# show huge_pages_status ;
huge_pages_status
-------------------
off
(1 row)
Crosschecking, shows it is true, no HP ended up being allocated:
$ grep -A 2 /memfd /proc/775241/smaps # postmaster shows no HP
usage (so just 4kB pages):
7f845b7f1000-7f8c5b7f3000 rw-s 00000000 00:01 28682
/memfd:buffers (deleted)
Size: 33554440 kB
KernelPageSize: 4 kB
--
7f8c5b7f3000-7f8c9941f000 rw-s 00000000 00:01 28680
/memfd:main (deleted)
Size: 1011888 kB
KernelPageSize: 4 kB
strace shows silent failure on startup:
[pid 775320] mmap(NULL, 35399925760, PROT_NONE,
MAP_SHARED|MAP_ANONYMOUS|MAP_HUGETLB, -1, 0) = -1 ENOMEM (Cannot
allocate memory)
[pid 775320] memfd_create("main", 0) = 4
[pid 775320] mmap(NULL, 1036173312, PROT_NONE,
MAP_SHARED|MAP_NORESERVE, 4, 0) = 0x7fc3faff3000
[pid 775320] ftruncate(4, 1036173312) = 0
[pid 775320] memfd_create("buffers", 0) = 5
[pid 775320] mmap(NULL, 34359746560, PROT_NONE,
MAP_SHARED|MAP_NORESERVE, 5, 0) = 0x7fbbfaff1000
[pid 775320] ftruncate(5, 34359746560) = 0
and further tracking nailed it down to bug2 / silent failure in
CreateSharedMemoryAndSemaphores()->PrepareHugePages()
That must be some logic error there in the patch, because if I have
huge_pages=on I want it to fail to start instead of silenty fallback
to off in huge_pages_status.
And this shows another problem with calculating
shared_memory_size_in_huge_pages - it's wrong right now I think [bug3]. I have
used postgres -C shared_memory_size_in_huge_pages and have put this value in the
proper sysctl. It told me to use 16879 (which gives *2MB huge page size =
33758MBs = 35397828608 bytes, but mmap() wanted less 34359746560 and that's
~989MB difference and still it failed)
OK, sure, I'll throw some more huge pages huge pages (17879) instead but then
it still it cries with bug4 with even more strange another error:
2026-02-10 15:05:45.897 CET [775386] DEBUG: reserving space: probe
mmap(35399925760) with MAP_HUGETLB
2026-02-10 15:05:45.897 CET [775386] DEBUG: segment[main]: mmap(1038090240)
2026-02-10 15:05:46.142 CET [775386] DEBUG: segment[buffers]: mmap(34361835520)
2026-02-10 15:05:46.388 CET [775386] FATAL: segment[buffers]: could
not allocate space for anonymous file: No space left on device
so this time it was
[pid 775426] memfd_create("main", MFD_HUGETLB) = 4
[pid 775426] mmap(NULL, 1038090240, PROT_NONE,
MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 4, 0) = 0x7f8b6b200000
[pid 775426] ftruncate(4, 1038090240) = 0
[pid 775426] fallocate(4, 0, 0, 1038090240) = 0
[pid 775426] memfd_create("buffers", MFD_HUGETLB) = 5
[pid 775426] mmap(NULL, 34361835520, PROT_NONE,
MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 5, 0) = 0x7f836b000000
[pid 775426] ftruncate(5, 34361835520) = 0
[pid 775426] fallocate(5, 0, 0, 34361835520) = -1 ENOSPC (No space
left on device)
34361835520+1038090240 = 35399925760 bytes total = 16880, but I had
more than HPs free:
/sys/devices/system/node/node0/hugepages/hugepages-2048kB/free_hugepages:4750
/sys/devices/system/node/node1/hugepages/hugepages-2048kB/free_hugepages:4750
/sys/devices/system/node/node2/hugepages/hugepages-2048kB/free_hugepages:4750
/sys/devices/system/node/node3/hugepages/hugepages-2048kB/free_hugepages:4750
Run out of time to track down those HP bugs, just letting You know.
-J.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
@ 2026-02-10 15:21 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2026-02-10 15:21 UTC (permalink / raw)
To: Jakub Wartak <jakub.wartak@enterprisedb.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
Hi Jakub,
On Tue, Feb 10, 2026 at 8:07 PM Jakub Wartak
<jakub.wartak@enterprisedb.com> wrote:
>
> > I see the bug. Fixed in the attached diff. Please apply it on top of
> > 20260209 and let me know if it fixes the issue for you. I will include
> > it in the next set of patches.
>
> Yes, it fixes that "bug1", given
> shared_buffers = '32 GB'
> max_shared_buffers = '32 GB'
> max_connections = 1000
> huge_pages = 'on'
>
> without it , it was:
> mmap(NULL, 35399925760, PROT_NONE,
> MAP_SHARED|MAP_ANONYMOUS|MAP_HUGETLB, -1, 0) = 0x7fad34c00000
> mmap(NULL, 1038090240, PROT_NONE,
> MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 4, 0) = 0x7facf6e00000
> ftruncate(4, 2074263552) = -1 EINVAL (Invalid argument)
>
> and with it:
> mmap(NULL, 1038090240, PROT_NONE,
> MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 4, 0) = 0x7fbf49a00000
> ftruncate(4, 1038090240) = 0
> mmap(NULL, 34361835520, PROT_NONE,
> MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 5, 0) = 0x7fb749800000
> ftruncate(5, 34361835520) = 0
>
> So bug1 should be fixed. However there's something odd afterwards:
>
> postgres=# show huge_pages;
> huge_pages
> ------------
> on
> (1 row)
> postgres=# show huge_pages_status ;
> huge_pages_status
> -------------------
> off
> (1 row)
>
> Crosschecking, shows it is true, no HP ended up being allocated:
> $ grep -A 2 /memfd /proc/775241/smaps # postmaster shows no HP
> usage (so just 4kB pages):
> 7f845b7f1000-7f8c5b7f3000 rw-s 00000000 00:01 28682
> /memfd:buffers (deleted)
> Size: 33554440 kB
> KernelPageSize: 4 kB
> --
> 7f8c5b7f3000-7f8c9941f000 rw-s 00000000 00:01 28680
> /memfd:main (deleted)
> Size: 1011888 kB
> KernelPageSize: 4 kB
>
> strace shows silent failure on startup:
> [pid 775320] mmap(NULL, 35399925760, PROT_NONE,
> MAP_SHARED|MAP_ANONYMOUS|MAP_HUGETLB, -1, 0) = -1 ENOMEM (Cannot
> allocate memory)
> [pid 775320] memfd_create("main", 0) = 4
> [pid 775320] mmap(NULL, 1036173312, PROT_NONE,
> MAP_SHARED|MAP_NORESERVE, 4, 0) = 0x7fc3faff3000
> [pid 775320] ftruncate(4, 1036173312) = 0
> [pid 775320] memfd_create("buffers", 0) = 5
> [pid 775320] mmap(NULL, 34359746560, PROT_NONE,
> MAP_SHARED|MAP_NORESERVE, 5, 0) = 0x7fbbfaff1000
> [pid 775320] ftruncate(5, 34359746560) = 0
>
> and further tracking nailed it down to bug2 / silent failure in
> CreateSharedMemoryAndSemaphores()->PrepareHugePages()
PrepareHugePages() seems like a kludge but I haven't yet gotten time
to do something about it.
>
> That must be some logic error there in the patch, because if I have
> huge_pages=on I want it to fail to start instead of silenty fallback
> to off in huge_pages_status.
>
> And this shows another problem with calculating
> shared_memory_size_in_huge_pages - it's wrong right now I think [bug3]. I have
> used postgres -C shared_memory_size_in_huge_pages and have put this value in the
> proper sysctl. It told me to use 16879 (which gives *2MB huge page size =
> 33758MBs = 35397828608 bytes, but mmap() wanted less 34359746560 and that's
> ~989MB difference and still it failed)
>
> OK, sure, I'll throw some more huge pages huge pages (17879) instead but then
> it still it cries with bug4 with even more strange another error:
> 2026-02-10 15:05:45.897 CET [775386] DEBUG: reserving space: probe
> mmap(35399925760) with MAP_HUGETLB
> 2026-02-10 15:05:45.897 CET [775386] DEBUG: segment[main]: mmap(1038090240)
> 2026-02-10 15:05:46.142 CET [775386] DEBUG: segment[buffers]: mmap(34361835520)
> 2026-02-10 15:05:46.388 CET [775386] FATAL: segment[buffers]: could
> not allocate space for anonymous file: No space left on device
>
> so this time it was
> [pid 775426] memfd_create("main", MFD_HUGETLB) = 4
> [pid 775426] mmap(NULL, 1038090240, PROT_NONE,
> MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 4, 0) = 0x7f8b6b200000
> [pid 775426] ftruncate(4, 1038090240) = 0
> [pid 775426] fallocate(4, 0, 0, 1038090240) = 0
> [pid 775426] memfd_create("buffers", MFD_HUGETLB) = 5
> [pid 775426] mmap(NULL, 34361835520, PROT_NONE,
> MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 5, 0) = 0x7f836b000000
> [pid 775426] ftruncate(5, 34361835520) = 0
> [pid 775426] fallocate(5, 0, 0, 34361835520) = -1 ENOSPC (No space
> left on device)
>
> 34361835520+1038090240 = 35399925760 bytes total = 16880, but I had
> more than HPs free:
> /sys/devices/system/node/node0/hugepages/hugepages-2048kB/free_hugepages:4750
> /sys/devices/system/node/node1/hugepages/hugepages-2048kB/free_hugepages:4750
> /sys/devices/system/node/node2/hugepages/hugepages-2048kB/free_hugepages:4750
> /sys/devices/system/node/node3/hugepages/hugepages-2048kB/free_hugepages:4750
>
> Run out of time to track down those HP bugs, just letting You know.
Thanks a lot for all your tests. I will come around to fixing HP bugs
once I have tackled the high level items mentioned in [1] and shared
memory management rewrite that Heikki is suggesting. May I request you
to keep these tests with you till then. Or if you could investigate
these bugs and provide patches containing fixes, that will help as
well. But it may not be as easy as the bug 1 fix.
[1] https://www.postgresql.org/message-id/CAExHW5s8s=UhjqNa_Tz1PFCRLzt3=5nvd5vD1wFdKWMQCmFySQ@mail.gmail...
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-02-12 14:12 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Jakub Wartak @ 2026-02-12 14:12 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On Tue, Feb 10, 2026 at 4:21 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> Hi Jakub,
>
> On Tue, Feb 10, 2026 at 8:07 PM Jakub Wartak
> <jakub.wartak@enterprisedb.com> wrote:
> >
> > > I see the bug. Fixed in the attached diff. Please apply it on top of
> > > 20260209 and let me know if it fixes the issue for you. I will include
> > > it in the next set of patches.
> >
> > Yes, it fixes that "bug1", given
> > shared_buffers = '32 GB'
> > max_shared_buffers = '32 GB'
> > max_connections = 1000
> > huge_pages = 'on'
>
> PrepareHugePages() seems like a kludge but I haven't yet gotten time
> to do something about it.
>
> >
> > That must be some logic error there in the patch, because if I have
> > huge_pages=on I want it to fail to start instead of silenty fallback
> > to off in huge_pages_status.
>
> Thanks a lot for all your tests. I will come around to fixing HP bugs
> once I have tackled the high level items mentioned in [1] and shared
> memory management rewrite that Heikki is suggesting. May I request you
> to keep these tests with you till then. Or if you could investigate
> these bugs and provide patches containing fixes, that will help as
> well. But it may not be as easy as the bug 1 fix.
Hi Ashutosh,
OK, so with huge_page_fix.diff.no_ci with just shared_buffers=1GB,
max_shared_buffers=2GB and sysctl hugepages=634 (that's what
shared_memory_size_in_huge_page told me) was failing on fallocate for
small main ~240MB (not even the big buffers):
2026-02-12 14:08:42.924 CET [314850] DEBUG: segment[main]: mmap(241172480)
2026-02-12 14:08:42.936 CET [314850] FATAL: segment[main]: could not
allocate space for anonymous file: No space left on device
[pid 314850] mmap(NULL, 1335885824, PROT_NONE,
MAP_SHARED|MAP_ANONYMOUS|MAP_HUGETLB, -1, 0) = 0x7ce247a00000
[pid 314850] openat(AT_FDCWD, "/proc/meminfo", O_RDONLY) = 4
[pid 314850] close(4) = 0
[pid 314850] memfd_create("main", MFD_HUGETLB) = 4
[pid 314850] mmap(NULL, 241172480, PROT_NONE,
MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 4, 0) = 0x7ce239400000
[pid 314850] ftruncate(4, 241172480) = 0
[pid 314850] fallocate(4, 0, 0, 241172480) = -1 ENOSPC (No space left on device)
but before fallocate() failure we were having proper allocation of:
1335885824/1024/1024 = 1274MB
241172480/1024/1024 = 230MB
so rougly (1274+230)/2MB HPs are needed, so = 752 , if I raise it to that brings
us to that there are three (!) mmap() calls:
[pid 317348] mmap(NULL, 1335885824, PROT_NONE,
MAP_SHARED|MAP_ANONYMOUS|MAP_HUGETLB, -1, 0) = 0x722157800000
[pid 317348] openat(AT_FDCWD, "/proc/meminfo", O_RDONLY) = 4
[pid 317348] close(4) = 0
[pid 317348] memfd_create("main", MFD_HUGETLB) = 4
[pid 317348] mmap(NULL, 241172480, PROT_NONE,
MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 4, 0) = 0x722149200000
[pid 317348] ftruncate(4, 241172480) = 0
[pid 317348] fallocate(4, 0, 0, 241172480) = 0
[pid 317348] openat(AT_FDCWD, "postmaster.pid", O_RDWR) = 5
[pid 317348] close(5) = 0
[pid 317348] openat(AT_FDCWD, "/proc/meminfo", O_RDONLY) = 5
[pid 317348] close(5) = 0
[pid 317348] memfd_create("buffers", MFD_HUGETLB) = 5
[pid 317348] mmap(NULL, 2149580800, PROT_NONE,
MAP_SHARED|MAP_NORESERVE|MAP_HUGETLB, 5, 0) = 0x7220c9000000
[pid 317348] ftruncate(5, 1075838976) = 0
[pid 317348] fallocate(5, 0, 0, 1075838976) = -1 ENOSPC (No space left
on device)
So it looks like the patch requests huge-pages like that:
- it allocated 1335885824 = 1274MB with reservation from start
- it adds 230MB for memfd "main" (starts with without reservation,
lazy reserves and then full use via ftruncate -- all ok)
- then it adds not reserved 2149580800 = 2050MB (ok as it is max_shared_buffers)
and then lazily allocated 1GB, but when really trying to touch
it failed on fallocate() because pages were already consumed by
first mmap() probe call
That gives us already: 1.2 + 0.2 + 1 = 2.4GB for just shared_buffers=1GB
I don't know the rationale, but it appears that PrepareHugePages() should simply
free the huge page memory once it has validated it is there:
/* Map total amount of memory to test its availability. */
The server starts properly with HPs if I just add this munmap() there
based on that above comment (assuming PrepareHugePages() is just testing
stuff):
@@ -927,8 +913,7 @@ PrepareHugePages()
#else
if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
{
- Size hugepagesize,
- total_size = 0;
+ Size hugepagesize;
int huge_mmap_flags;
GetHugePageSize(&hugepagesize, &huge_mmap_flags, NULL);
@@ -964,6 +949,8 @@ PrepareHugePages()
SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on",
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
huge_pages_on = ptr != MAP_FAILED;
+ if(ptr != MAP_FAILED)
+ munmap(ptr, total_size);
Hope that helps a little! (I have run out of time when trying to see if and why
shared_memory_size_in_huge_page is wrongly calculated).
To sum I think those are issues as patchset stands:
- huge_page_fix.diff.no_ci (wrong calculation)
- lack of munmap() of HPs as per above
- probably logic for huge_pages=on vs huge_pages_status=off
should be reviewed (it should fallback only with try)
TBH, I haven't really looked at the code outside of that region, I'm just
trespasser that was interested in memfd ;)
-J.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
@ 2026-02-13 11:52 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2026-02-13 11:52 UTC (permalink / raw)
To: Jakub Wartak <jakub.wartak@enterprisedb.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>
On Thu, Feb 12, 2026 at 7:43 PM Jakub Wartak
<jakub.wartak@enterprisedb.com> wrote:
>
>
> TBH, I haven't really looked at the code outside of that region, I'm just
> trespasser that was interested in memfd ;)
Your trespassing has been very helpful. I have started a separate
thread to discuss resizable shared structures at [1]. Once the
implementation there is somewhat finalized, it will be good to try
your huge page tests again.
[1] https://www.postgresql.org/message-id/CAExHW5vM1bneLYfg0wGeAa=52UiJ3z4vKd3AJ72X8Fw6k3KKrg@mail.gmail...
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-07-24 12:56 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2026-07-24 12:56 UTC (permalink / raw)
To: pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi,
On Fri, Feb 13, 2026 at 5:22 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> On Thu, Feb 12, 2026 at 7:43 PM Jakub Wartak
> <jakub.wartak@enterprisedb.com> wrote:
> >
> >
> > TBH, I haven't really looked at the code outside of that region, I'm just
> > trespasser that was interested in memfd ;)
>
> Your trespassing has been very helpful. I have started a separate
> thread to discuss resizable shared structures at [1]. Once the
> implementation there is somewhat finalized, it will be good to try
> your huge page tests again.
Here's the next version of the patch implementing shared buffer pool
resizing. The patch is based on the latest master. Here's the summary
of changes since the last version:
1. The patch now uses the new shared memory infrastructure that was
introduced in PG 19. Patch 0006 enhances that infrastructure to
support resizable shared structures. I will also post the same patch
to [1]. I am fine to discuss the patch in that thread or here. The
APIs for registering and resizing the structures are documented in the
programming interface documentation.
2. Patch 0007 implements the shared buffer pool resizing using the
resizable shared structures infrastructure. It has a lot of code
improvements, including better documentation in comments, READMEs,
user-facing documentation and more TAP tests. The
storage/buffer/README has a section on buffer resizing. The buffer
resizing is implemented in buf_resize.c, which also has detailed
comments about the implementation. I suggest starting the review with
the user documentation, README and buf_resize.c.
Both patches have detailed commit messages and also list open items
for discussion and TODOs. Opinions and suggestions on those are
welcome.
The remaining patches are:
0001 - Adds sanity Asserts in the background writer code; see the
commit message. This could be committed separately.
0002 - Decouples the use of the NBuffers variable as a GUC from its
use as the size of the shared buffer pool. Prepares for the buffer
pool size to differ from the GUC value, as required by the resizing
feature.
0003 - Adds a diagnostic view to show the contents of the buffer
lookup table. Useful for debugging and testing, but not necessarily
for final commit.
0004 - Small change to pass use_units down to the per-GUC show
callbacks. Required if we accept the new output of the SHOW
shared_buffers command.
0005 - Adds backend_pid as a member of background_psql (rationale in
the commit message). Used by the buffer pool resizing tests.
Haoyu shared his version of the buffer pool resizing patch in [2].
What follows is a review of that patch. That thread is for the
resizable shared memory work, whereas these patches are for shared
buffer pool resizing. Hence, I am replying here.
The patch isn't applicable as-is since it does not use the new shared
memory infrastructure. But I have picked up some ideas and code as
explained below.
1. Buffer-manager infrastructure: a two-water-mark scheme
(lowNBuffers/highNBuffers) protected by AccessNBuffersLock, a new
Shared / local variable mapping in my patch:
- shared activeNBuffers (was lowNBuffers in your patch) <-> local activeNBuffers
- shared currentNBuffers (was highNBuffers in your patch) <-> local NBuffers
The locals are shadows of the shared values, kept in sync via the
ProcSignalBarrier mechanism; detailed comments are in the patch.
I should probably have named the shared and local versions of the
"current" value the same, but the names have stuck. Once we decide the
final names, I will change them to be consistent. I would prefer
allocNBuffers and NBuffers respectively, but I am open to suggestions.
* enable_dynamic_shared_buffers (PGC_POSTMASTER): off by default; when
off, all of the new code paths are no-ops and the server behaves as
before.
If max_shared_buffers = shared_buffers OR if the buffer pool
structures are registered as fixed-size, the feature is disabled. I
don't think an additional GUC is needed.
Patch rebased onto upstream master from the v18-based development
branch.
@@ -88,16 +89,84 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
The latest version of this function adapts the scan to the current
number of buffers. My patch does not touch this function at all.
+ BEGIN_NBUFFERS_ACCESS(localNBuffers);
This will block buffer pool resizing while a scan is in progress and
vice versa. Specifically, the changes in CheckPointBuffers() are
problematic. This function runs for minutes when there are several
buffers to sync. A resize will need to wait for that long. Instead, my
patch takes a different approach which does not require either of them
to wait too long. Please let me know what you think of that approach.
Another problem with this approach is that it adds a few steps to
acquire the lock which are unnecessary when no resize is in progress.
Since resizing is a rare event, I think we should penalize resizing
rather than normal buffer pool operations.
I think we need to fix pg_buffercache_os_pages_internal() to not rely
on NBuffers being constant throughout the duration of the function
execution. That's true for all the functions that allocate memory
before scanning the buffer pool. I have added TODOs in the current
patches to all such places. But I might have missed some. Please let
me know if you find any such places. We should first decide between
blocking buffer pool resizing while a scan is in progress, the
approach taken by my patch, or some different approach. Once we decide
that, we will fix all such places accordingly.
@@ -967,19 +976,29 @@ CheckpointerShmemRequest(void *arg)
{
The changes here seem to be specific to Neon. In my patch the requests
array is sized to the initial size of the buffer pool, capped by
MAX_CHECKPOINT_REQUESTS. Alternatively, we can size it to MaxNBuffers,
but still cap it to MAX_CHECKPOINT_REQUESTS to avoid wasting too much
memory. What do you think about that?
Size size;
+ size = offsetof(CheckpointerShmemStruct, requests);
+
/*
- * The size of the requests[] array is arbitrarily set equal to NBuffers.
- * But there is a cap of MAX_CHECKPOINT_REQUESTS to prevent accumulating
- * too many checkpoint requests in the ring buffer.
+ * The size of the requests[] array is arbitrarily set equal to the
+ * initial size of buffer pool. But there is a cap of
+ * MAX_CHECKPOINT_REQUESTS to prevent accumulating too many checkpoint
+ * requests in the ring buffer.
+ *
+ * Under dynamic_shared_buffers we use a small fixed cap instead --
+ * sizing the queue on MaxNBuffers would waste a lot of shmem under
+ * auto-scale, but a real (non-zero) queue is still required so that
+ * SYNC_UNLINK_REQUEST can be forwarded to the checkpointer for delayed
+ * unlink processing.
*/
- size = offsetof(CheckpointerShmemStruct, requests);
- size = add_size(size, mul_size(Min(NBuffers,
- MAX_CHECKPOINT_REQUESTS),
- sizeof(CheckpointerRequest)));
- ShmemRequestStruct(.name = "Checkpointer Data",
- .size = size,
- .ptr = (void **) &CheckpointerShmem,
- );
+ if (enable_dynamic_shared_buffers)
+ size = add_size(size, mul_size(DSB_CHECKPOINT_REQUESTS,
+ sizeof(CheckpointerRequest)));
+ else
+ size = add_size(size, mul_size(Min(NBuffersGUC,
+ MAX_CHECKPOINT_REQUESTS),
+ sizeof(CheckpointerRequest)));
+
+ return size;
}
diff --git a/src/backend/storage/buffer/buf_init.c
b/src/backend/storage/buffer/buf_init.c
index 1407c930c56..757c00d03d6 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -14,12 +14,18 @@
*/
#include "postgres.h"
+#include <unistd.h>
+#ifdef __linux__
+#include <sys/mman.h>
+#endif
+
+#include "miscadmin.h"
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
-#include "storage/proclist.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
-#include "storage/subsystems.h"
+#include "utils/memdebug.h"
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
@@ -69,6 +75,208 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
* multiple times. Check the PrivateRefCount infrastructure in bufmgr.c.
*/
+static Size
+buffer_pool_madvise_alignment(void)
... snip ...
+static Size
+BufferPoolArrayPhysicalExpand(void *baseptr, Size elem_size,
+ int lowNBuffers, int highNBuffers,
+ bool *success)
... snip ...
+static Size
+BufferPoolArrayPhysicalShrink(void *baseptr, Size elem_size,
+ int lowNBuffers, int highNBuffers,
+ bool *success)
... snip ...
@@ -76,26 +284,52 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
static void
BufferManagerShmemRequest(void *arg)
{
and other changes to buf_init.c
... snip ...
These changes are not required in buf_init.c. With the resizable
shared structures infrastructure, the memory allocation and
deallocation is handled in shmem.c. Please review those changes and
let me know if I am missing something.
+
+ /*
+ * Reset the clock-sweep cursor before lowering the low water mark. The
+ * existing cursor may point above new_size. Once we publish the new
+ * lowNBuffers, ClockSweepTick() may otherwise immediately wrap past
+ * the new buffers via modulo arithmetic. Resetting to 0 means the
+ * next sweep starts from the bottom of the surviving range.
+ */
+ StrategyReset(old_size, new_size);
+
+ elog(LOG, "[Shrink Barrier]: restricting allocations to %d buffers",
new_size);
+ INSTR_TIME_SET_CURRENT(phase_start);
+ /*
+ * Wait for all backends to acknowledge the new lowNBuffers. After the
+ * barrier returns, all new buffer allocations will land in [0, lowNBuffers)
+ * range. For buffers in [lowNBuffers, highNBuffers), backends can
+ * hold pins and create new pins on buffers already pinned.
+ * The EvictExtraBuffers() loop below will wait for all buffers in
+ * [lowNBuffers, highNBuffers) to be unpinned.
+ */
+ SharedBufferResizeBarrier(PROCSIGNAL_BARRIER_SHBUF_RESIZE,
CppAsString(PROCSIGNAL_BARRIER_SHBUF_RESIZE));
Two issues here. First, StrategyReset isn't bullet-proof:
nextVictimBuffer can wrap past the new size via concurrent allocations
while this function is still lowering lowNBuffers to new_size. Second,
it does nothing about the complete passes. In my patches I have
revised the reset to wrap nextVictimBuffer around and increment
complete passes when shrinking. Even so, we still need to reset the
background writer statistics, which I would like to avoid -- but I
think this is a step in the right direction.
+/*
+ * Cleanup callback. Runs from the transaction-abort PG_CATCH path *and* from
+ * before_shmem_exit() if the backend dies while holding the resize slot.
+ *
+ * Rollback policy:
+ * - Partial shrink (lowNBuffers < highNBuffers): restore lowNBuffers to
+ * highNBuffers so the buffer pool is consistent at the larger size.
+ * Memory for [lowNBuffers, highNBuffers) is still mapped, so rolling
+ * back is safe.
+ * - Partial expand: BufferManagerShmemExpand() may have populated some of
+ * [old_size, inflight_expand_target) without publishing the new water
+ * marks. This is wasteful but harmless. We surface a WARNING so operators
+ * know to retry the resize.
+ */
+static void
+ResetResizeInProgress(int code, Datum arg)
+{
... snip ...
+}
My patches currently resort to a server restart to recover from an
interrupted resize; I would like to improve on that. This function's
intentions seem to be to recover more gracefully. However, I think the
implementation is not bullet-proof. I have a TAP test which induces
all kinds of interruptions during resize. It will be good to see how
this implementation fares against that test. If you could provide an
incremental patch on top of my patches to implement a graceful
recovery, I will be happy to incorporate it into my patch after
review.
+
+ if (new_size < old_size)
+ DoShrink(rsinfo, old_size, new_size);
+ else
+ DoExpand(rsinfo, old_size, new_size);
In my earlier patches I didn't separate the C implementation function
and the actual resize function. But I see that it's useful now, and
have done it that way in the attached patches. I think I should write
separate functions for shrink and expand, as your patch does. I will
do that once the protocol is agreed upon.
I am also not sure whether it's useful to print the function progress
or statistics as the output of this function. A better approach might
be to plug into the progress tracker infrastructure. But resizing
isn't a long-running operation like vacuum or index build. So I am not
sure whether it is worth the effort.
+ static int saved_low_nbuffers = 0;
+ int current_low_nbuffers;
+
+ BEGIN_NBUFFERS_ACCESS(localNBuffers);
Locking the background writer for the duration of resizing means new
allocators may not find victims easily, and creates more work for the
checkpointer. Similarly, making resize wait for the background writer
to finish its work means resize may take longer. In my patch they do
not block each other. Like any other backend, the background writer
adapts to changes in the buffer pool through the ProcSignalBarrier
mechanism.
+
+ strategy_buf_id = StrategySyncStart(&strategy_passes, &recent_alloc,
+ ¤t_low_nbuffers);
+ if (current_low_nbuffers != saved_low_nbuffers)
+ {
+#ifdef BGW_DEBUG
+ elog(DEBUG2, "invalidated background writer state after pool resize:
%d -> %d buffers",
+ saved_low_nbuffers, current_low_nbuffers);
+#endif
+ saved_info_valid = false;
+ saved_low_nbuffers = current_low_nbuffers;
+ }
I have similar code in my patch, but I would like to avoid that as
explained above.
CheckPointBuffers(int flags)
{
- BufferSync(flags);
+ BEGIN_NBUFFERS_ACCESS(localNBuffers);
+ BufferSync(flags, localNBuffers);
+ END_NBUFFERS_ACCESS(localNBuffers);
As mentioned earlier, this could block buffer pool resizing for minutes.
@@ -5391,6 +5392,13 @@ ShowGUCOption(const struct config_generic
*record, bool use_units)
{
const struct config_int *conf = &record->_int;
+ /*
+ * Set NBuffersGUC here so that both SHOW shared_buffers (use_units==true)
+ * and pg_settings (use_units==false) reflect the current shared
buffer pool size.
+ */
+ if (conf->variable == &NBuffersGUC)
+ NBuffersGUC = GetLowNBuffers();
+
Should this be highNBuffers? That's the value that actually reflects
the current pool size -- lowNBuffers only equals it when no resize is
in progress. My patch instead prints both the current and the pending
size. Let me know what you think about that approach.
I didn't see any APIs to make the address range against the freed
shared memory inaccessible. That capability is added in the resizable
shared memory infrastructure.
The big missing piece in your patch is tests: there is nothing
exercising the resizing functionality itself or its synchronization
with other buffer pool operations. My patch has some, but I would like
to add more. Any help in that regard would be appreciated.
[1] https://www.postgresql.org/message-id/CAExHW5vM1bneLYfg0wGeAa=52UiJ3z4vKd3AJ72X8Fw6k3KKrg@mail.gmail...
[2] https://www.postgresql.org/message-id/CAM1e6U5XDwKYZo6Jj3yD3xpCB4qkhRSQn8upauHt%3DWhEbK9VZA%40mail.g...
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] v20260724-0007-Allow-to-resize-shared-buffers-without-res.patch (178.8K, ../../CAExHW5tYUYY13b4S1fv-zdxNsEEJzTy1EW6mRUj+FFEpA4Q0=g@mail.gmail.com/2-v20260724-0007-Allow-to-resize-shared-buffers-without-res.patch)
download | inline diff:
From be0fcdb2ee5eb21e63cbfe185fc9bf4c7522578d Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Sat, 6 Jun 2026 13:21:50 +0530
Subject: [PATCH v20260724 7/7] Allow to resize shared buffers without restart
User interface
==============
shared_buffers is now PGC_SIGHUP instead of PGC_POSTMASTER.
When a server is running, the new value of GUC (set using ALTER SYSTEM
... SET shared_buffers = ...; followed by SELECT pg_reload_conf()) does
not come into effect immediately. Instead a function
pg_resize_shared_buffers() is used to resize the buffer pool. The
function uses the current value of GUC in the backend where it is
executed. The function also coordinates the buffer access in other
backends while resizing the buffer pool.
SHOW shared_buffers now shows the current size of the shared buffer pool
but it also shows pending size of shared buffers, if any.
A new GUC max_shared_buffers is introduced to control the maximum value
of shared_buffers that can be set. By default it is 0. When explicitly
set, it needs to be higher than 'shared_buffers'. When max_shared_buffers
is set to 0, it assumes the same value as GUC shared_buffers. This GUC
determines the size of address space reserved for future buffer pool
sizes and the size of buffer look up table.
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking. If a buffer being evicted is pinned, we abort the
resizing operation. There are other alternative which are not
implemented in the current patches 1. to wait for the pinned buffer to
get unpinned, 2. the backend is killed or it itself cancels the query
or 3. rollback the operation which is implemented currently. Note that
option 1 and 2 would require the pinning related local and shared
records to be accessed. But we need infrastructure to do either of this
right now.
Passing current buffer pool state to a new backend
==================================================
So far the buffer pool metdata (NBuffers and the shared memory segment
address space) is saved in process local heap memory since it's static
for the life of a server. It is passed to a new backend through
Postmaster. But with buffer pool being resized while the server running,
we need Postmaster to update its buffer pool metadata as the resizing
progresses and pass it to the new backend. This has few complications:
1. Postmaster does not receive ProcSignalBarrier. So we need to signal
it separately.
2. Postmaster's local state is inherited by the new backend when
fork()ed. But we need more complex implementation to pass it to an
exec()ed backend.
3. If Postmaster can not attend to its core functionality while it is
busy responding to the resizing signal
Instead, we maintain the buffer manager state in the shared memory.
Every backend maintains its own local copy of the state. When a backend
starts, it fetches syncs its local state with the global state before it
accesses any shared buffers. During run time it updates the local state
in response to the ProcSignalBarriers. The resizing does not advance
until the recent changes to the buffer pool state have been absorbed by
all the concurrent backends. This allows the backends to continue with
their regular activity without bothering about the buffer pool state
becoming inconsistent underneath.
Discussion points
=================
Removing the evicted buffers from buffer ring
---------------------------------------------
If the buffer pool has been shrunk, the buffers in the buffer ring may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition and masks any bug that
introduces an invalid buffer into the ring. The alternative fix is more
complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure hence inaccessible when processing
buffer resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work. So defering it
to v2.
Buffer lookup table
-------------------
In order to shrink the buffer lookup table, we need to compact the hash
table directory and the hash table entries so that the free space is
moved to the end of the memory allocated to the buffer lookup table.
This requires exclusively locking the shared hash table for a longer
duration, freezing the server for that duration. Furthermore the
compaction operation itself requires significant code. Hence we setup
the buffer lookup table considering the maximum possible size of the
buffer pool which is MaxAvailableMemory only once at the beginning. It
is not resized even though buffer pool is resized. We will need separate
effort later to implement a hash table which can be resized without
locking it for a longer duration.
BgWriter reset
--------------
The background writer makes sure to free buffers ahead of the clock
hand. For this it keeps track of the clock hand and jump forward if it
fells behind the clock hand. The position of clock hand is maintained as
a couple (size of buffer pool, position of next victim buffer). When
buffer pool size changes, the position of clock hand is adjusted
according to the new size. The background's knowledge of clock hand goes
out of sync when resize happens. Hence when resize happens the
background writer resets its knowledge of clock hand and jumps to the
clock hand's position. In case the background writer was ahead of clock
hand when resize happens, it will loose that advantage causing a
momentary glitch which may not be noticeable. It may be possible to
compare background writer's past knowledge of clock hand with the
current position by adjusting the previous according to the new size of
the buffer pool and avoid reset. But it requires more investigation
into background worker and clock hand synchronization. Hence deferred it
to a future version.
Fault tolerance of pg_resize_shared_buffers()
---------------------------------------------
If the backend executing pg_resize_shared_buffers() is interrupted
because of an ERROR, query cancellation, terminate signal, timeout etc.
the server is restarted to avoid leaving the buffer pool in an
inconsistent state. The server restarts with the buffer pool setup with
the new size. If the chances of such interruption are very low, this
solution might work for the first version. However, the restart can be
avoided by in following ways:
1. roll back the resize operation, if ERRORs are recoverable
2. In case of timeout and query cancellation, leave the resizing
operation with the buffer pool in degenerate but functioning state.
Re-attempting the resize would complete the operation or roll it
back.
3. In case of non-recoverable errors or terminate signal, it's better to
restart the operation.
But this needs more thought and implementation.
GUC type of shared_buffers
--------------------------
With the current code, on the platforms where have_resizable_shmem is
OFF, the users will be able to load new value of shared_buffers but not
resize the buffer pool. Probably we should change that set the type of
GUC to be POSTMASTER on those platforms. However, that still leaves the
server running one of the platforms which usually supports resizable
shared memory structure, but can not do so at run time because the
shared memory type does not support it, we can not do the same. Will
tackle this as the patches get finalized.
Calling BufferManagerInitProc()
-------------------------------
This function gets called twice in a backend startup sequence. Once so that
BufferManagerInitalizeAccess can access the buffer pool and second time after
ProcSignalInit() for the reasons mentioned there. I think we need a better place
so that we avoid calling it twice and set it only once properly. Also it's not
clear whether the function has been called at all the right places.
I am actually not sure why don't we call ProcSignalInit() right after
InitProcess() in a backend startup sequence. If that happens, we don't need to
worry about it.
Updating local activeNBuffers when allocating new buffer
--------------------------------------------------------
Local activeNBuffers is updated in response to ProcSignalBarrier. It is also
updated by ClockSweepTick() when choosing a victim buffer for the reasons
mentioned there. The GetStrategyBuffer() function which allocates a new buffer
from the buffer ring doesn't update the local activeNBuffers. I think it should
be harmless to just rely on the ProcSignalBarrier to update the local
activeNBuffers and everyone uses the same value. But it might interfere with the
way Clock hand is wrapped.
Barrier optimization
--------------------
Each resize operation uses three barriers to synchronize the buffer pool state
across all the backends. As noted at a few places in the code, we may be able to
use lesser number of barriers to reduce waiting time. However, we need to be
sure that we are not introducing any hazard by doing so. Hence we will keep the
current implementation for now and optimize it after we have enough tests to
cover all the scenarios.
Testing
-------
We have added a new test suite to test the buffer pool resizing. There is a
stress test exercising the buffer pool resizing with pgbench. There are also
some white box tests which use injection points to trigger specific scenarios.
However, we require more tests to cover all the scenarios.
Author: Ashutosh Bapat
Initial patches by: Dmitry Dolgov <9erthalion6@gmail.com>,
Inspired by patches from: Haoyu Huang <haoyu.huang.68@gmail.com>
Author of some tests: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
Reviewed-by: Tomas Vondra
---
contrib/pg_buffercache/pg_buffercache_pages.c | 7 +
contrib/pg_prewarm/autoprewarm.c | 5 +
doc/src/sgml/config.sgml | 62 +-
doc/src/sgml/func/func-admin.sgml | 69 ++
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/postmaster/auxprocess.c | 7 +
src/backend/postmaster/postmaster.c | 5 +
src/backend/storage/buffer/Makefile | 3 +-
src/backend/storage/buffer/README | 62 +-
src/backend/storage/buffer/buf_init.c | 299 ++++++-
src/backend/storage/buffer/buf_resize.c | 523 +++++++++++
src/backend/storage/buffer/buf_table.c | 27 +-
src/backend/storage/buffer/bufmgr.c | 176 +++-
src/backend/storage/buffer/freelist.c | 150 +++-
src/backend/storage/buffer/meson.build | 1 +
src/backend/storage/ipc/procsignal.c | 10 +
src/backend/storage/lmgr/proc.c | 31 +
src/backend/tcop/postgres.c | 3 +
src/backend/utils/init/globals.c | 2 +
src/backend/utils/init/postinit.c | 77 +-
src/backend/utils/misc/guc.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 15 +-
src/include/catalog/pg_proc.dat | 15 +
src/include/miscadmin.h | 3 +
src/include/storage/buf_internals.h | 48 +-
src/include/storage/bufmgr.h | 15 +
src/include/storage/procsignal.h | 5 +
src/include/utils/guc.h | 2 +
src/include/utils/guc_hooks.h | 2 +
src/test/Makefile | 3 +-
src/test/README | 3 +
src/test/buffermgr/Makefile | 35 +
src/test/buffermgr/README | 26 +
src/test/buffermgr/buffermgr_test.conf | 11 +
src/test/buffermgr/expected/buffer_resize.out | 290 ++++++
src/test/buffermgr/meson.build | 25 +
src/test/buffermgr/sql/buffer_resize.sql | 106 +++
src/test/buffermgr/t/001_resize_buffer.pl | 193 ++++
.../buffermgr/t/003_resize_fault_tolerance.pl | 839 ++++++++++++++++++
.../t/004_client_join_buffer_resize.pl | 221 +++++
src/test/buffermgr/t/005_resize_failures.pl | 171 ++++
.../buffermgr/t/006_resize_with_syslogger.pl | 58 ++
src/test/meson.build | 1 +
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 76 ++
src/tools/pgindent/typedefs.list | 1 +
45 files changed, 3564 insertions(+), 123 deletions(-)
create mode 100644 src/backend/storage/buffer/buf_resize.c
create mode 100644 src/test/buffermgr/Makefile
create mode 100644 src/test/buffermgr/README
create mode 100644 src/test/buffermgr/buffermgr_test.conf
create mode 100644 src/test/buffermgr/expected/buffer_resize.out
create mode 100644 src/test/buffermgr/meson.build
create mode 100644 src/test/buffermgr/sql/buffer_resize.sql
create mode 100644 src/test/buffermgr/t/001_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/003_resize_fault_tolerance.pl
create mode 100644 src/test/buffermgr/t/004_client_join_buffer_resize.pl
create mode 100644 src/test/buffermgr/t/005_resize_failures.pl
create mode 100644 src/test/buffermgr/t/006_resize_with_syslogger.pl
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 9512f1efa2f..312343fd7bf 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -289,6 +289,13 @@ pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
HeapTuple tuple;
Datum result;
+ /*
+ * TODO: This allocates memory using NBuffers which may change while this
+ * function is executed. We need to change this function so that it
+ * doesn't rely on NBuffers being static throughout the execution of this
+ * function.
+ */
+
if (SRF_IS_FIRSTCALL())
{
int i,
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index deb4c2671b5..e12323fb7c6 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -702,6 +702,11 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
return 0;
}
+ /*
+ * TODO: we need to modify this function to not rely on NBuffers being
+ * constant.
+ */
+
/*
* With sufficiently large shared_buffers, allocation will exceed 1GB, so
* allow for a huge allocation to prevent outright failure.
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 5b6cc870093..7be9378bde8 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -1802,7 +1802,6 @@ include_dir 'conf.d'
that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
(Non-default values of <symbol>BLCKSZ</symbol> change the minimum
value.)
- This parameter can only be set at server start.
</para>
<para>
@@ -1825,6 +1824,67 @@ include_dir 'conf.d'
appropriate, so as to leave adequate space for the operating system.
</para>
+ <para>
+ The shared memory consumed by the buffer pool is allocated and
+ initialized according to the value of the GUC at the time of starting
+ the server. A desired new value of GUC can be loaded while the server is
+ running using <systemitem>SIGHUP</systemitem>. But the buffer pool will
+ not be resized immediately. Use
+ <function>pg_resize_shared_buffers()</function> to dynamically resize
+ the shared buffer pool (see <xref linkend="functions-admin"/> for
+ details). If the running server has
+ <varname>have_resizable_shmem</varname> set to OFF,
+ <function>pg_resize_shared_buffers()</function> throws an error since
+ resizing a shared memory structure is not supported on that server. In
+ such a case the new value of <varname>shared_buffers</varname> gets
+ loaded using <function>pg_reload_conf</function> but the new size of the
+ pool takes effect only after restarting the server. <command>SHOW
+ shared_buffers</command> shows the currently effective value and any
+ pending value of the GUC. Please note that when the GUC is changed, the
+ other GUCS which use this GUCs value to set their defaults will not be
+ changed. They may still require a server restart to consider new value.
+ </para>
+
+ <para>
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-max-shared-buffers" xreflabel="max_shared_buffers">
+ <term><varname>max_shared_buffers</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>max_shared_buffers</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Sets the upper limit for the <varname>shared_buffers</varname> value.
+ The default value is <literal>0</literal>,
+ which means no explicit limit is set and <varname>max_shared_buffers</varname>
+ will be automatically set to the value of <varname>shared_buffers</varname>
+ at server startup.
+ If this value is specified without units, it is taken as blocks,
+ that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
+ This parameter can only be set at server start.
+ </para>
+
+ <para>
+ This parameter determines the amount of memory address space to reserve
+ in each backend for expanding the buffer pool in future. While the
+ memory for buffer pool is allocated on demand as it is resized, the
+ memory required for the buffer lookup table and the array used to sort
+ buffers during a checkpoint is allocated at the server start
+ considering the largest buffer pool size allowed by this parameter.
+ <!-- TODO: Provide a numeric example of how much extra memory say max_shared_buffers = 1GB consume. -->
+ </para>
+
+ <para>
+ When <varname>have_resizable_shmem</varname> is OFF, this parameter does
+ not have any effect except limiting the maximum value of
+ <varname>shared_buffers</varname>. It does not decide the address space
+ to be reserved or the size of the buffer lookup table and the array
+ used to sort buffers during a checkpoint.
+ </para>
</listitem>
</varlistentry>
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 0eae1c1f616..2a621fff060 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -99,6 +99,75 @@
<returnvalue>off</returnvalue>
</para></entry>
</row>
+
+ <row>
+ <entry role="func_table_entry"><para role="func_signature">
+ <indexterm>
+ <primary>pg_resize_shared_buffers</primary>
+ </indexterm>
+ <function>pg_resize_shared_buffers</function> ()
+ <returnvalue>boolean</returnvalue>
+ </para>
+ <para>
+ Dynamically resizes the shared buffer pool to match the current value of
+ the <varname>shared_buffers</varname> parameter in the client backend
+ where it is run. This function implements a coordinated resize process
+ that ensures all backend processes continue to operate without causing
+ any hazard. The resize happens in multiple phases to maintain data
+ consistency and system stability. Returns <literal>true</literal> if the
+ resize was successful, otherwise <literal>false</literal>. In the latter
+ case, the buffer pool is left at its previous size; this can happen if a
+ buffer being evicted during shrink could not be released, or if the
+ operating system could not supply enough shared memory to expand the
+ pool. Consult the server log for the underlying reason. This function
+ can only be called by superusers.
+ </para>
+ <para>
+ To resize shared buffers, first update the <varname>shared_buffers</varname>
+ setting and reload the configuration, then verify the new value is loaded
+ before calling this function. For example:
+<programlisting>
+postgres=# ALTER SYSTEM SET shared_buffers = '256MB'; -- Step 1
+ALTER SYSTEM
+postgres=# SELECT pg_reload_conf(); -- Step 2
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers; -- Step 3
+ shared_buffers
+-------------------------
+ 128MB (pending: 256MB)
+(1 row)
+
+postgres=# SELECT pg_resize_shared_buffers(); -- Step 4
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers; -- Step 5 (verification)
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+</programlisting>
+ The <command>SHOW shared_buffers</command> at Step 3 is important to
+ verify that the configuration reload was successful and the new value is
+ available to the current session before attempting the resize. The
+ output shows both the current and pending values when the GUC change is
+ pending to be applied.
+ </para>
+ <para>
+ <!-- TODO: Document behaviour when the function is called on platforms that do not support resizable shared memory -->
+ linkend="functions-admin-signal-table"/> send control signals to
+ other server processes. Use of these functions is restricted to
+ superusers by default but access may be granted to others using
+ <command>GRANT</command>, with noted exceptions.
+ </para>
+ </entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index a678f345230..fa9a01a7ec3 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -376,6 +376,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ InitializeMaxNBuffers();
+
ShmemCallRequestCallbacks();
CreateSharedMemoryAndSemaphores();
diff --git a/src/backend/postmaster/auxprocess.c b/src/backend/postmaster/auxprocess.c
index 07a3b5c5923..a09d376d333 100644
--- a/src/backend/postmaster/auxprocess.c
+++ b/src/backend/postmaster/auxprocess.c
@@ -19,6 +19,7 @@
#include "miscadmin.h"
#include "pgstat.h"
#include "postmaster/auxprocess.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/proc.h"
@@ -106,6 +107,12 @@ AuxiliaryProcessMainCommon(void)
*/
ShmemReprotectResizableStructs();
+ /*
+ * Update buffer manager's local state, which might have been changed by
+ * an online resize after startup.
+ */
+ BufferManagerInitProc();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 90c7c4528e8..9a9e60e5fb1 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -959,6 +959,11 @@ PostmasterMain(int argc, char *argv[])
*/
InitializeFastPathLocks();
+ /*
+ * Calculate MaxNBuffers after NBuffersGUC has been set.
+ */
+ InitializeMaxNBuffers();
+
/*
* Also call any legacy shmem request hooks that might've been installed
* by preloaded libraries.
diff --git a/src/backend/storage/buffer/Makefile b/src/backend/storage/buffer/Makefile
index fd7c40dcb08..3bc9aee85de 100644
--- a/src/backend/storage/buffer/Makefile
+++ b/src/backend/storage/buffer/Makefile
@@ -17,6 +17,7 @@ OBJS = \
buf_table.o \
bufmgr.o \
freelist.o \
- localbuf.o
+ localbuf.o \
+ buf_resize.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index b332e002ba1..9ee05d562eb 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -181,8 +181,11 @@ buffer header spinlock, which would have to be taken anyway to increment the
buffer reference count, so it's nearly free.)
The "clock hand" is a buffer index, nextVictimBuffer, that moves circularly
-through all the available buffers. nextVictimBuffer is protected by the
-buffer_strategy_lock.
+through all the available buffers. Usually a victim can be chosen from the whole
+buffer pool, except when resizing the buffer pool. During resizing the victim
+can be chosen from a range of buffer pool which will not be affected by the
+resizing. See "Resizing shared buffers section" below. nextVictimBuffer is
+protected by the buffer_strategy_lock.
The algorithm for a process that needs to obtain a victim buffer is:
@@ -275,3 +278,58 @@ As of 8.4, background writer starts during recovery mode when there is
some form of potentially extended recovery to perform. It performs an
identical service to normal processing, except that checkpoints it
writes are technically restartpoints.
+
+Resizing shared buffers at runtime
+----------------------------------
+
+Before PostgreSQL 20, the size of the shared buffer pool (i.e. the number of
+shared buffers) was given by the global variable NBuffers and was fixed at
+server start time using GUC 'shared_buffers'. In order to change the size of the
+shared buffer pool, one needed to change the GUC and restart the server.
+Starting PostgreSQL 20, PostgreSQL supports resizing the buffer pool without a
+server restart. The value of the GUC 'shared_buffers' is stored in a new GUC
+variable called NBuffersGUC. The old variable NBuffers now strictly reflects the
+number of buffers in the buffer pool. The new GUC variable max_shared_buffers
+defines the maximum size of the shared buffer pool. See configure.sgml for more
+details about these GUCs.
+
+Resizing buffer pool involves resizing the data structures used by the buffer
+manager and coordinating the resize with all the backends.
+
+The buffer manager maintains following data structures in shared memory.
+1. Buffer Descriptors: An array of BufferDesc structures, one per buffer.
+2. Buffer Blocks: An array of buffer blocks, forming the buffer pool.
+3. Buffer Lookup Table: A hash table mapping a page to the buffer containing
+ that page.
+4. IO conditional variables: An array of conditional variables, one per buffer.
+5. Checkpoint buffer ids: An array of buffer ids used during checkpointing.
+6. Buffer Control: A structure containing information about the size of the
+ buffer pool and whether it is being resized.
+
+Except for the Buffer Lookup Table, Checkpoint buffer ids and Buffer Control
+all other structures are registered as resizable structures. We use
+ShmemResizeStruct() described in doc/src/sgml/xfunc.sgml to resize them. The
+hash table contains multiple shared structures, each of which needs to be
+resized separately. Since these substructures are not registered as separate
+structures, they can not be turned into resizable structures and hence the hash
+table is not registered as a resizable structure. Instead we set it up for
+maximal buffer pool at the server startup. The Checkpoint buffer ids array is
+filled in by the checkpointer with buffers to be written out during a
+checkpoint; if it were resized alongside the buffer pool, a shrink concurrent
+with an ongoing checkpoint could discard entries the checkpointer still needs
+to process. To avoid that, it is also sized for the maximal buffer pool at
+server startup.
+
+For ease of resizing, we differentiate between the number of buffers in the
+whole buffer pool (BufferControl::currentNBuffers) and the number of buffers at
+the start of the buffer pool from which a victim can be chosen for replacement
+(BufferControl::activeNBuffers). Each backend maintains a local copy of these
+two numbers. These copies provide a local view of the buffer pool that the
+backend can rely upon without worrying about the concurrent changes happening to
+the buffer pool. These copies are updated through a barrier mechanism when a
+resize is performed.
+
+To resize shared buffers at runtime a user performs the steps mentioned in the
+description of shared_buffers GUC variable in configure.sgml. Actual resizing
+protocol is documented in the prologue of function pg_resize_shared_buffers() in
+src/backend/storage/buffer/buf_resize.c
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 9ddf6551fcd..91671bce888 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,10 +17,15 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proclist.h"
#include "storage/shmem.h"
#include "storage/subsystems.h"
+#include "utils/guc.h"
+#include "utils/guc_hooks.h"
+#include "utils/injection_point.h"
+BufferControlBlock *BufferControl;
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
ConditionVariableMinimallyPadded *BufferIOCVArray;
@@ -37,6 +42,31 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
.attach_fn = BufferManagerShmemAttach,
};
+/*
+ * Resizable shared memory structures backing the buffer pool.
+ */
+static const struct
+{
+ const char *name;
+ size_t element_size;
+ size_t alignment;
+ void **ptr;
+} BufferManagerResizableStructs[] = {
+
+ {"Buffer Descriptors",
+ sizeof(BufferDescPadded),
+ PG_CACHE_LINE_SIZE,
+ (void **) &BufferDescriptors},
+ {"Buffer Blocks",
+ BLCKSZ,
+ PG_IO_ALIGN_SIZE,
+ (void **) &BufferBlocks},
+ {"Buffer IO Condition Variables",
+ sizeof(ConditionVariableMinimallyPadded),
+ PG_CACHE_LINE_SIZE,
+ (void **) &BufferIOCVArray},
+};
+
/*
* Data Structures:
* buffers live in a freelist and a lookup data structure.
@@ -69,6 +99,27 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
* multiple times. Check the PrivateRefCount infrastructure in bufmgr.c.
*/
+/*
+ * Initialize a single buffer.
+ */
+static void
+InitializeBuffer(int buf_id)
+{
+ /*
+ * Do not use GetBufferDescriptor here since it relies on the buffer
+ * descriptor being initialized.
+ */
+ BufferDesc *buf = &(BufferDescriptors[buf_id]).bufferdesc;
+
+ ClearBufferTag(&buf->tag);
+ pg_atomic_init_u64(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ buf->buf_id = buf_id;
+ pgaio_wref_clear(&buf->io_wref);
+ proclist_init(&buf->lock_waiters);
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+}
+
/*
* Register shared memory area for the buffer pool.
@@ -76,25 +127,21 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
static void
BufferManagerShmemRequest(void *arg)
{
- ShmemRequestStruct(.name = "Buffer Descriptors",
- .size = NBuffersGUC * sizeof(BufferDescPadded),
- /* Align descriptors to a cacheline boundary. */
- .alignment = PG_CACHE_LINE_SIZE,
- .ptr = (void **) &BufferDescriptors,
- );
-
- ShmemRequestStruct(.name = "Buffer Blocks",
- .size = NBuffersGUC * (Size) BLCKSZ,
- /* Align buffer pool on IO page size boundary. */
- .alignment = PG_IO_ALIGN_SIZE,
- .ptr = (void **) &BufferBlocks,
- );
+ /*
+ * Fall back to fixed sized shared buffer pool if resizable shared memory
+ * is not supported on this platform.
+ */
+#ifdef HAVE_RESIZABLE_SHMEM
+ bool resizable = (shared_memory_type == SHMEM_TYPE_MMAP);
+#else
+ bool resizable = false;
+#endif
+ int min_nbuffers = resizable ? MIN_NUM_BUFFERS : 0;
+ int max_nbuffers = resizable ? MaxNBuffers : 0;
- ShmemRequestStruct(.name = "Buffer IO Condition Variables",
- .size = NBuffersGUC * sizeof(ConditionVariableMinimallyPadded),
- /* Align descriptors to a cacheline boundary. */
- .alignment = PG_CACHE_LINE_SIZE,
- .ptr = (void **) &BufferIOCVArray,
+ ShmemRequestStruct(.name = "Buffer Control",
+ .size = sizeof(BufferControl),
+ .ptr = (void **) &BufferControl,
);
/*
@@ -103,11 +150,29 @@ BufferManagerShmemRequest(void *arg)
* memory at runtime. As that'd be in the middle of a checkpoint, or when
* the checkpointer is restarted, memory allocation failures would be
* painful.
+ *
+ * When the buffer pool is resizable, it is sized for MaxNBuffers up front
+ * so that the entries filled in by the checkpointer are not freed even
+ * when the buffer pool is shrunk.
*/
ShmemRequestStruct(.name = "Checkpoint BufferIds",
- .size = NBuffersGUC * sizeof(CkptSortItem),
+ .size = (size_t) (resizable ? MaxNBuffers : NBuffersGUC) * sizeof(CkptSortItem),
+ .alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &CkptBufferIds,
);
+
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
+ {
+ size_t elem_size = BufferManagerResizableStructs[i].element_size;
+
+ ShmemRequestStruct(.name = BufferManagerResizableStructs[i].name,
+ .minimum_size = min_nbuffers * elem_size,
+ .size = NBuffersGUC * elem_size,
+ .maximum_size = max_nbuffers * elem_size,
+ .alignment = BufferManagerResizableStructs[i].alignment,
+ .ptr = BufferManagerResizableStructs[i].ptr,
+ );
+ }
}
/*
@@ -129,34 +194,194 @@ BufferManagerShmemInit(void *arg)
* Initialize all the buffer headers.
*/
for (int i = 0; i < NBuffers; i++)
+ InitializeBuffer(i);
+
+ /* Initialize BufferControl */
+ pg_atomic_init_u32(&BufferControl->currentNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->activeNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->targetNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->resizer_pid, 0);
+
+ /* Need to perform per backend steps in this backend too. */
+ BufferManagerShmemAttach(arg);
+}
+
+static void
+BufferManagerShmemAttach(void *arg)
+{
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+
+ BufferManagerInitProc();
+}
+
+/*
+ * Fetch latest buffer pool sizes (NBuffers and activeNBuffers) shared state
+ * into process local globals.
+ */
+void
+BufferManagerInitProc(void)
+{
+ NBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ elog(DEBUG1, "setting process local buffer pool sizes: currentNBuffers = %d, activeNBuffers = %d", NBuffers, activeNBuffers);
+}
+
+/*
+ * Protect unused shared memory reserved address space.
+ *
+ * Protect the parts of the shared memory address space reserved by the buffer
+ * manager which are not used by current structures from being accessed by
+ * backends.
+ *
+ * Unused address spaces of all resizable shared structures, including the
+ * buffer manager ones, are protected at the server startup together using
+ * ShmemProtectResizableStructs(). We need this function only during resizing of
+ * the buffer pool when we specifically adjust protections of buffer manager
+ * structures.
+ */
+void
+BufferManagerShmemProtect(void)
+{
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
{
- BufferDesc *buf = GetBufferDescriptor(i);
+ if (i == 2)
+ INJECTION_POINT("buffer-mgr-protect-struct", NULL);
+ ShmemProtectStruct(BufferManagerResizableStructs[i].name);
+ }
+}
+
+/*
+ * Resize and reinitialize shared buffer manager structures when resizing the
+ * buffer pool.
+ *
+ * Returns true if all the structures were resized successfully, false
+ * otherwise. We expect shrink to always succeed, but expansion may fail if the
+ * system is out of memory.
+ *
+ * The caller will always see all the structures resized consistently. If
+ * expanding a structure fails, all the expanded structures are shrunk back to
+ * their original sizes and no barrier is sent to the other backends.
+ */
+bool
+BufferManagerShmemResize(int currentNBuffers, int targetNBuffers)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizing shared buffer pool is not supported on this platform"));
+ pg_unreachable();
+#else
+ int resized = 0;
- ClearBufferTag(&buf->tag);
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
- pg_atomic_init_u64(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
+ {
+ const char *name = BufferManagerResizableStructs[i].name;
+ size_t elem_size = BufferManagerResizableStructs[i].element_size;
+ bool resize_ok = true;
+
+#ifdef USE_INJECTION_POINTS
+ if (i == 2)
+ {
+ /* Injection point to simulate an interruption in this function. */
+ INJECTION_POINT("buffer-mgr-resize-struct", NULL);
- buf->buf_id = i;
+ /*
+ * Injection point to simulate a failure in resizing a structure
+ * like memory allocation failure without actually running out of
+ * memory.
+ */
+ if (IS_INJECTION_POINT_ATTACHED("buffer-mgr-resize-struct-fail"))
+ resize_ok = false;
+ }
+#endif
- pgaio_wref_clear(&buf->io_wref);
+ if (resize_ok)
+ resize_ok = ShmemResizeStruct(name, (size_t) targetNBuffers * elem_size);
- proclist_init(&buf->lock_waiters);
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ if (!resize_ok)
+ {
+ Assert(targetNBuffers > currentNBuffers);
+ for (int j = 0; j < resized; j++)
+ ShmemResizeStruct(BufferManagerResizableStructs[j].name,
+ (size_t) currentNBuffers * BufferManagerResizableStructs[j].element_size);
+ return false;
+ }
+ resized++;
}
- /* Initialize per-backend file flush context */
- WritebackContextInit(&BackendWritebackContext,
- &backend_flush_after);
+ /* Initialize the headers for new buffers. */
+ for (int i = currentNBuffers; i < targetNBuffers; i++)
+ InitializeBuffer(i);
+
+ return true;
+#endif
}
-static void
-BufferManagerShmemAttach(void *arg)
+/*
+ * check_shared_buffers
+ * GUC check_hook for shared_buffers
+ *
+ * When reloading the configuration, shared_buffers should not be set to a value
+ * higher than max_shared_buffers fixed at the boot time.
+ */
+bool
+check_shared_buffers(int *newval, void **extra, GucSource source)
{
- /* Update the size of the buffer pool. */
- NBuffers = NBuffersGUC;
+ if (finalMaxNBuffers && *newval > MaxNBuffers)
+ {
+ GUC_check_errdetail("\"shared_buffers\" must be less than \"max_shared_buffers\".");
+ return false;
+ }
+ return true;
+}
- /* Initialize per-backend file flush context */
- WritebackContextInit(&BackendWritebackContext,
- &backend_flush_after);
+/*
+ * show_shared_buffers
+ * GUC show_hook for shared_buffers
+ *
+ * Shows both current and pending buffer counts with proper unit formatting.
+ */
+const char *
+show_shared_buffers(bool use_units)
+{
+ static char buffer[128];
+ int64 current_value;
+ const char *current_unit;
+ int currentNBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+
+ if (use_units)
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ else
+ {
+ current_unit = "";
+ current_value = currentNBuffers;
+ }
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s", current_value, current_unit);
+
+ if (currentNBuffers != NBuffersGUC)
+ {
+ int64 pending_value;
+ const char *pending_unit;
+
+ /*
+ * Shared buffer pool is pending to be resized, show both current and
+ * pending sizes.
+ */
+ if (use_units)
+ convert_int_from_base_unit(NBuffersGUC, GUC_UNIT_BLOCKS, &pending_value, &pending_unit);
+ else
+ {
+ pending_value = NBuffersGUC;
+ pending_unit = "";
+ }
+ snprintf(buffer + strlen(buffer), sizeof(buffer) - strlen(buffer), " (pending: " INT64_FORMAT "%s)",
+ pending_value, pending_unit);
+ }
+
+ return buffer;
}
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
new file mode 100644
index 00000000000..c758493f353
--- /dev/null
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -0,0 +1,523 @@
+/*-------------------------------------------------------------------------
+ *
+ * buf_resize.c
+ * shared buffer pool resizing functionality
+ *
+ * This module contains the implementation of shared buffer pool resizing,
+ * including the main resize coordination function and barrier processing
+ * functions that synchronize all backends during resize operations.
+ *
+ * Portions Copyright (c) 1996-2026, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/storage/buffer/buf_resize.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "access/htup_details.h"
+#include "fmgr.h"
+#include "funcapi.h"
+#include "miscadmin.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
+#include "utils/builtins.h"
+#include "utils/injection_point.h"
+
+static volatile sig_atomic_t safe_exit = true;
+
+static bool resize_shared_buffers_internal(void);
+static void buf_resize_shmem_exit(int code, Datum arg);
+
+/*
+ * Set the new buffer allocation pool size, broadcast it to all the backends
+ * and wait for them to acknowledge it.
+ */
+static void
+buf_resize_set_new_alloc_size(int alloc_size)
+{
+ uint64 generation;
+
+ pg_atomic_write_u32(&BufferControl->activeNBuffers, alloc_size);
+ StrategyAdjustNewBufAllocSize();
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC);
+ INJECTION_POINT("pgrsb-new-buffer-alloc-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC barrier");
+}
+
+/*
+ * Update the buffer pool size, broadcast it to all the backends and wait for
+ * them to acknowledge the change.
+ */
+static void
+buf_resize_set_buffer_pool_size(int new_size)
+{
+ uint64 generation;
+
+ pg_atomic_write_u32(&BufferControl->currentNBuffers, new_size);
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE);
+ INJECTION_POINT("pgrsb-buffer-pool-size-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE barrier");
+}
+
+/*
+ * Resize the shared buffer manager structures, broadcast the change to all
+ * the backends and wait for them to acknowledge it.
+ *
+ * If memory is not available when expanding the buffer pool, this function
+ * returns false without sending the barrier. When shrinking the buffer pool, we
+ * don't expect any failure, so this function always returns true.
+ */
+static bool
+buf_resize_shmem_resize(int currentNBuffers, int targetNBuffers)
+{
+ uint64 generation;
+
+ if (!BufferManagerShmemResize(currentNBuffers, targetNBuffers))
+ {
+ Assert(targetNBuffers > currentNBuffers);
+ return false;
+ }
+
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE);
+ INJECTION_POINT("pgrsb-buffer-pool-resize-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE barrier");
+ return true;
+}
+
+/*
+ * C implementation of SQL interface to update the shared buffers according to
+ * the current values of shared_buffers GUC.
+ *
+ * Atomic BufferControl::resizer_pid holds the PID of the backend currently
+ * performing a resize, or 0 when no resize is in progress. Using
+ * compare-and-exchange to set and reset this field, we make sure that only one
+ * resize is in progress at a time.
+ *
+ * Shrinking the buffer pool involves the following steps:
+ * - s1: Set BufferControl::activeNBuffers to the new size of the buffer pool
+ * and send SHBUF_NEW_BUFFER_ALLOC barrier to all backends. Every backend is
+ * expected to update their local buffer allocation pool size and acknowledge
+ * the barrier.
+ * - s2: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, new buffer allocations will be restricted
+ * to the new size of the buffer pool.
+ * - s3: Evict the buffers beyond the new size. A backend which still requires
+ * a previously allocated buffer which is being evicted, must have pinned it.
+ * If a pinned buffer is encountered, the resize operation is rolled back and
+ * the function returns false.
+ * - s4: If eviction succeeds, no backend should be using the buffers beyond
+ * the new size of the buffer pool and no new buffers can be allocated in
+ * that range. Update BufferControl::currentNBuffers to the new size of the
+ * buffer pool and send SHBUF_BUFFER_POOL_SIZE barrier to all backends. In
+ * response, all the backends should update their local buffer pool size and
+ * acknowledge the barrier.
+ * - s5: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, no backend will be accessing the shared
+ * buffer manager structures beyond the new size of the buffer pool.
+ * - s6: Resize the shared buffer manager structures to the new size (using
+ * ShmemResizeStruct()) and send SHBUF_BUFFER_POOL_RESIZE barrier to all
+ * backends. In response, all the backends should call ShmemProtectStruct()
+ * to update the memory address space protection of the shared buffer
+ * manager structures.
+ * - s7: Wait for all backends to acknowledge the barrier, before returning
+ * true to indicate successful resizing.
+ *
+ * Expanding the buffer pool involves the following steps:
+ * - e1: Resize the shared buffer manager structures to the new size (using
+ * ShmemResizeStruct()) and send SHBUF_BUFFER_POOL_RESIZE barrier to all the
+ * backends. In response, all the backends should call ShmemProtectStruct()
+ * to update the memory address space protection of the shared buffer
+ * manager structures. If expanding the shared buffer manager structures
+ * fails because of lack of memory, the function returns false without
+ * sending the barrier.
+ * - e2: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, every backend will be able to access the
+ * shared buffer manager structures beyond the old size of the buffer pool.
+ * - e3: Update BufferControl::currentNBuffers to the new size of the buffer
+ * pool and send SHBUF_BUFFER_POOL_SIZE barrier to all backends. In response,
+ * all the backends should update their local buffer pool size and
+ * acknowledge the barrier.
+ * - e4: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, every backend is setup to use the new size
+ * of the buffer pool.
+ * - e5: Update BufferControl::activeNBuffers to the new size of the buffer
+ * pool, so that backends can start allocating from the new area of the
+ * buffer pool. Send SHBUF_NEW_BUFFER_ALLOC barrier to all backends. In
+ * response, all the backends should update their local buffer allocation
+ * pool size and acknowledge the barrier.
+ * - e6: Wait for all backends to acknowledge the barrier, before returning
+ * true to indicate successful resizing.
+ *
+ * Reason we introduce e4:
+ * Once we expand the new buffer allocation area, all the backends will start
+ * allocating buffers from the new area. Since this happens asynchronously,
+ * there is a chance that some backends may see buffers from outside their
+ * known buffer pool size. To avoid that, first set the new buffer pool size,
+ * broadcast it to all the backends and wait for them to update their
+ * knowledge of buffer pool size. We may be able to avoid sending the barrier
+ * after the first step or avoid them altogether but it's not clear that that
+ * is completely hazard free. It feels safer this way, even though it takes
+ * longer.
+ *
+ * If a timeout happens or request to cancel query arrives while the function is
+ * being executed, we need to abort the operation immediately. The state of the
+ * buffer pool and its state as viewed by the backends may not be consistent at
+ * that point. Hence we escalate it to PANIC to restart the server and avoid
+ * inconsistent state. We may improve this situation by leaving the buffer pool
+ * in a consistent but degenerate state and allowing a subsequent resize
+ * operation to rollback or continue the operation.
+ *
+ * If an ERROR is raised while the function is being executed, we may have
+ * already entered an inconsistent state. Hence we escalate it to PANIC to
+ * restart the server and avoid inconsistent state.
+ */
+Datum
+pg_resize_shared_buffers(PG_FUNCTION_ARGS)
+{
+ bool success = false;
+
+ /*
+ * Register the exit hook before claiming resizer_pid, so that if we exit
+ * after claiming resizer_pid, the hook is in place to reset it.
+ */
+ before_shmem_exit(buf_resize_shmem_exit, 0);
+
+ PG_TRY();
+ {
+ uint32 expected_pid = 0;
+
+ if (!pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, MyProcPid))
+ {
+ /*
+ * Another backend holds resizer_pid; expected_pid was updated by
+ * the CAS to reflect its PID.
+ */
+ elog(LOG, "shared buffer resize already in progress in backend %u",
+ expected_pid);
+ /* No shared memory was touched, so it should be safe to exit. */
+ Assert(safe_exit);
+ }
+ else
+ {
+ INJECTION_POINT("pg-resize-shared-buffers-flag-set", NULL);
+
+ /*
+ * We are about to make changes to the shared memory which can not
+ * be rolled back easily since we need all the backends to
+ * acknowledge these changes. Indicate that a sudden exit in this
+ * state can leave the server in an inconsistent state.
+ */
+ safe_exit = false;
+ success = resize_shared_buffers_internal();
+
+ /*
+ * The changes to shared memory are in a consistent state across
+ * all the backends, so it should be safe to exit.
+ */
+ safe_exit = true;
+ }
+ }
+ PG_FINALLY();
+ {
+ uint32 expected_pid = MyProcPid;
+
+ /*
+ * We are in the middle of resizing and caught an error. Without
+ * knowing the reason and exact state of resizing it's not safe to
+ * continue or to exit. Restarting the server is the safest option
+ * here. Emit the error to the server log and raise PANIC to restart
+ * the server.
+ */
+ if (!safe_exit)
+ {
+ HOLD_INTERRUPTS();
+ errcontext("during shared buffer resize");
+ EmitErrorReport();
+ ereport(PANIC,
+ errmsg("shared buffer resize caught an error when shared memory was in an inconsistent state"));
+ pg_unreachable();
+ }
+
+ /*
+ * Reset the PID, if we set it before removing the shmem_exit hook so
+ * as not to leave it set after the backend has exited.
+ */
+ (void) pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, 0);
+ cancel_before_shmem_exit(buf_resize_shmem_exit, 0);
+ }
+ PG_END_TRY();
+
+ if (success)
+ elog(LOG, "shared buffer resizing to %d buffers completed successfully", NBuffersGUC);
+ else
+ elog(WARNING, "shared buffer resizing to %d buffers failed", NBuffersGUC);
+
+ PG_RETURN_BOOL(success);
+}
+
+/*
+ * Workhorse function for the C implementation.
+ */
+static bool
+resize_shared_buffers_internal(void)
+{
+ int currentNBuffers;
+ int targetNBuffers;
+ bool resize_success;
+
+ currentNBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+ targetNBuffers = NBuffersGUC;
+ if (currentNBuffers == targetNBuffers)
+ {
+ elog(LOG, "shared buffers are already at %d, no need to resize", currentNBuffers);
+ return true;
+ }
+
+ /*
+ * TODO: What if the NBuffersGUC value seen here is not the desired one
+ * because somebody did a pg_reload_conf() between the last
+ * pg_reload_conf() and execution of this function?
+ */
+
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, targetNBuffers);
+ elog(LOG, "resizing shared buffers from %d to %d", currentNBuffers, targetNBuffers);
+
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * step s1, s2: Restrict new buffer allocations to the new buffer pool
+ * size.
+ *
+ * TODO: Alternate design idea by Andres (as I understand it): Set
+ * BufferControl::activeNBuffers and send the barrier. Instead of
+ * waiting for barrier, start evicting buffers but don't unpin the
+ * evicted buffers so that they will not be considered for new
+ * allocations. Once all the buffers are evicted wait for the barrier
+ * to be acknowledged. This will reduce the time taken to shrink the
+ * buffer pool.
+ */
+ elog(LOG, "shrinking buffer pool, restricting allocations to %d buffers", targetNBuffers);
+ buf_resize_set_new_alloc_size(targetNBuffers);
+
+ /* Step s3: Evict buffers in the area being shrunk */
+ elog(LOG, "evicting buffers %u..%u", targetNBuffers + 1, currentNBuffers);
+ if (!EvictExtraBuffers(targetNBuffers, currentNBuffers))
+ {
+ elog(WARNING, "failed to evict extra buffers during shrinking");
+
+ /* Eviction failed, rollback the buffer resize operation. */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, currentNBuffers);
+ buf_resize_set_new_alloc_size(currentNBuffers);
+ return false;
+ }
+
+ /* Step s4, s5: Update the buffer pool size. */
+ buf_resize_set_buffer_pool_size(targetNBuffers);
+ }
+
+ /* Step s6, s7 or e1, e2: Resize the buffer manager structures. */
+ resize_success = buf_resize_shmem_resize(currentNBuffers, targetNBuffers);
+
+ if (targetNBuffers > currentNBuffers)
+ {
+ if (!resize_success)
+ {
+ elog(WARNING, "failed to expand buffer pool structures");
+
+ /* Revert any changes to the shared memory in this function. */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, currentNBuffers);
+ return false;
+ }
+
+ /* Step e3, e4: Declare new buffer pool size. */
+ buf_resize_set_buffer_pool_size(targetNBuffers);
+
+ /* Step e5, e6: Let expanded buffer pool be used by all backends. */
+ buf_resize_set_new_alloc_size(targetNBuffers);
+ }
+
+ return true;
+}
+
+/*
+ * Function to handle process exit when buffer resizing is in progress.
+ */
+static void
+buf_resize_shmem_exit(int code, Datum arg)
+{
+ uint32 expected_pid;
+
+ /*
+ * If resizer_pid does not match our PID, either we never claimed it or we
+ * have already released it. Nothing to do.
+ */
+ if (pg_atomic_read_u32(&BufferControl->resizer_pid) != MyProcPid)
+ return;
+
+ /*
+ * Resize is in progress and the process crashed. We do not know exactly
+ * at which step of the resizing we are. Just restart the server to be
+ * safe.
+ *
+ * TODO: If we can perform heavy operations in this callback like waiting
+ * for barriers, we could set the current status of resizing in the
+ * process local memory and use this callback to rollback every operation
+ * that was performed, except buffer eviction.
+ */
+ if (!safe_exit)
+ ereport(PANIC,
+ errmsg("buffer resize operation interrupted, restarting to avoid inconsistent state"));
+
+ /*
+ * safe_exit should be set to true when new allocations are not using the
+ * whole buffer pool, so the following condition should never happen. But
+ * be on the safer side.
+ */
+ if (pg_atomic_read_u32(&BufferControl->currentNBuffers) != pg_atomic_read_u32(&BufferControl->activeNBuffers))
+ ereport(PANIC,
+ (errmsg("buffer resize operation interrupted at an unexpected stage, restarting to avoid inconsistent state")));
+
+ /*
+ * Reset targetNBuffers before releasing resizer_pid, so that a backend
+ * claiming ownership immediately afterwards does not have its own
+ * targetNBuffers overwritten by us.
+ */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, pg_atomic_read_u32(&BufferControl->currentNBuffers));
+
+ expected_pid = MyProcPid;
+ (void) pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, 0);
+}
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC.
+ */
+bool
+ProcessBarrierNewBufferAlloc(void)
+{
+ elog(DEBUG2, "processing barrier to restrict new buffer allocations to %d buffers (target = %d)",
+ pg_atomic_read_u32(&BufferControl->activeNBuffers), pg_atomic_read_u32(&BufferControl->targetNBuffers));
+
+ INJECTION_POINT("pgrsb-handle-new-buffer-alloc-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(NBuffers == pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ return true;
+}
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE.
+ */
+bool
+ProcessBarrierBufferPoolResize(void)
+{
+ elog(DEBUG2, "processing barrier to propagate resized shared buffer pool structures");
+
+ INJECTION_POINT("pgrsb-handle-buffer-pool-resize-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(NBuffers == pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ Assert(activeNBuffers == pg_atomic_read_u32(&BufferControl->activeNBuffers));
+
+ /*
+ * Access permissions to address range covered by a resizable structure is
+ * maintained consistently across all the backends right from the time a
+ * backend is started. We maintain that consistency as the buffer pool is
+ * resized.So ideally modifying the access permissions in this backend
+ * should not fail. But if it does, the address space accessible to this
+ * backend may be inconsistent with the new buffer pool size and also with
+ * the other backends. This may cause data corruption and other memory
+ * access issues, if we let this backend continue to run and access the
+ * buffer pool. Better to quit from the faulty backend.
+ */
+ PG_TRY();
+ {
+ BufferManagerShmemProtect();
+ }
+ PG_CATCH();
+ {
+ /*
+ * We don't know what caused the error and so avoid using further
+ * resources. Emit the original error to the server log so that it's
+ * not lost and raise a FATAL to terminate this backend.
+ */
+ HOLD_INTERRUPTS();
+ errcontext("during shared buffer pool resize barrier");
+ EmitErrorReport();
+ ereport(FATAL,
+ (errmsg("shared buffer pool resize barrier caught an error while updating buffer pool protection")));
+ pg_unreachable();
+ }
+ PG_END_TRY();
+
+ return true;
+}
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE.
+ */
+bool
+ProcessBarrierBufferPoolSize(void)
+{
+ elog(DEBUG2, "processing barrier to establish new size of the buffer pool to %d", pg_atomic_read_u32(&BufferControl->currentNBuffers));
+
+ INJECTION_POINT("pgrsb-handle-buffer-pool-size-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(activeNBuffers == pg_atomic_read_u32(&BufferControl->activeNBuffers));
+ NBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+
+ return true;
+}
+
+/*
+ * SQL-callable function reporting the current shared buffer pool resize
+ * status.
+ */
+Datum
+pg_get_buffer_resize_status(PG_FUNCTION_ARGS)
+{
+#define PG_GET_BUFFER_RESIZE_STATUS_COLS 4
+ TupleDesc tupdesc;
+ Datum values[PG_GET_BUFFER_RESIZE_STATUS_COLS];
+ bool nulls[PG_GET_BUFFER_RESIZE_STATUS_COLS] = {0};
+ HeapTuple tuple;
+
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+ tupdesc = BlessTupleDesc(tupdesc);
+
+ values[0] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->activeNBuffers));
+ values[1] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ values[2] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->targetNBuffers));
+ values[3] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->resizer_pid));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ PG_RETURN_DATUM(HeapTupleGetDatum(tuple));
+#undef PG_GET_BUFFER_RESIZE_STATUS_COLS
+}
+
+/*
+ * TODO: add progress report facility if required.
+ */
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index c82b71deaa6..85664990b3c 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -28,6 +28,7 @@
#include "utils/builtins.h"
#include "storage/lwlock.h"
#include "storage/subsystems.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -58,15 +59,25 @@ BufTableShmemRequest(void *arg)
* Request the shared buffer lookup hashtable.
*
* Since we can't tolerate running out of lookup table entries, we must be
- * sure to specify an adequate table size here. The maximum steady-state
- * usage is of course as many entries as the number of buffers in the
- * pool, but BufferAlloc() tries to insert a new entry before deleting the
- * old. In principle this could be happening in each partition
- * concurrently, so we could need as many as (number of buffers in the
- * pool) + NUM_BUFFER_PARTITIONS entries. Since we are still requesting
- * shared memory, use the GUC value instead of the actual size.
+ * sure to specify an adequate table size here. The maximum number of
+ * entries we could need is the number of buffers, but BufferAlloc() tries
+ * to insert a new entry before deleting the old. In principle this could
+ * be happening in each partition concurrently, so we could need as many
+ * as (number of buffers) + NUM_BUFFER_PARTITIONS entries. The number of
+ * buffers can be increased upto MaxNBuffers at run time. So we need to
+ * make sure that the table can accommodate MaxNBuffers +
+ * NUM_BUFFER_PARTITIONS entries.
+ *
+ * See storage/buffer/README for reasons why we don't register the hash
+ * table as a resizable structure.
+ *
*/
- size = NBuffersGUC + NUM_BUFFER_PARTITIONS;
+#ifdef HAVE_RESIZABLE_SHMEM
+ size = (shared_memory_type == SHMEM_TYPE_MMAP) ? MaxNBuffers : NBuffersGUC;
+#else
+ size = NBuffersGUC;
+#endif
+ size = size + NUM_BUFFER_PARTITIONS;
ShmemRequestHash(.name = "Shared Buffer Lookup Table",
.nelems = size,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index db7f34d2274..ff45a171cfb 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -223,7 +223,10 @@ int io_max_combine_limit = DEFAULT_IO_COMBINE_LIMIT;
int checkpoint_flush_after = DEFAULT_CHECKPOINT_FLUSH_AFTER;
int bgwriter_flush_after = DEFAULT_BGWRITER_FLUSH_AFTER;
int backend_flush_after = DEFAULT_BACKEND_FLUSH_AFTER;
+
int NBuffers = 0; /* number of buffers in the buffer pool */
+int activeNBuffers = 0; /* number of buffers at the start of the pool
+ * from which new buffer allocations happnen. */
/* local state for LockBufferForCleanup */
static BufferDesc *PinCountWaitBuf = NULL;
@@ -814,6 +817,15 @@ PrefetchBuffer(Relation reln, ForkNumber forkNum, BlockNumber blockNum)
* Compared to ReadBuffer(), this avoids a buffer mapping lookup when it's
* successful. Return true if the buffer is valid and still has the expected
* tag. In that case, the buffer is pinned and the usage count is bumped.
+ *
+ * The callers of this function should make sure that the buffer is valid even
+ * if the shared buffer pool has undergone a resize. Buffer pool resizing waits
+ * for all backends to acknowledge the barrier before changing the buffer pool
+ * size. Hence the caller should call this function after validation without an
+ * intervening ProcSignalBarrier processing.
+ *
+ * TODO: This function could perform the validation in this function itself
+ * instead of relying on the two callers who do it currently.
*/
bool
ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockNum,
@@ -3655,6 +3667,11 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_START(NBuffers, num_to_scan);
+ /*
+ * TODO: Test the case when buffer pool is shrunk after CkptBufferIds is
+ * filled and num_to_scan is higher than the new NBuffers?
+ */
+
/*
* Sort buffers that need to be written to reduce the likelihood of random
* IO. The sorting is also important for the implementation of balancing
@@ -3768,6 +3785,21 @@ BufferSync(int flags)
buf_id = CkptBufferIds[ts_stat->index].buf_id;
Assert(buf_id != -1);
+ /*
+ * TODO: We need to test the scenario when the buffer pool is shrunk
+ * after checkpointer has collected the buffer ids and one or more of
+ * the buffer ids is out of range.
+ */
+
+ /*
+ * The buffer pool might have been shrunk between the time the
+ * checkpoint collected the buffer ids and now. Ignore any buffers
+ * that are out of range now. Those buffers must have been written
+ * when they were evicted during resizing.
+ */
+ if (buf_id >= NBuffers)
+ continue;
+
bufHdr = GetBufferDescriptor(buf_id);
num_processed++;
@@ -3843,6 +3875,13 @@ BufferSync(int flags)
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
+ * Usually new buffers are allocated from the whole buffer pool. However, when
+ * resizing the buffer pool, the new buffer allocations are restricted to some
+ * initial portion of the buffer pool. The rest of pool is either being evicted
+ * when shrinking or not utilized yet when growing, hence not interesting to the
+ * background writer. Hence the background writer always restricts its activity
+ * to the same portion of the buffer pool as new buffer allocations.
+ *
* This is called periodically by the background writer process.
*
* Returns true if it's appropriate for the bgwriter process to go into
@@ -3894,6 +3933,7 @@ BgBufferSync(WritebackContext *wb_context)
/* Variables for final smoothed_density update */
long new_strategy_delta;
uint32 new_recent_alloc;
+ static int prev_activeNBuffers;
Assert(AmBackgroundWriterProcess());
@@ -3903,6 +3943,24 @@ BgBufferSync(WritebackContext *wb_context)
*/
strategy_buf_id = StrategySyncStart(&strategy_passes, &recent_alloc);
+ if (prev_activeNBuffers != activeNBuffers)
+ {
+ /*
+ * Previous clock sweep position does not make sense if the size of
+ * the active buffer pool has changed.
+ *
+ * TODO: actually we have added a fix in
+ * StrategyAdjustNewBufAllocSize() to adjust the complete_passes so
+ * that we don't have to reset the saved info here, but still the
+ * Assert(StrategyDelta >= 0) below fails without resetting the saved
+ * info. Resetting the saved info means we loose the current position
+ * of the bgwriter and thus may have to unnecessarily scan already
+ * scanned buffers and the new allocations may have to find victim
+ * themselves. Needs further investigation.
+ */
+ saved_info_valid = false;
+ }
+
/* Report buffer alloc counts to pgstat */
PendingBgWriterStats.buf_alloc += recent_alloc;
@@ -3930,7 +3988,7 @@ BgBufferSync(WritebackContext *wb_context)
int32 passes_delta = strategy_passes - prev_strategy_passes;
strategy_delta = strategy_buf_id - prev_strategy_buf_id;
- strategy_delta += (long) passes_delta * NBuffers;
+ strategy_delta += (long) passes_delta * activeNBuffers;
if ((int32) (next_passes - strategy_passes) > 0)
{
@@ -3947,7 +4005,7 @@ BgBufferSync(WritebackContext *wb_context)
next_to_clean >= strategy_buf_id)
{
/* on same pass, but ahead or at least not behind */
- bufs_to_lap = NBuffers - (next_to_clean - strategy_buf_id);
+ bufs_to_lap = activeNBuffers - (next_to_clean - strategy_buf_id);
#ifdef BGW_DEBUG
elog(DEBUG2, "bgwriter ahead: bgw %u-%u strategy %u-%u delta=%ld lap=%d",
next_passes, next_to_clean,
@@ -3969,7 +4027,7 @@ BgBufferSync(WritebackContext *wb_context)
#endif
next_to_clean = strategy_buf_id;
next_passes = strategy_passes;
- bufs_to_lap = NBuffers;
+ bufs_to_lap = activeNBuffers;
}
/*
@@ -3996,12 +4054,13 @@ BgBufferSync(WritebackContext *wb_context)
strategy_delta = 0;
next_to_clean = strategy_buf_id;
next_passes = strategy_passes;
- bufs_to_lap = NBuffers;
+ bufs_to_lap = activeNBuffers;
}
/* Update saved info for next time */
prev_strategy_buf_id = strategy_buf_id;
prev_strategy_passes = strategy_passes;
+ prev_activeNBuffers = activeNBuffers;
saved_info_valid = true;
/*
@@ -4022,7 +4081,7 @@ BgBufferSync(WritebackContext *wb_context)
* strategy point and where we've scanned ahead to, based on the smoothed
* density estimate.
*/
- bufs_ahead = NBuffers - bufs_to_lap;
+ bufs_ahead = activeNBuffers - bufs_to_lap;
reusable_buffers_est = (float) bufs_ahead / smoothed_density;
/*
@@ -4060,7 +4119,7 @@ BgBufferSync(WritebackContext *wb_context)
* the BGW will be called during the scan_whole_pool time; slice the
* buffer pool into that many sections.
*/
- min_scan_buffers = (int) (NBuffers / (scan_whole_pool_milliseconds / BgWriterDelay));
+ min_scan_buffers = (int) (activeNBuffers / (scan_whole_pool_milliseconds / BgWriterDelay));
if (upcoming_alloc_est < (min_scan_buffers + reusable_buffers_est))
{
@@ -4082,13 +4141,24 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
+ /*
+ * Execute the LRU scan on the part of the buffer pool from where the new
+ * allocations will happen.
+ *
+ * Note that activeNBuffers may change during this loop if buffer pool
+ * gets resized concurrently. This may invalidate the num_to_scan count.
+ * If the buffer pool was shrunk this would make background writing more
+ * aggressive, which might be desired since the buffer pool is smaller. If
+ * the buffer pool was grown this would make the background writer not
+ * scan the newly added buffers, which should not have any dirty buffers
+ * yet.
+ */
while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
- if (++next_to_clean >= NBuffers)
+ if (++next_to_clean >= activeNBuffers)
{
next_to_clean = 0;
next_passes++;
@@ -8053,8 +8123,6 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
uint64 buf_state;
bool buffer_flushed;
- CHECK_FOR_INTERRUPTS();
-
buf_state = pg_atomic_read_u64(&desc->state);
if (!(buf_state & BM_VALID))
continue;
@@ -8071,6 +8139,13 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
if (buffer_flushed)
(*buffers_flushed)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8105,8 +8180,6 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
uint64 buf_state = pg_atomic_read_u64(&(desc->state));
bool buffer_flushed;
- CHECK_FOR_INTERRUPTS();
-
/* An unlocked precheck should be safe and saves some cycles. */
if ((buf_state & BM_VALID) == 0 ||
!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
@@ -8133,6 +8206,13 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
if (buffer_flushed)
(*buffers_flushed)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8250,8 +8330,6 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
uint64 buf_state = pg_atomic_read_u64(&(desc->state));
bool buffer_already_dirty;
- CHECK_FOR_INTERRUPTS();
-
/* An unlocked precheck should be safe and saves some cycles. */
if ((buf_state & BM_VALID) == 0 ||
!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
@@ -8277,6 +8355,13 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
(*buffers_already_dirty)++;
else
(*buffers_skipped)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8304,8 +8389,6 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
uint64 buf_state;
bool buffer_already_dirty;
- CHECK_FOR_INTERRUPTS();
-
buf_state = pg_atomic_read_u64(&desc->state);
if (!(buf_state & BM_VALID))
continue;
@@ -8321,6 +8404,13 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
(*buffers_already_dirty)++;
else
(*buffers_skipped)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -9017,3 +9107,57 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ *
+ * If this function encounters a pinned buffer towards the end of the buffer
+ * pool, we would have evicted most of the buffers and yet rollback the resize
+ * operation. If we could find pinned buffers before evicting any, we could save
+ * wasted work and avoid performance impact because of evicted buffers. But
+ * there is no guarantee that a buffer won't be pinned after we check it. So we
+ * have to check for pinned buffers while evicting and rollback if we encounter
+ * any.
+ */
+bool
+EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
+{
+ bool result = true;
+
+ for (Buffer buf = targetNBuffers + 1; buf <= currentNBuffers; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint64 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u64(&desc->state);
+
+ /*
+ * Nobody is expected to allocate new buffers while resizing is going
+ * on hence unlocked precheck should be safe and saves some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 4d5ee52ddc0..e397c5ce47a 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -37,8 +37,9 @@ typedef struct
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo size of the buffer
- * pool.
+ * get an actual buffer, it needs to be used modulo the size of the buffer
+ *
+ * allocation area.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -101,11 +102,52 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
static void AddBufferToRing(BufferAccessStrategy strategy,
BufferDesc *buf);
+/*
+ * StrategyWrapAround - Wrap around the clock-sweep hand.
+ *
+ * `new_pos` is the new_pos position of the clock hand after wrap around.
+ * `num_passes` is the number of completed passes when wrapping around.
+ */
+static void
+StrategyWrapAround(uint32 new_pos, uint32 num_passes)
+{
+ bool success = false;
+
+ while (!success)
+ {
+ uint32 wrapped;
+
+ /*
+ * Acquire the spinlock while increasing completePasses. That allows
+ * other readers to read nextVictimBuffer and completePasses in a
+ * consistent manner which is required for StrategySyncStart(). In
+ * theory delaying the increment could lead to an overflow of
+ * nextVictimBuffers, but that's highly unlikely and wouldn't be
+ * particularly harmful.
+ */
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ wrapped = new_pos % activeNBuffers;
+
+ success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
+ &new_pos, wrapped);
+ if (success)
+ StrategyControl->completePasses += num_passes;
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+ }
+}
+
/*
* ClockSweepTick - Helper routine for StrategyGetBuffer()
*
- * Move the clock hand one buffer ahead of its current position and return the
- * id of the buffer now under the hand.
+ * Move the clock hand one buffer ahead of its current position and return the id
+ * of the buffer now under the hand.
+ *
+ * We use the same value of activeNBuffers through out the function. Hence, even
+ * if the multiple backends end up wrapping around nextVictimBuffer using
+ * different activeNBuffers, they end up increasing completePasses incrementally
+ * and consistent to the respective victims. Use the latest activeNBuffers so as
+ * to be as consistent with the other allocators as possible.
*/
static inline uint32
ClockSweepTick(void)
@@ -119,13 +161,14 @@ ClockSweepTick(void)
*/
victim =
pg_atomic_fetch_add_u32(&StrategyControl->nextVictimBuffer, 1);
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
- if (victim >= NBuffers)
+ if (victim >= activeNBuffers)
{
uint32 originalVictim = victim;
/* always wrap what we look up in BufferDescriptors */
- victim = victim % NBuffers;
+ victim = victim % activeNBuffers;
/*
* If we're the one that just caused a wraparound, force
@@ -134,35 +177,9 @@ ClockSweepTick(void)
* value consisting of nextVictimBuffer and completePasses.
*/
if (victim == 0)
- {
- uint32 expected;
- uint32 wrapped;
- bool success = false;
-
- expected = originalVictim + 1;
-
- while (!success)
- {
- /*
- * Acquire the spinlock while increasing completePasses. That
- * allows other readers to read nextVictimBuffer and
- * completePasses in a consistent manner which is required for
- * StrategySyncStart(). In theory delaying the increment
- * could lead to an overflow of nextVictimBuffers, but that's
- * highly unlikely and wouldn't be particularly harmful.
- */
- SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
-
- wrapped = expected % NBuffers;
-
- success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
- &expected, wrapped);
- if (success)
- StrategyControl->completePasses++;
- SpinLockRelease(&StrategyControl->buffer_strategy_lock);
- }
- }
+ StrategyWrapAround(originalVictim + 1, 1);
}
+
return victim;
}
@@ -180,6 +197,11 @@ ClockSweepTick(void)
*
* The buffer is pinned and marked as owned, using TrackNewBufferPin(),
* before returning.
+ *
+ * We do not process a ProcSignalBarrier between choosing a victim and pinning
+ * it, so the buffer will remain valid even if the buffer pool is shrunk. For
+ * better safety we may want to disable interrupt handling explicitly during
+ * this time.
*/
BufferDesc *
StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
@@ -238,7 +260,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
pg_atomic_fetch_add_u32(&StrategyControl->numBufferAllocs, 1);
/* Use the "clock sweep" algorithm to find a free buffer */
- trycounter = NBuffers;
+ trycounter = activeNBuffers;
for (;;)
{
uint64 old_buf_state;
@@ -291,7 +313,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
local_buf_state))
{
- trycounter = NBuffers;
+ trycounter = activeNBuffers;
break;
}
}
@@ -334,9 +356,15 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
uint32 nextVictimBuffer;
int result;
+ /*
+ * Update backend local activeNBuffers for the same reason as in
+ * StrategyGetBuffer.
+ */
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
- result = nextVictimBuffer % NBuffers;
+ result = nextVictimBuffer % activeNBuffers;
if (complete_passes)
{
@@ -346,7 +374,7 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
* Additionally add the number of wraparounds that happened before
* completePasses could be incremented. C.f. ClockSweepTick().
*/
- *complete_passes += nextVictimBuffer / NBuffers;
+ *complete_passes += nextVictimBuffer / activeNBuffers;
}
if (num_buf_alloc)
@@ -392,6 +420,29 @@ StrategyCtlShmemRequest(void *arg)
);
}
+/*
+ * StrategyAdjustNewBufAllocSize
+ *
+ * Adjust clock hand when resizing the buffer pool.
+ */
+void
+StrategyAdjustNewBufAllocSize(void)
+{
+ int num_passes;
+ uint32 nextVictimBuffer;
+
+ /* Should be called only when resizing is in progress. */
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ /* Consistently wrap around the clock sweep hand, if necessary. */
+ nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
+ num_passes = nextVictimBuffer / activeNBuffers;
+ if (num_passes > 0)
+ StrategyWrapAround(nextVictimBuffer, num_passes);
+}
+
/*
* StrategyCtlShmemInit -- initialize the buffer cache replacement strategy.
*/
@@ -634,12 +685,27 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot is outside
+ * the buffer allocation area tell the caller to allocate a new buffer
+ * with the normal allocation strategy. He will then fill this slot by
+ * calling AddBufferToRing with the new buffer. Usually the buffers in the
+ * ring will be within the buffer allocation area, but if the buffer pool
+ * has been shrunk since the last time the ring was filled, some of the
+ * buffers in the ring may be outside the new buffer allocation area.
+ *
+ * TODO: buffer ids in the ring will never be greater than the size of
+ * buffer pool, except maybe the first time ring is accessed after
+ * shrinking the buffer pool. Checking the upper bound on buffer id always
+ * may mask a bug bugs that introduces buffer ids higher than the size of
+ * buffer pool in the ring. But performing that check only once after
+ * shrinking seems impossible. The BufferAccessStrategy objects are not
+ * accessible outside the ScanState. Hence we can not purge the buffers
+ * while evicting the buffers. After the resizing is finished, it's not
+ * possible to notice when we touch the first of those objects and the
+ * last of objects. See if this can fixed.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer || bufnum > activeNBuffers)
return NULL;
buf = GetBufferDescriptor(bufnum - 1);
diff --git a/src/backend/storage/buffer/meson.build b/src/backend/storage/buffer/meson.build
index ed84bf08971..f219e29d5ef 100644
--- a/src/backend/storage/buffer/meson.build
+++ b/src/backend/storage/buffer/meson.build
@@ -6,4 +6,5 @@ backend_sources += files(
'bufmgr.c',
'freelist.c',
'localbuf.c',
+ 'buf_resize.c',
)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 21a77f98c1d..f97215dae9c 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -28,6 +28,7 @@
#include "replication/logicalworker.h"
#include "replication/slotsync.h"
#include "replication/walsender.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
@@ -598,6 +599,15 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_CHECKSUM_OFF:
processed = AbsorbDataChecksumsBarrier(type);
break;
+ case PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC:
+ processed = ProcessBarrierNewBufferAlloc();
+ break;
+ case PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE:
+ processed = ProcessBarrierBufferPoolResize();
+ break;
+ case PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE:
+ processed = ProcessBarrierBufferPoolSize();
+ break;
}
/*
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 780cdb5e9ff..8332bc0b252 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -43,6 +43,7 @@
#include "postmaster/autovacuum.h"
#include "replication/slotsync.h"
#include "replication/syncrep.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
@@ -583,6 +584,21 @@ InitProcess(void)
* the reasons mentioned there.
*/
ShmemReprotectResizableStructs();
+
+#ifndef EXEC_BACKEND
+
+ /*
+ * Pick up the current buffer pool size from shared memory. Fork'ed
+ * backends would otherwise inherit the postmaster's potentially stale
+ * values. EXEC_BACKEND children do this via the buffer manager attach
+ * callback above.
+ *
+ * We also call this function after ProcSignalInit() for the reasons
+ * specified there, but we need it here so that InitBufferManagerAccess()
+ * can use the current buffer pool size.
+ */
+ BufferManagerInitProc();
+#endif
}
/*
@@ -772,6 +788,21 @@ InitAuxiliaryProcess(void)
* the reasons mentioned there.
*/
ShmemReprotectResizableStructs();
+
+#ifndef EXEC_BACKEND
+
+ /*
+ * Pick up the current buffer pool size from shared memory. Fork'ed
+ * backends would otherwise inherit the postmaster's potentially stale
+ * values. EXEC_BACKEND children do this via the buffer manager attach
+ * callback above.
+ *
+ * We also call this function after ProcSignalInit() for the reasons
+ * specifid there, but we need it here so that InitBufferManagerAccess()
+ * can use the current buffer pool size.
+ */
+ BufferManagerInitProc();
+#endif
}
/*
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index b6bdfe213fe..2586b7087ac 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4288,6 +4288,9 @@ PostgresSingleUserMain(int argc, char *argv[],
/* Initialize size of fast-path lock cache. */
InitializeFastPathLocks();
+ /* Initialize MaxNBuffers for buffer pool resizing. */
+ InitializeMaxNBuffers();
+
/*
* Also call any legacy shmem request hooks that might'be been installed
* by preloaded libraries.
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index ccf845e87b9..4c5cb3c3908 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -142,6 +142,8 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffersGUC = 16384;
+bool finalMaxNBuffers = false;
+int MaxNBuffers = 0;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 815b865aa38..c129d3a8a70 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -610,6 +610,55 @@ InitializeFastPathLocks(void)
pg_nextpower2_32(FastPathLockGroupsPerBackend));
}
+/*
+ * Initialize MaxNBuffers variable with validation.
+ *
+ * This must be called after GUCs have been loaded but before shared memory size
+ * is determined.
+ *
+ * Since MaxNBuffers limits the size of the buffer pool, it must be at least as
+ * much as NBuffersGUC. If MaxNBuffers is 0 (default), set it to
+ * NBuffersGUC. Otherwise, validate that MaxNBuffers is not less than
+ * NBuffersGUC.
+ */
+void
+InitializeMaxNBuffers(void)
+{
+ if (MaxNBuffers == 0) /* default/boot value */
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", NBuffersGUC);
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+
+ /*
+ * We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
+ * However, if the DBA explicitly set max_shared_buffers = 0 in the
+ * config file, then PGC_S_DYNAMIC_DEFAULT will fail to override that
+ * and we must force the matter with PGC_S_OVERRIDE.
+ */
+ if (MaxNBuffers == 0) /* failed to apply it? */
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+ else
+ {
+ if (MaxNBuffers < NBuffersGUC)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("max_shared_buffers (%d) cannot be less than current shared_buffers (%d)",
+ MaxNBuffers, NBuffersGUC),
+ errhint("Increase max_shared_buffers or decrease shared_buffers.")));
+ }
+ }
+
+ Assert(MaxNBuffers > 0);
+ Assert(!finalMaxNBuffers);
+ finalMaxNBuffers = true;
+}
+
/*
* Early initialization of a backend (either standalone or under postmaster).
* This happens even before InitPostgres.
@@ -760,32 +809,38 @@ InitPostgres(const char *in_dbname, Oid dboid,
SharedInvalBackendInit(false);
/*
- * Prevent consuming interrupts between setting ProcSignalInit and setting
- * the initial local data checksum value. If a barrier is emitted, and
- * absorbed, before local cached state is initialized the state transition
- * can be invalid.
+ * Prevent consuming interrupts between ProcSignalInit() and the
+ * initialization of state that is kept in sync with shared memory via
+ * procsignal-based barriers (currently the data_checksum_version cache
+ * and the local NBuffers/activeNBuffers cache). If a barrier is emitted,
+ * and absorbed, before that local cached state is initialized the state
+ * transition can be invalid.
*/
HOLD_INTERRUPTS();
ProcSignalInit(MyCancelKey, MyCancelKeyLength);
/*
- * Initialize a local cache of the data_checksum_version, to be updated by
- * the procsignal-based barriers.
+ * Initialize the per-backend caches that are kept in sync via
+ * procsignal-based barriers: currently the local data_checksum_version
+ * and the local copy of the buffer pool size (NBuffers/activeNBuffers).
*
- * This intentionally happens after initializing the procsignal, otherwise
- * we might miss a state change. This means we can get a barrier for the
- * state we've just initialized.
+ * These initializations intentionally happen after ProcSignalInit(),
+ * otherwise we might miss a state change. This means we may also receive
+ * a barrier for the state we've just initialized.
*
* The postmaster (which is what gets forked into the new child process)
* does not handle barriers, therefore it may not have the current value
- * of LocalDataChecksumState value (it'll have the value read from the
- * control file, which may be arbitrarily old).
+ * of LocalDataChecksumState (it'll have the value read from the control
+ * file, which may be arbitrarily old) or NBuffers/activeNBuffers (which
+ * may have been changed by an online resize after the postmaster
+ * started).
*
* NB: Even if the postmaster handled barriers, the value might still be
* stale, as it might have changed after this process forked.
*/
InitLocalDataChecksumState();
+ BufferManagerInitProc();
/*
* Refresh per-backend protections for resizable shmem structures. Usually
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 1a5a168bc1a..5f1dd338f56 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -2631,7 +2631,7 @@ convert_to_base_unit(double value, const char *unit,
* the value without loss. For example, if the base unit is GUC_UNIT_KB, 1024
* is converted to 1 MB, but 1025 is represented as 1025 kB.
*/
-static void
+void
convert_int_from_base_unit(int64 base_value, int base_unit,
int64 *value, const char **unit)
{
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index 2cbec1981c5..d5099fa9659 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2146,6 +2146,15 @@
max => 'MAX_BACKENDS /* XXX? */',
},
+{ name => "max_shared_buffers", type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+ short_desc => 'Sets the upper limit for the shared_buffers value.',
+ flags => 'GUC_UNIT_BLOCKS',
+ variable => 'MaxNBuffers',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX / 2',
+},
+
{ name => 'max_slot_wal_keep_size', type => 'int', context => 'PGC_SIGHUP', group => 'REPLICATION_SENDING',
short_desc => 'Sets the maximum WAL size that can be reserved by replication slots.',
long_desc => 'Replication slots will be marked as failed, and segments released for deletion or recycling, if this much space is occupied by WAL on disk. -1 means no maximum.',
@@ -2720,13 +2729,15 @@
# We sometimes multiply the number of shared buffers by two without
# checking for overflow, so we mustn't allow more than INT_MAX / 2.
-{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
variable => 'NBuffersGUC',
boot_val => '16384',
- min => '16',
+ min => 'MIN_NUM_BUFFERS',
max => 'INT_MAX / 2',
+ check_hook => 'check_shared_buffers',
+ show_hook => 'show_shared_buffers',
},
{ name => 'shared_memory_initial_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 712172760b3..0487576ffd7 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12715,4 +12715,19 @@
proname => 'hashoid8extended', prorettype => 'int8',
proargtypes => 'oid8 int8', prosrc => 'hashoid8extended' },
+{ oid => '9999', descr => 'resize shared buffers according to the value of GUC `shared_buffers`',
+ proname => 'pg_resize_shared_buffers',
+ provolatile => 'v',
+ prorettype => 'bool',
+ proargtypes => '',
+ prosrc => 'pg_resize_shared_buffers'},
+{ oid => '9998', descr => 'report shared buffer pool resize status',
+ proname => 'pg_get_buffer_resize_status',
+ provolatile => 'v',
+ prorettype => 'record',
+ proargtypes => '',
+ proallargtypes => '{int4,int4,int4,int4}',
+ proargmodes => '{o,o,o,o}',
+ proargnames => '{active_nbuffers,current_nbuffers,target_nbuffers,resizer_pid}',
+ prosrc => 'pg_get_buffer_resize_status'},
]
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index b9449e9aced..04fb70086ad 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -176,6 +176,8 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffersGUC;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
@@ -511,6 +513,7 @@ extern PGDLLIMPORT ProcessingMode Mode;
extern void pg_split_opts(char **argv, int *argcp, const char *optstr);
extern void InitializeMaxBackends(void);
extern void InitializeFastPathLocks(void);
+extern void InitializeMaxNBuffers(void);
extern void InitPostgres(const char *in_dbname, Oid dboid,
const char *username, Oid useroid,
uint32 flags,
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 678065b3de2..a4971cebde4 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -264,6 +264,33 @@ BufMappingPartitionLockByIndex(uint32 index)
return &MainLWLockArray[BUFFER_MAPPING_LWLOCK_OFFSET + index].lock;
}
+/*
+ * BufferControl -- shared area controlling buffer pool
+ *
+ * This structure stores information about the size of the buffer pool and
+ * whether it is being resized.
+ */
+typedef struct BufferControlBlock
+{
+ /*
+ * size of the part of the buffer pool from where buffers are being
+ * allocated to new requests.
+ */
+ pg_atomic_uint32 activeNBuffers;
+
+ /* current size of the buffer pool, in number of buffers */
+ pg_atomic_uint32 currentNBuffers;
+
+ /* target size of the buffer pool, in number of buffers */
+ pg_atomic_uint32 targetNBuffers;
+
+ /*
+ * PID of the backend currently performing a resize, or 0 when no resize
+ * is in progress. Also acts as a lock prohibiting concurrent resizes.
+ */
+ pg_atomic_uint32 resizer_pid;
+} BufferControlBlock;
+
/*
* BufferDesc -- shared descriptor/state data for a single shared buffer.
*
@@ -411,6 +438,7 @@ typedef struct WritebackContext
} WritebackContext;
/* in buf_init.c */
+extern PGDLLIMPORT BufferControlBlock *BufferControl;
extern PGDLLIMPORT BufferDescPadded *BufferDescriptors;
extern PGDLLIMPORT ConditionVariableMinimallyPadded *BufferIOCVArray;
extern PGDLLIMPORT WritebackContext BackendWritebackContext;
@@ -422,9 +450,26 @@ extern PGDLLIMPORT BufferDesc *LocalBufferDescriptors;
static inline BufferDesc *
GetBufferDescriptor(int id)
{
+ BufferDesc *bdesc;
+
Assert(id >= 0 && id < NBuffers);
- return &(BufferDescriptors[id]).bufferdesc;
+ bdesc = &(BufferDescriptors[id]).bufferdesc;
+
+ /*
+ * TODO: This assertion was proposed in
+ * https://www.postgresql.org/message-id/CAExHW5uzRMYVZsXXS3HXXT0fG_sNrpUhUqwP4NorhaCqH9JDhA@mail.gmail.com,
+ * but was ultimately removed since there was no adequate reason to keep
+ * it in the code without shared buffer resizing. With resizing we may
+ * write and rewrite parts of the buffer descriptor array. So it's better
+ * to make sure that the buffer descriptor is initialized correctly. For
+ * now just make sure that the id in the buffer descriptor is the same as
+ * the id used to access it. Later we may want to expand the assertion to
+ * check the BufferDesc invariants or remove this assertion.
+ */
+ Assert(bdesc->buf_id == id);
+
+ return bdesc;
}
static inline BufferDesc *
@@ -594,6 +639,7 @@ extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
extern int StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc);
extern void StrategyNotifyBgWriter(int bgwprocno);
+extern void StrategyAdjustNewBufAllocSize(void);
/* buf_table.c */
extern uint32 BufTableHashCode(BufferTag *tagPtr);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index f1f6e601f51..187e9b29a9e 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -14,6 +14,7 @@
#ifndef BUFMGR_H
#define BUFMGR_H
+#include "fmgr.h"
#include "port/pg_iovec.h"
#include "storage/aio_types.h"
#include "storage/block.h"
@@ -159,6 +160,7 @@ typedef struct ReadBuffersOperation ReadBuffersOperation;
typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
+#define MIN_NUM_BUFFERS 16
extern PGDLLIMPORT int NBuffersGUC;
/* in bufmgr.c */
@@ -167,6 +169,7 @@ extern PGDLLIMPORT int bgwriter_lru_maxpages;
extern PGDLLIMPORT double bgwriter_lru_multiplier;
extern PGDLLIMPORT bool track_io_timing;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int activeNBuffers;
#define DEFAULT_EFFECTIVE_IO_CONCURRENCY 16
#define DEFAULT_MAINTENANCE_IO_CONCURRENCY 16
@@ -371,6 +374,12 @@ extern void MarkDirtyRelUnpinnedBuffers(Relation rel,
extern void MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
int32 *buffers_already_dirty,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int targetNBuffers, int currentNBuffers);
+
+/* in buf_init.c */
+extern bool BufferManagerShmemResize(int currentNBuffers, int targetNBuffers);
+extern void BufferManagerShmemProtect(void);
+extern void BufferManagerInitProc(void);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
@@ -473,4 +482,10 @@ BufferGetPage(Buffer buffer)
#endif /* FRONTEND */
+/* buf_resize.c */
+extern Datum pg_resize_shared_buffers(PG_FUNCTION_ARGS);
+extern bool ProcessBarrierNewBufferAlloc(void);
+extern bool ProcessBarrierBufferPoolResize(void);
+extern bool ProcessBarrierBufferPoolSize(void);
+
#endif /* BUFMGR_H */
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index aaa158bfd66..78dda9c8d21 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,11 @@ typedef enum
PROCSIGNAL_BARRIER_CHECKSUM_INPROGRESS_ON,
PROCSIGNAL_BARRIER_CHECKSUM_INPROGRESS_OFF,
PROCSIGNAL_BARRIER_CHECKSUM_ON,
+ PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC, /* New buffer allocation pool size
+ * changed */
+ PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE, /* Buffer pool shared structures
+ * resized */
+ PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE, /* Buffer pool size updated */
} ProcSignalBarrierType;
/*
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 2a6e2ed18b3..2e3dc919018 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -462,6 +462,8 @@ extern config_handle *get_config_handle(const char *name);
extern void AlterSystemSetConfigFile(AlterSystemStmt *altersysstmt);
extern char *GetConfigOptionByName(const char *name, const char **varname,
bool missing_ok);
+extern void convert_int_from_base_unit(int64 base_value, int base_unit,
+ int64 *value, const char **unit);
extern void TransformGUCArray(ArrayType *array, List **names,
List **values);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index df048517a0f..7e549a36530 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -181,4 +181,6 @@ extern void assign_synchronized_standby_slots(const char *newval, void *extra);
extern bool check_log_min_messages(char **newval, void **extra, GucSource source);
extern void assign_log_min_messages(const char *newval, void *extra);
+extern const char *show_shared_buffers(bool use_units);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
#endif /* GUC_HOOKS_H */
diff --git a/src/test/Makefile b/src/test/Makefile
index 3eb0a06abb4..7a0d74086c1 100644
--- a/src/test/Makefile
+++ b/src/test/Makefile
@@ -20,7 +20,8 @@ SUBDIRS = \
postmaster \
recovery \
regress \
- subscription
+ subscription \
+ buffermgr
ifeq ($(with_icu),yes)
SUBDIRS += icu
diff --git a/src/test/README b/src/test/README
index afdc7676519..77f11607ff7 100644
--- a/src/test/README
+++ b/src/test/README
@@ -15,6 +15,9 @@ examples/
Demonstration programs for libpq that double as regression tests via
"make check"
+buffermgr/
+ Tests for resizing buffer pool without restarting the server
+
isolation/
Tests for concurrent behavior at the SQL level
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
new file mode 100644
index 00000000000..24c245c900a
--- /dev/null
+++ b/src/test/buffermgr/Makefile
@@ -0,0 +1,35 @@
+#-------------------------------------------------------------------------
+#
+# Makefile for src/test/buffermgr
+#
+# Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+# Portions Copyright (c) 1994, Regents of the University of California
+#
+# src/test/buffermgr/Makefile
+#
+#-------------------------------------------------------------------------
+
+EXTRA_INSTALL = contrib/pg_buffercache \
+ src/test/modules/injection_points \
+ src/test/modules/test_shmem
+
+REGRESS = buffer_resize
+
+# Custom configuration for buffer manager tests
+TEMP_CONFIG = $(srcdir)/buffermgr_test.conf
+
+export enable_injection_points
+
+subdir = src/test/buffermgr
+top_builddir = ../../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+
+check:
+ $(prove_check)
+
+installcheck:
+ $(prove_installcheck)
+
+clean distclean:
+ rm -rf tmp_check
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
new file mode 100644
index 00000000000..c375ad80989
--- /dev/null
+++ b/src/test/buffermgr/README
@@ -0,0 +1,26 @@
+src/test/buffermgr/README
+
+Regression tests for buffer manager
+===================================
+
+This directory contains a test suite for resizing buffer manager without restarting the server.
+
+
+Running the tests
+=================
+
+NOTE: You must have given the --enable-tap-tests argument to configure.
+
+Run
+ make check
+or
+ make installcheck
+You can use "make installcheck" if you previously did "make install".
+In that case, the code in the installation tree is tested. With
+"make check", a temporary installation tree is built from the current
+sources and then tested.
+
+Either way, this test initializes, starts, and stops a test Postgres
+cluster.
+
+See src/test/perl/README for more info about running these tests.
diff --git a/src/test/buffermgr/buffermgr_test.conf b/src/test/buffermgr/buffermgr_test.conf
new file mode 100644
index 00000000000..a15f3e442a5
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.conf
@@ -0,0 +1,11 @@
+# Configuration for buffer manager regression tests
+
+# Even if max_shared_buffers is set multiple times only the last one is used to
+# as the limit on shared_buffers.
+max_shared_buffers = 128kB
+# Set initial shared_buffers as expected by test
+shared_buffers = 128MB
+# Set a larger value for max_shared_buffers to allow testing resize operations
+max_shared_buffers = 300MB
+# Turn huge pages off, since that affects the size of memory segments
+huge_pages = off
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
new file mode 100644
index 00000000000..64430b4a3ae
--- /dev/null
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -0,0 +1,290 @@
+-- Test buffer pool resizing and shared memory allocation tracking This test
+-- resizes the buffer pool multiple times and monitors shared memory allocations
+-- related to buffer management
+-- TODOs
+--
+-- 1. The test sets shared_buffers values in MBs. Instead it could use values in
+-- kBs so that the test runs on very small machines.
+--
+-- 2. The size, minimum_size and maximum_size columns in pg_shmem_allocations
+-- for "Buffer Blocks" should be same as the value of GUC shared_buffers. We
+-- should test that.
+--
+-- 3. We should make sure that when the shared_buffers value is increased, the
+-- size and allocated_size for all buffer related shared memory allocations
+-- increases and when the shared_buffers value is decreased, the size and
+-- allocated_size for all buffer related shared memory allocations decreases
+-- proportionately.
+--
+-- 4. allocated_size for allocations should be greater than or equal to size for
+-- all buffer related shared memory allocations. Similarly reserved_space should
+-- be greater than or equal to maximum_size for all buffer related shared memory
+-- allocations. We should test these conditions as well.
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+-- Load test_shmem for test_shmem_pagesize().
+CREATE EXTENSION IF NOT EXISTS test_shmem;
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, size,
+ allocated_size >= size AS alloc_size_cmp,
+ allocated_size - size < 2 * test_shmem_pagesize() AS alloc_size_diff,
+ minimum_size, maximum_size, reserved_space
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 134217728 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1048576 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 262144 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 134217728 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1048576 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 262144 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 128MB (pending: 64MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 64MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 67108864 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 524288 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 131072 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 8192
+(1 row)
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 64MB (pending: 256MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 268435456 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 2097152 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 524288 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 32768
+(1 row)
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 256MB (pending: 100MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 100MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 104857600 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 819200 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 204800 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 12800
+(1 row)
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 100MB (pending: 128kB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+--------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 131072 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1024 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 256 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16
+(1 row)
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+ERROR: invalid value for parameter "shared_buffers": 51200
+DETAIL: "shared_buffers" must be less than "max_shared_buffers".
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+-- TODO: Test that a non-superuser can not invoke pg_resize_shared_buffers()
+-- function.
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
new file mode 100644
index 00000000000..7a6d5e29f8d
--- /dev/null
+++ b/src/test/buffermgr/meson.build
@@ -0,0 +1,25 @@
+# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+
+tests += {
+ 'name': 'buffermgr',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'buffer_resize',
+ ],
+ 'regress_args': ['--temp-config', files('buffermgr_test.conf')],
+ },
+ 'tap': {
+ 'env': {
+ 'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
+ },
+ 'tests': [
+ 't/001_resize_buffer.pl',
+ 't/003_resize_fault_tolerance.pl',
+ 't/004_client_join_buffer_resize.pl',
+ 't/005_resize_failures.pl',
+ 't/006_resize_with_syslogger.pl',
+ ],
+ },
+}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
new file mode 100644
index 00000000000..4219e33a805
--- /dev/null
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -0,0 +1,106 @@
+-- Test buffer pool resizing and shared memory allocation tracking This test
+-- resizes the buffer pool multiple times and monitors shared memory allocations
+-- related to buffer management
+
+-- TODOs
+--
+-- 1. The test sets shared_buffers values in MBs. Instead it could use values in
+-- kBs so that the test runs on very small machines.
+--
+-- 2. The size, minimum_size and maximum_size columns in pg_shmem_allocations
+-- for "Buffer Blocks" should be same as the value of GUC shared_buffers. We
+-- should test that.
+--
+-- 3. We should make sure that when the shared_buffers value is increased, the
+-- size and allocated_size for all buffer related shared memory allocations
+-- increases and when the shared_buffers value is decreased, the size and
+-- allocated_size for all buffer related shared memory allocations decreases
+-- proportionately.
+--
+-- 4. allocated_size for allocations should be greater than or equal to size for
+-- all buffer related shared memory allocations. Similarly reserved_space should
+-- be greater than or equal to maximum_size for all buffer related shared memory
+-- allocations. We should test these conditions as well.
+
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+
+-- Load test_shmem for test_shmem_pagesize().
+CREATE EXTENSION IF NOT EXISTS test_shmem;
+
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, size,
+ allocated_size >= size AS alloc_size_cmp,
+ allocated_size - size < 2 * test_shmem_pagesize() AS alloc_size_diff,
+ minimum_size, maximum_size, reserved_space
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+
+-- TODO: Test that a non-superuser can not invoke pg_resize_shared_buffers()
+-- function.
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
new file mode 100644
index 00000000000..fb5a42be26a
--- /dev/null
+++ b/src/test/buffermgr/t/001_resize_buffer.pl
@@ -0,0 +1,193 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Minimal test testing shared_buffer resizing under load
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Function to check if pgbench is still running.
+#
+# Relying on IPC::Run's pumpable status to check if pgbench is still running has
+# been proven unreliable. Instead we rely on existence of pgbench processes in
+# pg_stat_activity. Since we use -C with pgbench, there can be a non-zero
+# chance that no pgbench process is running even thought pgbench is running. But
+# that's a very rare possibility that can be ignored.
+sub pgbench_processes_active
+{
+ my ($node, $application_name) = @_;
+
+ my $result = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_stat_activity WHERE application_name = '$application_name';");
+ return int($result) > 0;
+}
+
+my $resize_sql_func_def = q{
+create or replace function pg_resize_shared_buffers_sql(new_size int, out num_tries int) returns int as $$
+declare
+ success boolean := false;
+ tries int := 0;
+ cur_setting text;
+ pending_pattern text;
+ target text := new_size::text;
+begin
+ -- Wait until pg_settings reports the new value as pending,
+ -- i.e. "<old value> (pending: <new value>)".
+ pending_pattern := '%(pending: ' || target || ')%';
+ loop
+ select setting into cur_setting
+ from pg_settings where name = 'shared_buffers';
+ exit when cur_setting like pending_pattern or cur_setting = target;
+ perform pg_sleep(0.1);
+ raise notice 'Current setting: %', cur_setting;
+ end loop;
+
+ -- pg_resize_shared_buffers() returns true on success; retry until it succeeds.
+ while not success loop
+ tries := tries + 1;
+ select pg_resize_shared_buffers() into success;
+ if not success then
+ perform pg_sleep(0.1);
+ end if;
+ raise notice 'pg_resize_shared_buffers() attempt %: success = %', tries, success;
+ end loop;
+
+ -- Confirm the new value is in effect (no longer pending).
+ select setting into cur_setting
+ from pg_settings where name = 'shared_buffers';
+ if cur_setting <> target then
+ raise exception 'shared_buffers resize did not take effect: expected %, got %',
+ target, cur_setting;
+ end if;
+
+ num_tries := tries;
+ return;
+end;
+$$ language plpgsql;
+};
+
+# Function to resize buffer pool and verify the change.
+sub apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ # Use the new pg_resize_shared_buffers() interface which handles everything synchronously
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers_sql($new_size)");
+
+ # Any failure in resizing the buffer pool will cause the test to timeout. So
+ # if we reach here, the resize was successful. Just declare it as a
+ # successful test so that we can see progress in the test output.
+ ok(1, "Buffer pool resized to $new_size");
+}
+
+my @buffer_sizes = (128, 28, 16 * 1024, 32 * 1024, 1024, 512, 16, 24, 256, 128 * 1024, 16 * 1204);
+
+# Initialize a cluster and start pgbench in the background for concurrent load.
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Permit resizing up to 1GB for this test and let the server start with 128MB.
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = } . (sort { $b <=> $a } @buffer_sizes)[0] . qq{
+shared_buffers = 16
+log_statement = none
+restart_after_crash = off
+});
+
+$node->start;
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+$node->safe_psql('postgres', $resize_sql_func_def);
+
+my $pgb_scale = 10;
+my $pgb_duration = 120;
+my $pgb_num_clients = 10;
+# make it easy to identify pgbench processes in pg_stat_activity
+my $application_name = 'pgbench_buffer_resize_test';
+$node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$pgb_scale --quiet",
+ 0,
+ [qr{^$}],
+ [ # stderr patterns to verify initialization stages
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$pgb_scale)"
+);
+my ($pgbench_stdin, $pgbench_stdout, $pgbench_stderr) = ('', '', '');
+# Use --exit-on-abort so that the test stops on the first server crash or error,
+# thus making it easy to debug the failure. Use -C to increase the chances of a
+# new backend being created while resizing the buffer pool.
+my $pgbench_process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-h', $node->host,
+ '-T', $pgb_duration,
+ '-c', $pgb_num_clients,
+ '-C',
+ '--exit-on-abort',
+ '--continue-on-error',
+ "dbname=postgres application_name=$application_name"
+ ],
+ '<' => \$pgbench_stdin,
+ '>' => \$pgbench_stdout,
+ '2>' => \$pgbench_stderr
+);
+
+ok($pgbench_process, "pgbench started successfully");
+
+# Resize buffer pool to various sizes while pgbench is running in the
+# background. We use smaller sizes to induce frequent buffer eviction and
+# allocation. Also smaller buffer pool means frequent wraparound in background
+# writer, default buffer allocation strategy and checkpointer.
+#
+# TODO: These are pseudo-randomly picked sizes, but we can do better.
+my $tests_completed = 0;
+
+# Reset background writer stats before starting the resize cycle
+$node->safe_psql('postgres', "SELECT pg_stat_reset_shared('bgwriter')");
+
+# Resize as many times as possible while pgbench is running.
+while (pgbench_processes_active($node, $application_name))
+{
+ for my $target_size (@buffer_sizes)
+ {
+ # Stop if pgbench finished
+ if (!pgbench_processes_active($node, $application_name))
+ {
+ last;
+ }
+
+ apply_and_verify_buffer_change($node, $target_size);
+ $tests_completed++;
+
+ # Wait for the resized buffer pool to stabilize.
+ sleep(1);
+ }
+}
+
+ok($tests_completed > scalar(@buffer_sizes), "All buffer size transitions were tested");
+note "Completed $tests_completed buffer resize operations while pgbench was running";
+
+# Check that the background writer did some work during the resize cycle
+is($node->safe_psql('postgres', "SELECT buffers_clean > 0 FROM pg_stat_bgwriter"), 't', "Background writer ran during resize cycle");
+
+# Make sure that pgbench finishes
+$pgbench_process->signal('TERM');
+ok((IPC::Run::finish $pgbench_process), "pgbench finished successfully");
+
+# Log any error output from pgbench for debugging
+diag("pgbench stderr:\n$pgbench_stderr");
+diag("pgbench stdout:\n$pgbench_stdout");
+
+# Ensure database is still functional after all the buffer changes
+$node->connect_ok("dbname=postgres",
+ "Database remains accessible after $tests_completed buffer resize operations");
+
+done_testing();
diff --git a/src/test/buffermgr/t/003_resize_fault_tolerance.pl b/src/test/buffermgr/t/003_resize_fault_tolerance.pl
new file mode 100644
index 00000000000..366929f4c45
--- /dev/null
+++ b/src/test/buffermgr/t/003_resize_fault_tolerance.pl
@@ -0,0 +1,839 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test that only one pg_resize_shared_buffers() call succeeds when multiple
+# sessions attempt to resize buffers concurrently
+
+use strict;
+use warnings;
+use Config;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# =============================================================================
+# Initialization
+# =============================================================================
+my $initial_nbuffers = 16;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
+$node->append_conf('postgresql.conf', 'max_shared_buffers = 32');
+$node->append_conf('postgresql.conf', 'restart_after_crash = on');
+$node->start;
+
+# Load injection points extension for test coordination
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# =============================================================================
+# Helper functions
+# =============================================================================
+
+# Setup resize operation to be interrupted.
+#
+# Prepare to resize the buffer pool to a target size. Start a resize session
+# through a background psql session. Adjust GUCs for the mode of interruption.
+# If injection point is provided, setup it up with the injection point and wait
+# for the resize session to reach the injection point. The resize session is
+# returned to the caller.
+sub start_resize_session
+{
+ my ($target_nbuffers, $mode, $injection_point) = @_;
+
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$target_nbuffers'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ my $session = $node->background_psql('postgres', on_error_stop => 0);
+
+ my $injection_action;
+ if (defined $injection_point)
+ {
+ $injection_action = ($mode eq 'error') ? 'error' : 'wait';
+ $session->query_safe('SELECT injection_points_set_local()', verbose => 0);
+ $session->query_safe(
+ "SELECT injection_points_attach('$injection_point', '$injection_action')",
+ verbose => 0);
+ }
+
+ apply_session_gucs_for_mode($session, $mode);
+
+ $session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ ));
+
+ # Wait for the pg_resize_shared_buffers to start waiting at the injection
+ # point.
+ if (defined $injection_point && $injection_action eq 'wait')
+ {
+ my $resize_pid = $session->{backend_pid};
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = '$injection_point' FROM pg_stat_activity WHERE pid = $resize_pid")
+ or die "timed out waiting for resize backend $resize_pid at $injection_point";
+ }
+
+ return $session;
+}
+
+# Start a backend which can be used to test the barrier handler fault tolerance.
+# We use a long pg_sleep() to simulate a load that checks for interrupts
+# regularly. The given injection point is attached to the peer backend locally
+# to induce a fault in barrier handler.
+sub start_peer_session_with_injection_point
+{
+ my ($injection_point, $action) = @_;
+
+ my $session = $node->background_psql('postgres', on_error_stop => 0);
+
+ $session->query_safe('SELECT injection_points_set_local()', verbose => 0);
+ $session->query_safe("SELECT injection_points_attach('$injection_point', '$action')",
+ verbose => 0);
+ $session->query_until(
+ qr/starting_sleep/,
+ q(
+ \echo starting_sleep
+ SELECT pg_sleep(60);
+ ));
+
+ my $peer_pid = $session->{backend_pid};
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = 'PgSleep' FROM pg_stat_activity WHERE pid = $peer_pid")
+ or die "timed out waiting for peer $peer_pid to enter pg_sleep";
+
+ return $session;
+}
+
+# Apply per-mode session GUCs locally in the given session if required.
+sub apply_session_gucs_for_mode
+{
+ my ($session, $mode) = @_;
+
+ # In timeout mode, set a statement timeout long enough for the resizing
+ # session to reach the injection point and stay there but short enough that
+ # the test doesn't take too long to fail if something goes wrong.
+ if ($mode eq 'timeout')
+ {
+ $session->query_safe("SET statement_timeout = '500ms'", verbose => 0);
+ }
+
+ # Let the resize session detect a client disconnection when testing client
+ # disconnections.
+ if ($mode eq 'disconnect')
+ {
+ $session->query_safe("SET client_connection_check_interval = '100ms'",
+ verbose => 0);
+ }
+}
+
+# Administer the interrupt corresponding to $mode against a resize session
+# that is waiting to be interrupted while resizing the buffer pool.
+sub interrupt_resize_session
+{
+ my ($mode, $session) = @_;
+
+ if ($mode eq 'terminate')
+ {
+ $node->safe_psql('postgres', "SELECT pg_terminate_backend(" . $session->{backend_pid} . ")");
+ }
+ elsif ($mode eq 'cancel')
+ {
+ $node->safe_psql('postgres', "SELECT pg_cancel_backend(" . $session->{backend_pid} . ")");
+ }
+ elsif ($mode eq 'disconnect')
+ {
+ $session->{run}->kill_kill;
+ }
+ elsif ($mode eq 'timeout')
+ {
+ # Nothing to do; statement_timeout will fire from within the resize
+ # session itself.
+ }
+ elsif ($mode eq 'error')
+ {
+ # Nothing to do; the injection point raised ERROR from within the
+ # resize backend itself.
+ }
+ else
+ {
+ die "interrupt_resize_session: unknown mode '$mode'";
+ }
+}
+
+# Function to perform checks after the resize operation has been interrupted. As
+# a result of the interruption, the resize function may finish rolling back the
+# resize or the backend executing that function may exit rolling back the resize
+# or the postmaster may restart all the backends. Perform appropriate checks by
+# detecting the post-interrupt state.
+#
+# - sentinel_session: a background psql session that is used to detect whether the
+# postmaster restarted all backends or not.
+# - resize_session: the background psql session that was executing the resize
+# operation and was interrupted.
+# - log_offset: the offset in the server log file before the resize operation was
+# initiated.
+# - injection_point and mode: the injection point and mode of interruption that
+# was used to interrupt the resize operation.
+# - orig_nbuffers and target_nbuffers: the original and target buffer sizes for
+# the resize operation.
+# - test_label: a label to create unique test names for different tests
+sub check_interrupted_resize
+{
+ my ($sentinel_session, $resize_session, $log_offset, $mode,
+ $injection_point, $orig_nbuffers, $target_nbuffers, $test_label) = @_;
+
+ my $resize_pid = $resize_session->{backend_pid};
+ my $sentinel_pid = $sentinel_session->{backend_pid};
+
+ # Wait until the resize backend is no longer running the resize query.
+ $node->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity "
+ . "WHERE pid = $resize_pid AND state = 'active' "
+ . "AND query LIKE '%pg_resize_shared_buffers%'")
+ or die
+ "timed out waiting for resize backend $resize_pid to finish";
+
+ # Wait for the postmaster to be ready in case it restarted the backends.
+ $node->poll_query_until('postgres', 'SELECT true')
+ or die "timed out waiting for postmaster liveliness check";
+
+ # Confirm the resize backend was interrupted by the intended signal.
+ # Match against the log line emitted by the resize PID so we don't
+ # accidentally pick up an unrelated message.
+ my %expected_msg = (
+ terminate => 'terminating connection due to administrator command',
+ cancel => 'canceling statement due to user request',
+ timeout => 'canceling statement due to statement timeout',
+ disconnect => 'connection to client lost',
+ error => "error triggered for injection point $injection_point",
+ );
+ my $log_pattern = qr/\[$resize_pid\][^\n]*\Q$expected_msg{$mode}\E/;
+
+ $node->wait_for_log($log_pattern, $log_offset);
+ ok($node->log_contains($log_pattern, $log_offset),
+ "$test_label: server log shows expected $mode message from pid $resize_pid"
+ );
+
+ my $server_restarted = $node->safe_psql('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE pid = $sentinel_pid"
+ ) eq 't';
+
+ if ($server_restarted)
+ {
+ # Postmaster restarted all backends; sentinel and resize sessions
+ # are dead, just reap their IPC::Run handles.
+ $sentinel_session->finish;
+ $resize_session->finish;
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after crash recovery");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after crash recovery");
+ }
+ else
+ {
+ $sentinel_session->quit;
+ # The resize session may have exited (e.g. on FATAL or disconnect).
+ if ($node->safe_psql('postgres',
+ "SELECT count(*) = 1 FROM pg_stat_activity WHERE pid = $resize_pid") eq 't')
+ {
+ $resize_session->quit;
+ }
+ else
+ {
+ $resize_session->finish;
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$orig_nbuffers|$orig_nbuffers|$orig_nbuffers|0",
+ "$test_label: buffer resize rolled back after $mode");
+
+ # TODO: Also check that the pg_shmem_allocations values are not changed
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$orig_nbuffers (pending: $target_nbuffers)",
+ "$test_label: pg_settings reports pending new value after $mode");
+
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't',
+ "$test_label: resize succeeds after interrupted resize is cleaned up");
+ }
+}
+
+# =============================================================================
+# Concurrent resize test functions
+#
+# Verify that only one pg_resize_shared_buffers() call can succeed at a time
+# using injection points.
+# =============================================================================
+
+# Workhorse function:
+#
+# Make the resize session wait at the given injection point and start another
+# concurrent resize session. The concurrent resize should fail.
+sub test_concurrent_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $session = start_resize_session($target_nbuffers, 'concurrent_resize',
+ $injection_point);
+ my $resize_pid = $session->{backend_pid};
+
+ is($node->safe_psql('postgres',
+ "SELECT resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$resize_pid", "$test_label: resizer_pid reports resize backend");
+
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 'f', "$test_label: concurrent resize fails");
+
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('$injection_point')");
+
+ $session->quit;
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool resized to target after wakeup");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after wakeup");
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points
+sub test_concurrent_resize
+{
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'pgrsb-buffer-pool-resize-barrier-sent',
+ );
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = (['expand', 24], ['shrink', $initial_nbuffers]);
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_concurrent_resize_at_injection_point($point, $target,
+ "$name: $point");
+ }
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test resize operation interruption
+#
+# Verify that an interruption in resize operation does not leave the buffer pool
+# in an inconsistent state.
+# =============================================================================
+
+# Workhorse function:
+#
+# Interrupt pg_resize_shared_buffers() when it is waiting on an injection point.
+# Check that the buffer pool is left in a consistent state as an aftermath.
+#
+# $mode selects how the resize session is interrupted.
+# - 'terminate' - SIGTERM via pg_terminate_backend() from another session.
+# - 'cancel' - SIGINT via pg_cancel_backend() from another session.
+# - 'timeout' - statement_timeout fires inside the resize session itself.
+# - 'disconnect' - the resize session's client connection is closed abruptly.
+# - 'error' - the injection point itself raises ERROR from within the
+# resize backend.
+sub test_interrupt_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $mode, $test_label) = @_;
+
+ my $orig_nbuffers = $node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+ my $log_offset = -s $node->logfile;
+
+ # Start a sentinel session that will be used to detect whether the
+ # postmaster restarted all backends or not after the resize session is
+ # interrupted.
+ my $sentinel_session = $node->background_psql('postgres', on_error_stop => 0);
+
+ my $resize_session = start_resize_session($target_nbuffers, $mode,
+ $injection_point);
+
+ interrupt_resize_session($mode, $resize_session);
+
+ check_interrupted_resize($sentinel_session, $resize_session, $log_offset,
+ $mode, $injection_point, $orig_nbuffers, $target_nbuffers,
+ $test_label);
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points passing it the
+# given mode of interruption.
+sub test_interrupt_resize_session
+{
+ my ($mode) = @_;
+
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'buffer-mgr-resize-struct',
+ 'pgrsb-buffer-pool-resize-barrier-sent',
+ );
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = (['expand', 24], ['shrink', $initial_nbuffers]);
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "$mode: buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_interrupt_resize_at_injection_point($point, $target, $mode,
+ "$mode $name: $point");
+ }
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "$mode: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test error handling in barrier handler
+#
+# Verify that an error in barrier handler does not cause a resize session to
+# fail. The barrier handler may run in a peer backend or the backend which is
+# performing the resize itself.
+# =============================================================================
+
+# Workhorse function:
+#
+# Make a peer session wait at the given injection point in the barrier handler
+# and simulate an error in the handler.
+sub test_error_in_barrier_handler_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $peer_session = start_peer_session_with_injection_point($injection_point, 'error');
+ my $peer_pid = $peer_session->{backend_pid};
+
+ my $log_offset = -s $node->logfile;
+
+ # Resize the buffer pool which will send a barrier to the peer backend
+ # simulating an error in the barrier handler.
+ $node->safe_psql('postgres',"ALTER SYSTEM SET shared_buffers = '$target_nbuffers'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"), 't');
+
+ # Confirm the peer raised the expected error from inside the handler.
+ my $log_pattern = qr/\[$peer_pid\][^\n]*\Qerror triggered for injection point $injection_point\E/;
+ $node->wait_for_log($log_pattern, $log_offset);
+ ok($node->log_contains($log_pattern, $log_offset),
+ "$test_label: server log shows error from peer pid $peer_pid at $injection_point"
+ );
+
+ # Check that the resize was completed as expected
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size");
+
+ # pg_resize_shared_buffers() returns only after every peer has
+ # acknowledged the barrier, so by this point the erroring peer has
+ # already left procArray. Assert that and then reap its IPC::Run handle.
+ is($node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_stat_activity WHERE pid = $peer_pid"),
+ '0',
+ "$test_label: peer pid $peer_pid exited after handler error");
+
+ $peer_session->finish;
+}
+
+# Driver function:
+#
+# Simulate a failure to change the protection on the shared memory. This should
+# cause the barrier handler to raise an error. The barrier handler may run in a
+# peer backend or the backend which is performing the resize itself. The resize
+# session should still complete successfully.
+sub test_error_in_barrier_handler
+{
+ my $injection_point = 'buffer-mgr-protect-struct';
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "error-in-handler: buffer pool size is $initial_nbuffers at start");
+
+ # Test error in barrier handler in a peer backend. Expand then shrink so
+ # the pool returns to its starting size.
+ for my $dir (['expand', 24], ['shrink', $initial_nbuffers])
+ {
+ my ($name, $target) = @$dir;
+ test_error_in_barrier_handler_at_injection_point($injection_point,
+ $target, "error-in-handler peer $name");
+ }
+
+ # Test the same error in the barrier handler in the resize backend itself.
+ # Expand then shrink so the pool returns to its starting size.
+ for my $dir (['expand', 24], ['shrink', $initial_nbuffers])
+ {
+ my ($name, $target) = @$dir;
+ test_interrupt_resize_at_injection_point($injection_point,
+ $target, 'error', "error-in-handler resize-backend $name");
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "error-in-handler: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test fault tolerance of resize operation waiting for barrier
+#
+# Test that, when interrupted, a resizing operation waiting for a barrier to be
+# acknowledged doesn't leave the buffer pool in an inconsistent state.
+# =============================================================================
+
+# Workhorse function:
+#
+# We start a peer session with the given injection point in the barrier handler
+# code attached locally. Once the resize operation starts, the peer session will
+# hit the injection point and wait there. Interrupt the resize session and check
+# that the buffer pool is left in a consistent state as an aftermath.
+#
+# - injection_point: the injection point to attach to the peer session.
+# - target_nbuffers: the target buffer size for the resize operation.
+# - mode: the mode of interruption to apply to the resize session.
+# - test_label: a label to create unique test names for different tests
+sub test_fault_resize_waiting_barrier
+{
+ my ($injection_point, $target_nbuffers, $mode, $test_label) = @_;
+
+ my $orig_nbuffers = $node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+ my $log_offset = -s $node->logfile;
+
+ my $peer_session = start_peer_session_with_injection_point($injection_point, 'wait');
+ my $peer_pid = $peer_session->{backend_pid};
+
+ # Sentinel session to detect a postmaster restart.
+ my $sentinel_session = $node->background_psql('postgres', on_error_stop => 0);
+
+ my $resize_session = start_resize_session($target_nbuffers, $mode);
+ my $resize_pid = $resize_session->{backend_pid};
+
+ # Wait for the peer to reach the injection point. At this point the resize
+ # backend should be blocked in WaitForProcSignalBarrier.
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = '$injection_point' FROM pg_stat_activity WHERE pid = $peer_pid")
+ or die "$test_label: timed out waiting for peer $peer_pid at $injection_point";
+ is($node->safe_psql('postgres',
+ "SELECT wait_event FROM pg_stat_activity WHERE pid = $resize_pid"),
+ 'ProcSignalBarrier',
+ "$test_label: resize $resize_pid is waiting at ProcSignalBarrier");
+
+ interrupt_resize_session($mode, $resize_session);
+
+ check_interrupted_resize($sentinel_session, $resize_session, $log_offset,
+ $mode, $injection_point, $orig_nbuffers, $target_nbuffers, $test_label);
+
+ # Cleanup peer session. If the postmaster restarted all backends, the peer
+ # backend is already gone.
+ if ($node->safe_psql('postgres',
+ "SELECT count(*) = 1 FROM pg_stat_activity WHERE pid = $peer_pid") eq 't')
+ {
+ $peer_session->quit;
+ }
+ else
+ {
+ $peer_session->finish;
+ }
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points passing it the
+# given mode of interruption.
+sub test_fault_resize_waiting_barrier_for_mode
+{
+ my ($mode) = @_;
+
+ my @injection_points = (
+ 'pgrsb-handle-new-buffer-alloc-barrier',
+ 'pgrsb-handle-buffer-pool-size-barrier',
+ 'pgrsb-handle-buffer-pool-resize-barrier',
+ );
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = (['expand', 24], ['shrink', $initial_nbuffers]);
+
+ is($node->safe_psql('postgres', "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "fault-resize-on-peer $mode: buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_fault_resize_waiting_barrier($point, $target, $mode,
+ "fault-resize-on-peer $mode $name: $point");
+ }
+ }
+
+ is($node->safe_psql('postgres', "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "fault-resize-on-peer $mode: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test server restart during a resize
+#
+# Verify that a server can be stopped and started while a resize operation is in
+# progress and the server is started with buffer pool in a consistent state that
+# reflects the target size.
+# =============================================================================
+
+# Workhorse function for fast/immediate shutdown:
+#
+# Make the resize session wait at the given injection point and restart the
+# server in the given mode.
+sub test_server_restart_during_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $stop_mode, $test_label) = @_;
+
+ my $resize_session = start_resize_session($target_nbuffers,
+ 'server_restart', $injection_point);
+
+ $node->stop($stop_mode);
+
+ # Cleanup resize session, the backend must have gone now.
+ $resize_session->finish;
+
+ $node->start;
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after $stop_mode restart");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after $stop_mode restart");
+}
+
+# Workhorse function for smart shutdown:
+#
+# Let the resize operation wait at the given injection point, send a smart
+# shutdown asynchronously. Once the postmaster enters smart shutdown, wakeup the
+# resize backend and let it complete. Verify that the pool reflects the target
+# size immediately and also after the restart.
+#
+# - injection_point: the injection point to park the resize backend at.
+# - target_nbuffers: the target buffer size for the resize operation.
+# - test_label: a label to create unique test names for different tests
+sub test_server_restart_smart_during_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $log_offset = -s $node->logfile;
+
+ my $resize_session = start_resize_session($target_nbuffers, 'server_restart', $injection_point);
+ my $resize_pid = $resize_session->{backend_pid};
+
+ # Open another session which can be used to wakeup the resize backend.
+ my $control_session = $node->background_psql('postgres', on_error_stop => 0);
+
+ # Start the process to stop the server in smart mode.
+ local %ENV = $node->_get_env();
+ my @stop_cmd = ('pg_ctl', '--pgdata' => $node->data_dir, '--mode' => 'smart', 'stop');
+ my ($stop_in, $stop_out, $stop_err) = ('', '', '');
+ my $stop_session = IPC::Run::start(\@stop_cmd,
+ \$stop_in, \$stop_out, \$stop_err);
+
+ # Confirm the postmaster entered smart shutdown.
+ $node->wait_for_log(qr/received smart shutdown request/, $log_offset);
+
+ # Make sure that the resize backend is still alive
+ is($control_session->query("SELECT wait_event FROM pg_stat_activity WHERE pid = $resize_pid"),
+ $injection_point,
+ "$test_label: resize backend alive during smart shutdown");
+
+ # Wake up the resize backnd and let it finish.
+ $control_session->query_safe("SELECT injection_points_wakeup('$injection_point')",
+ verbose => 0);
+
+ # Check that the resize finished successfully by querying from the same
+ # session. The queries won't return if the resize didn't finish. Accomodate
+ # the output 't' from pg_resize_shared_buffers() in the expected output of
+ # the first query.
+ is($resize_session->query("SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()",
+ verbose => 0),
+ "t\n$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: pg_resize_shared_buffers() succeeded and pool at target during smart shutdown");
+ is($resize_session->query("SELECT setting FROM pg_settings WHERE name = 'shared_buffers'",
+ verbose => 0),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size during smart shutdown");
+
+ $resize_session->quit;
+ $control_session->quit;
+
+ # Wait for server to stop
+ IPC::Run::finish($stop_session)
+ or die "$test_label: pg_ctl smart stop failed: $stop_err";
+
+ # Sync Cluster.pm internal state and start the cluster back.
+ $node->{_pid} = undef;
+ $node->start;
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after smart restart");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after smart restart");
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for every injection point in resize operation in
+# both directions for the given stop mode.
+sub test_server_restart_during_resize
+{
+ my ($stop_mode) = @_;
+
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'pgrsb-buffer-pool-resize-barrier-sent',
+ );
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = (['expand', 24], ['shrink', $initial_nbuffers]);
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "server-restart $stop_mode: buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+ my $label = "server-restart $stop_mode $name: $point";
+
+ if ($stop_mode eq 'smart')
+ {
+ test_server_restart_smart_during_resize_at_injection_point(
+ $point, $target, $label);
+ }
+ else
+ {
+ test_server_restart_during_resize_at_injection_point($point,
+ $target, $stop_mode, $label);
+ }
+ }
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "server-restart $stop_mode: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Run tests
+# =============================================================================
+test_concurrent_resize();
+test_error_in_barrier_handler();
+
+test_interrupt_resize_session('terminate');
+test_interrupt_resize_session('cancel');
+test_interrupt_resize_session('timeout');
+test_interrupt_resize_session('error');
+
+# A resize session waiting for a barrier to be acknowledged can not be
+# interrupted by an error. Hence don't test that mode.
+test_fault_resize_waiting_barrier_for_mode('terminate');
+test_fault_resize_waiting_barrier_for_mode('cancel');
+test_fault_resize_waiting_barrier_for_mode('timeout');
+
+test_server_restart_during_resize('immediate');
+test_server_restart_during_resize('fast');
+test_server_restart_during_resize('smart');
+
+# client_connection_check_interval is only effective on systems that expose
+# POLLRDHUP/EPOLLRDHUP (Linux, and a few other Unix variants). On other
+# platforms the GUC is silently a no-op, so the disconnect test would hang.
+if ($Config::Config{osname} eq 'linux')
+{
+ test_interrupt_resize_session('disconnect');
+ test_fault_resize_waiting_barrier_for_mode('disconnect');
+}
+else
+{
+ diag("skipping disconnect interrupt test on $Config::Config{osname} "
+ . "(requires POLLRDHUP support)");
+}
+
+done_testing();
+
+# Few more tests to add but may be somewhere else
+# TODO: test when there are backends that have not attached to the shared memory
+# TODO: test that a non-superuser cannot run pg_resize_shared_buffers()
+# TODO: the resize_sql_func_def in 001_resize_buffer may be useful in other
+# tests (not necessarily this one). Maybe we can use it in other tests where we
+# are looping in TAP test code.
diff --git a/src/test/buffermgr/t/004_client_join_buffer_resize.pl b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
new file mode 100644
index 00000000000..fda0f01bb27
--- /dev/null
+++ b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
@@ -0,0 +1,221 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test shared_buffer resizing coordination with client connections joining using injection points
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+use Time::HiRes qw(sleep);
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Function to calculate the size of test table required to fill up maximum
+# buffer pool when populating it.
+sub calculate_test_sizes
+{
+ my ($node, $block_size) = @_;
+
+ # Get the maximum buffer pool size from configuration
+ my $max_shared_buffers = $node->safe_psql('postgres', "SHOW max_shared_buffers");
+ my ($max_val, $max_unit) = ($max_shared_buffers =~ /(\d+)(\w+)/);
+ my $max_size_bytes;
+ if (lc($max_unit) eq 'kb') {
+ $max_size_bytes = $max_val * 1024;
+ } elsif (lc($max_unit) eq 'mb') {
+ $max_size_bytes = $max_val * 1024 * 1024;
+ } elsif (lc($max_unit) eq 'gb') {
+ $max_size_bytes = $max_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $max_size_bytes = $max_val * 1024;
+ }
+
+ # Fill more pages than minimally required to increase the chances of pages
+ # from the test table filling the buffer cache.
+ $max_size_bytes = $max_size_bytes;
+ my $pages_needed = int($max_size_bytes / $block_size) + 10; # Add some extra to ensure buffers are filled
+ my $rows_to_insert = $pages_needed * 100; # Assuming roughly 100 rows per page for our table structure
+ return ($max_size_bytes, $pages_needed, $rows_to_insert);
+}
+
+# Function to calculate expected buffer count from size string
+sub calculate_buffer_count
+{
+ my ($size_string, $block_size) = @_;
+ # Parse size and convert to bytes
+ my ($size_val, $unit) = ($size_string =~ /(\d+)(\w+)/);
+ my $size_bytes;
+ if (lc($unit) eq 'kb') {
+ $size_bytes = $size_val * 1024;
+ } elsif (lc($unit) eq 'mb') {
+ $size_bytes = $size_val * 1024 * 1024;
+ } elsif (lc($unit) eq 'gb') {
+ $size_bytes = $size_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $size_bytes = $size_val * 1024;
+ }
+ return int($size_bytes / $block_size);
+}
+
+# Initialize cluster with very small buffer sizes for testing
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Configure for buffer resizing with very small buffer pool sizes for faster tests.
+# TODO: for some reason parallel workers try to load default number of shared_buffers which doesn't work with lower max_shared_buffers. We need to fix that - somewhere it's picking default value of shared buffers. For now disable parallelism
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 512kB
+shared_buffers = 320kB
+max_parallel_workers_per_gather = 0
+});
+$node->start;
+
+# Enable injection points
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Get the block size (this is fixed for the binary)
+my $block_size = $node->safe_psql('postgres', "SHOW block_size");
+
+# Try to create pg_buffercache extension for buffer analysis
+eval {
+ $node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+};
+if ($@) {
+ $node->stop;
+ plan skip_all => 'pg_buffercache extension not available - cannot verify buffer usage';
+}
+
+# Create a small test table, and fetch its properties for later reference if required.
+$node->safe_psql('postgres', qq{
+ CREATE TABLE client_test (c1 int, data char(50));
+});
+my $table_oid = $node->safe_psql('postgres', "SELECT oid FROM pg_class WHERE relname = 'client_test'");
+my $table_relfilenode = $node->safe_psql('postgres', "SELECT relfilenode FROM pg_class WHERE relname = 'client_test'");
+note("Test table client_test: OID = $table_oid, relfilenode = $table_relfilenode");
+my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $block_size);
+
+# Create dedicated sessions for injection point handling and test queries,
+# so that we don't create new backends for test operations after starting
+# resize operation. Only one backend, which tests new backend synchronization
+# with resizing operation, should start after resizing has commenced.
+my $injection_session = $node->background_psql('postgres');
+my $query_session = $node->background_psql('postgres');
+my $resize_session = $node->background_psql('postgres');
+
+# Function to run a single injection point test
+sub run_injection_point_test
+{
+ my ($test_name, $injection_point, $target_size, $operation_type) = @_;
+
+ # Silence the logging of the statements we run to avoid
+ # unnecessarily bloating the test logs. This runs before the
+ # upgrade we're testing, so the details should not be very
+ # interesting for debugging. But if needed, you can make it more
+ # verbose by setting this.
+ my $verbose = 0;
+
+ note("Test with $test_name ($operation_type)");
+
+ # Calculate test parameters before starting resize
+ my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $target_size, $block_size);
+
+ # Update buffer pool size and wait for it to reflect pending state
+ $resize_session->query_safe("ALTER SYSTEM SET shared_buffers = '$target_size'", verbose => $verbose);
+ $resize_session->query_safe("SELECT pg_reload_conf()", verbose => $verbose);
+ my $pending_size_str = "pending: $target_size";
+ $resize_session->poll_query_until("SELECT substring(current_setting('shared_buffers'), '$pending_size_str')", $pending_size_str, verbose => $verbose);
+
+ # Set up injection point in injection session
+ $injection_session->query_safe("SELECT injection_points_attach('$injection_point', 'wait')", verbose => $verbose);
+
+ # Trigger resize
+ $resize_session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+ );
+
+ # Wait until resize actually reaches the injection point using the query session
+ $query_session->wait_for_event('client backend', $injection_point, verbose => $verbose);
+
+ # Start a client while resize is paused
+ my $client = $node->background_psql('postgres');
+ note("Background client backend PID: " . $client->query_safe("SELECT pg_backend_pid()", verbose => $verbose));
+
+ # Wake up the injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_wakeup('$injection_point')", verbose => $verbose);
+
+ # Test buffer functionality immediately after waking up injection point
+ # Insert data to test buffer pool functionality during/after resize
+ $client->query_safe("INSERT INTO client_test SELECT i, 'test_data_' || i FROM generate_series(1, $rows_to_insert) i", verbose => $verbose);
+ # Verify the data was inserted correctly and can be read back
+ is($client->query_safe("SELECT COUNT(*) FROM client_test", verbose => $verbose), $rows_to_insert, "inserted $rows_to_insert during $test_name ($operation_type) successful");
+
+ # Verify table size is reasonable (should be substantial for testing)
+ ok($query_session->query_safe("SELECT pg_total_relation_size('client_test')", verbose => $verbose) >= $max_size_bytes,"table size is large enough to overflow buffer pool in test $test_name ($operation_type)");
+
+ # Wait for the resize operation to complete. There is no direct way to do so
+ # in background_psql. Hence fire a psql command and wait for it to finish
+ $resize_session->query(q(\echo 'done'), verbose => $verbose);
+
+ # Detach injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_detach('$injection_point')", verbose => $verbose);
+
+ # Verify resize completed successfully
+ is($query_session->query_safe("SELECT current_setting('shared_buffers')", verbose => $verbose), $target_size,
+ "resize completed successfully to $target_size");
+
+ # Check buffer pool size using pg_buffercache after resize completion
+ is($query_session->query_safe("SELECT COUNT(*) FROM pg_buffercache", verbose => $verbose), calculate_buffer_count($target_size, $block_size), "all buffers in the buffer pool used in $test_name ($operation_type)");
+
+ # Wait for client to complete
+ ok($client->quit, "client succeeded during $test_name ($operation_type)");
+
+ # Clean up for next test
+ $query_session->query_safe("DELETE FROM client_test", verbose => $verbose);
+}
+
+# Test new client joining during various phases of buffer resizing operation using injection points
+my @injection_tests = (
+ {
+ name => 'flag setting phase',
+ injection_point => 'pg-resize-shared-buffers-flag-set',
+ },
+ {
+ name => 'new buffer alloc barrier complete',
+ injection_point => 'pgrsb-new-buffer-alloc-barrier-sent',
+ },
+ {
+ name => 'buffer pool size barrier complete',
+ injection_point => 'pgrsb-buffer-pool-size-barrier-sent',
+ },
+ {
+ name => 'buffer pool resize barrier complete',
+ injection_point => 'pgrsb-buffer-pool-resize-barrier-sent',
+ },
+);
+
+foreach my $test (@injection_tests)
+{
+ # Test shrinking scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '272kB', 'shrinking');
+
+ # Test expanding scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '400kB', 'expanding');
+}
+
+$injection_session->quit;
+$query_session->quit;
+$resize_session->quit;
+
+done_testing();
diff --git a/src/test/buffermgr/t/005_resize_failures.pl b/src/test/buffermgr/t/005_resize_failures.pl
new file mode 100644
index 00000000000..820aa07f006
--- /dev/null
+++ b/src/test/buffermgr/t/005_resize_failures.pl
@@ -0,0 +1,171 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Test that pg_resize_shared_buffers() rolls back cleanly when resize fails.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $have_injection_points = ($ENV{enable_injection_points} eq 'yes');
+
+# Start the pool large enough that there is room to shrink below a pinned
+# buffer while still satisfying the shared_buffers GUC minimum.
+my $initial_nbuffers = 24;
+my $max_nbuffers = 32;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+if ($have_injection_points)
+{
+ $node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+}
+$node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
+$node->append_conf('postgresql.conf', "max_shared_buffers = $max_nbuffers");
+$node->start;
+
+# pg_buffercache lets us locate the bufferid holding a given page.
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+if ($have_injection_points)
+{
+ $node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+}
+
+# ---------------------------------------------------------------------------
+# Test the case when shrinking is aborted by a pinned buffer
+# ---------------------------------------------------------------------------
+
+my $min_nbuffers = $node->safe_psql('postgres',
+ "SELECT min_val::int FROM pg_settings WHERE name = 'shared_buffers'");
+
+# In order to reliably pin a buffer above $min_nbuffers, we create as many
+# tables $min_nbuffers + 1, open a cursor on the tables and fetch one row from
+# each cursor one at a time. This will pin one buffer per table, guaranteeing
+# that at least one of the pinned buffers will be above $min_nbuffers.
+my $ntables = $min_nbuffers + 1;
+for my $i (1 .. $ntables)
+{
+ $node->safe_psql('postgres', "CREATE TABLE evict_target_$i AS SELECT generate_series(1, 2) AS i");
+}
+my $pinner = $node->background_psql('postgres', on_error_stop => 0);
+$pinner->query_safe("BEGIN", verbose => 0);
+my $pinned_buf = 0;
+for my $i (1 .. $ntables)
+{
+ $pinner->query_safe("DECLARE c_$i CURSOR FOR SELECT * FROM evict_target_$i", verbose => 0);
+ $pinner->query_safe("FETCH 1 FROM c_$i", verbose => 0);
+
+ my $buf = $node->safe_psql('postgres',
+ "SELECT min(bufferid) FROM pg_buffercache WHERE pinning_backends > 0 AND bufferid > $min_nbuffers");
+
+ if ($buf =~ /^\d+$/)
+ {
+ $pinned_buf = $buf;
+ last;
+ }
+}
+cmp_ok($pinned_buf, '>', $min_nbuffers, "pinned a buffer above $min_nbuffers");
+
+# Set the target so that the pinned buffer is in the range of buffers to be evicted.
+my $shrink_target = $pinned_buf - 1;
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$shrink_target'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+my $log_offset = -s $node->logfile;
+
+is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 'f',
+ "shrink returns false when a buffer to be evicted is pinned");
+ok($node->log_contains(qr/could not remove buffer $pinned_buf, it is pinned/, $log_offset),
+ "log reports the pinned buffer that blocked eviction");
+ok($node->log_contains(qr/failed to evict extra buffers during shrinking/, $log_offset),
+ "log reports the eviction failure");
+
+is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers|$initial_nbuffers|$initial_nbuffers|0",
+ "pool unchanged after eviction failure");
+
+is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$initial_nbuffers (pending: $shrink_target)",
+ "pg_settings reports pending shrink target");
+
+# Releasing all pins lets the retry succeed.
+$pinner->quit;
+
+is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't',
+ "shrink succeeds after pins released");
+
+is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$shrink_target|$shrink_target|$shrink_target|0",
+ "pool shrunk to $shrink_target after pins released");
+
+# ---------------------------------------------------------------------------
+# Test the case when memory allocation fails when expanding the buffer pool.
+# Uses an injection point to simulate the failure without exhausting real
+# memory.
+# ---------------------------------------------------------------------------
+
+SKIP:
+{
+ skip "injection points not supported by this build"
+ unless $have_injection_points;
+
+ # The buffer manager's resizable structures whose sizes must be rolled
+ # back if any one of them fails to grow.
+ my $resizable_structs =
+ q{('Buffer Descriptors', 'Buffer Blocks', 'Buffer IO Condition Variables', 'Checkpoint BufferIds')};
+ my $sizes_query = "SELECT name, size FROM pg_shmem_allocations WHERE name IN $resizable_structs ORDER BY name";
+ my $sizes_before = $node->safe_psql('postgres', $sizes_query);
+
+ my $resizer = $node->background_psql('postgres');
+ $resizer->query_safe("SELECT injection_points_set_local()", verbose => 0);
+ $resizer->query_safe("SELECT injection_points_attach('buffer-mgr-resize-struct-fail', 'notice')",
+ verbose => 0);
+
+ my $expand_target = $max_nbuffers;
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$expand_target'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ my $expand_log_offset = -s $node->logfile;
+
+ is($resizer->query("SELECT pg_resize_shared_buffers()"), 'f',
+ "expansion fails when a structure can not be expanded");
+
+ # Discard the expected WARNINGs so later query_safe calls do not die.
+ $resizer->{stderr} = '';
+
+ ok($node->log_contains(qr/failed to expand buffer pool structures/, $expand_log_offset),
+ "log reports the expansion failure");
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$shrink_target|$shrink_target|$shrink_target|0",
+ "buffer pool status after expansion failure");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$shrink_target (pending: $expand_target)",
+ "pg_settings reports pending expand target");
+
+ is($node->safe_psql('postgres', $sizes_query), $sizes_before,
+ "resizable buffer manager structures rolled back to previous sizes");
+
+ # Detach the injection point, to retry again. The retry should succeed.
+ $resizer->query_safe(
+ "SELECT injection_points_detach('buffer-mgr-resize-struct-fail')",
+ verbose => 0);
+ is($resizer->query("SELECT pg_resize_shared_buffers()"), 't',
+ "expand succeeds after the injection point is detached");
+
+ $resizer->quit;
+
+ is($node->safe_psql('postgres', "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$expand_target|$expand_target|$expand_target|0",
+ "pool expanded to $expand_target after detach");
+}
+
+done_testing();
diff --git a/src/test/buffermgr/t/006_resize_with_syslogger.pl b/src/test/buffermgr/t/006_resize_with_syslogger.pl
new file mode 100644
index 00000000000..75b047ad984
--- /dev/null
+++ b/src/test/buffermgr/t/006_resize_with_syslogger.pl
@@ -0,0 +1,58 @@
+# Copyright (c) 2026-2026, PostgreSQL Global Development Group
+#
+# Test that pg_resize_shared_buffers() works when a backend that never
+# attaches to shared memory is running.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $initial_nbuffers = 16;
+my $expanded_nbuffers = 24;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# When logging_collector is on, the server starts a syslogger process that never
+# attaches to the shared memory. We use that as a proxy for a backend that never
+# attaches to shared memory.
+$node->append_conf(
+ 'postgresql.conf', qq{
+shared_buffers = $initial_nbuffers
+max_shared_buffers = $expanded_nbuffers
+logging_collector = on
+});
+$node->start;
+
+# Check that the syslogger is running by writing a log marker and waiting for it
+# to appear in the log file.
+sub check_syslogger_running
+{
+ my ($marker) = @_;
+
+ $node->safe_psql('postgres', "DO \$\$ BEGIN RAISE LOG '$marker'; END \$\$");
+ return $node->poll_query_until('postgres', "SELECT pg_read_file(pg_current_logfile()) ~ '$marker'");
+}
+
+check_syslogger_running('syslogger_marker_before_resize')
+ or die "syslogger is not running";
+
+# Resize the buffer pool, and check that the syslogger continues to run while
+# the resize is in progress.
+# TODO: Instead of custom markers we could use the log line that is emitted when
+# the resize is complete, when we have frozen those.
+for my $dir (['expand', $expanded_nbuffers], ['shrink', $initial_nbuffers])
+{
+ my ($name, $target) = @$dir;
+
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$target'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't',
+ "$name to $target succeeds with syslogger running");
+ ok(check_syslogger_running("syslogger_marker_after_$name"),
+ "syslogger drains logs after $name");
+}
+
+done_testing();
diff --git a/src/test/meson.build b/src/test/meson.build
index cd45cbf57fb..e9550933063 100644
--- a/src/test/meson.build
+++ b/src/test/meson.build
@@ -4,6 +4,7 @@ subdir('regress')
subdir('isolation')
subdir('authentication')
+subdir('buffermgr')
subdir('postmaster')
subdir('recovery')
subdir('subscription')
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index 699334320d9..e02c5314f4b 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -61,6 +61,7 @@ use Config;
use IPC::Run;
use PostgreSQL::Test::Utils qw(pump_until);
use Test::More;
+use Time::HiRes qw(usleep);
=pod
@@ -403,4 +404,79 @@ sub set_query_timer_restart
return $self->{query_timer_restart};
}
+=pod
+
+=item $session->poll_query_until($query [, $expected ])
+
+Run B<$query> repeatedly in this background session, until it returns the
+B<$expected> result ('t', or SQL boolean true, by default).
+Continues polling if the query returns an error result.
+Times out after a reasonable number of attempts.
+Returns 1 if successful, 0 if timed out.
+
+=cut
+
+sub poll_query_until
+{
+ my ($self, $query, $expected, %params) = @_;
+
+ $expected = 't' unless defined($expected); # default value
+
+ my $max_attempts = 10 * $PostgreSQL::Test::Utils::timeout_default;
+ my $attempts = 0;
+ my ($stdout, $stderr_flag);
+
+ while ($attempts < $max_attempts)
+ {
+ ($stdout, $stderr_flag) = $self->query($query, %params);
+
+ chomp($stdout);
+
+ # If query succeeded and returned expected result
+ if (!$stderr_flag && $stdout eq $expected)
+ {
+ return 1;
+ }
+
+ # Wait 0.1 second before retrying.
+ usleep(100_000);
+
+ $attempts++;
+ }
+
+ # Give up. Print the output from the last attempt, hopefully that's useful
+ # for debugging.
+ my $stderr_output = $stderr_flag ? $self->{stderr} : '';
+ diag qq(poll_query_until timed out executing this query:
+$query
+expecting this output:
+$expected
+last actual query output:
+$stdout
+with stderr:
+$stderr_output);
+ return 0;
+}
+
+=item $session->wait_for_event(backend_type, wait_event_name)
+
+Poll pg_stat_activity until backend_type reaches wait_event_name using this
+background session.
+
+=cut
+
+sub wait_for_event
+{
+ my ($self, $backend_type, $wait_event_name, %params) = @_;
+
+ $self->poll_query_until(qq[
+ SELECT count(*) > 0 FROM pg_stat_activity
+ WHERE backend_type = '$backend_type' AND wait_event = '$wait_event_name'
+ ], undef, %params)
+ or die
+ qq(timed out when waiting for $backend_type to reach wait event '$wait_event_name');
+
+ return;
+}
+
1;
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 65f0fa4d5cb..c0323d7f170 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -359,6 +359,7 @@ BufferAccessStrategy
BufferAccessStrategyType
BufferCacheOsPagesContext
BufferCacheOsPagesRec
+BufferControlBlock
BufferDesc
BufferDescPadded
BufferHeapTupleTableSlot
--
2.34.1
[text/x-patch] v20260724-0005-PID-of-the-backend-process-backing-the-Bac.patch (3.0K, ../../CAExHW5tYUYY13b4S1fv-zdxNsEEJzTy1EW6mRUj+FFEpA4Q0=g@mail.gmail.com/3-v20260724-0005-PID-of-the-backend-process-backing-the-Bac.patch)
download | inline diff:
From c13580f8e36583fa63230ab5fbbbdc1f61324aba Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 23 Jun 2026 09:19:06 +0530
Subject: [PATCH v20260724 5/7] PID of the backend process backing the
BackgroundPsql session
Many tests which use background psql sessions also fetch the pid of the backend
process as a separate step. They end up maintaing a separate variable for the
pid and pass it to the subroutines along with the BackgroundPsql object when
needed. Further the PID can not be fetched if the backend is blocked in a query
or an injection point or if the backend has died. This can be avoided by
proactively fetching the PID of the backend process when the BackgroundPsql is
started and storing it in the BackgroundPsql object. The PID can then be fetched
from the BackgroundPsql object when needed.
The pid is reset when the BackgroundPsql session is finished. Many of the tests
which use BackgroundPsql explicitly call {run}->finish to finish the session
when the backend is gone. They will have stale PID in the BackgroundPsql object.
The commit introduces a finish method in BackgroundPsql which will finish the
session and reset the PID.
Note to the reviewer:
This commit adds code to set the PID proactively and also the finish method.
These will be used in a future buffer resize test. But I have not modified the
existing tests to use the new capabilities. If we find this change useful, we
can modify the existing tests to use the new capabilities before committing this
change.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 29 ++++++++++++++++++-
1 file changed, 28 insertions(+), 1 deletion(-)
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index d7797225451..699334320d9 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -173,6 +173,14 @@ sub wait_connect
$self->{stderr} = '';
die "psql startup timed out" if $self->{timeout}->is_expired;
+
+ # Many tests which use background psql sessions also fetch the pid of the
+ # backend process, so we capture it here. The Callers that need the pid
+ # after a blocking query or after the backend has died can read it from
+ # $self->{backend_pid}.
+ my $pid = $self->query('SELECT pg_backend_pid()', verbose => 0);
+ chomp $pid;
+ $self->{backend_pid} = $pid;
}
=pod
@@ -190,6 +198,25 @@ sub quit
$self->{stdin} .= "\\q\n";
+ return $self->finish;
+}
+
+=pod
+
+=item $session->finish
+
+Reap the underlying IPC::Run handle without sending \q. Intended for
+sessions whose psql process has already exited (e.g. after the server
+terminated the backend or the client connection was killed).
+
+=cut
+
+sub finish
+{
+ my ($self) = @_;
+
+ $self->{backend_pid} = undef;
+
return $self->{run}->finish;
}
@@ -212,7 +239,7 @@ sub reconnect_and_clear
{
$self->{stdin} .= "\\q\n";
}
- $self->{run}->finish;
+ $self->finish;
# restart
$self->{run}->run();
--
2.34.1
[text/x-patch] v20260724-0003-Add-a-view-to-read-contents-of-shared-buff.patch (14.6K, ../../CAExHW5tYUYY13b4S1fv-zdxNsEEJzTy1EW6mRUj+FFEpA4Q0=g@mail.gmail.com/4-v20260724-0003-Add-a-view-to-read-contents-of-shared-buff.patch)
download | inline diff:
From 2b67f25aaa6a80664acc1fc3626e919e3a591139 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 25 Aug 2025 19:23:50 +0530
Subject: [PATCH v20260724 3/7] Add a view to read contents of shared buffer
lookup table
The view exposes the contents of the shared buffer lookup table for
debugging, testing and investigation.
This helped me in debugging issues where the buffer descriptor array and
buffer lookup table were out of sync; either the buffer lookup table had
a mapping page->buffer which wasn't present in the buffer descriptor
array or a page in the buffer descriptor array didn't have corresponding
entry in the buffer lookup table. pg_buffercache doesn't help with those
kind of issues. Also doing that under the debugger in very painful.
I intend to keep this patch while the rest of the code matures. If it is
found useful as a debugging tool, we may consider make it committable
and commit it.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../expected/pg_buffercache.out | 39 ++++++++
.../pg_buffercache--1.5--1.6.sql | 24 +++++
contrib/pg_buffercache/pg_buffercache_pages.c | 17 ++++
contrib/pg_buffercache/sql/pg_buffercache.sql | 20 +++++
doc/src/sgml/system-views.sgml | 89 +++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 58 ++++++++++++
src/include/storage/buf_internals.h | 2 +
7 files changed, 249 insertions(+)
diff --git a/contrib/pg_buffercache/expected/pg_buffercache.out b/contrib/pg_buffercache/expected/pg_buffercache.out
index c52a8491ff9..1ce5a970334 100644
--- a/contrib/pg_buffercache/expected/pg_buffercache.out
+++ b/contrib/pg_buffercache/expected/pg_buffercache.out
@@ -33,6 +33,26 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
t
(1 row)
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -46,6 +66,10 @@ SELECT * FROM pg_buffercache_summary();
ERROR: permission denied for function pg_buffercache_summary
SELECT * FROM pg_buffercache_usage_counts();
ERROR: permission denied for function pg_buffercache_usage_counts
+SELECT * FROM pg_buffercache_lookup_table_entries();
+ERROR: permission denied for function pg_buffercache_lookup_table_entries
+SELECT * FROM pg_buffercache_lookup_table;
+ERROR: permission denied for view pg_buffercache_lookup_table
RESET role;
-- Check that pg_monitor is allowed to query view / function
SET ROLE pg_monitor;
@@ -81,6 +105,21 @@ FROM pg_buffercache_pages() AS p
LIMIT 1;
ERROR: function return row and query-specified return row do not match
DETAIL: Returned type boolean at ordinal position 7, but query expects text.
+RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
RESET role;
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
index 458f054a691..9bf58567878 100644
--- a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
+++ b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
@@ -44,3 +44,27 @@ CREATE FUNCTION pg_buffercache_evict_all(
OUT buffers_skipped int4)
AS 'MODULE_PATHNAME', 'pg_buffercache_evict_all'
LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Add the buffer lookup table function
+CREATE FUNCTION pg_buffercache_lookup_table_entries(
+ OUT tablespace oid,
+ OUT database oid,
+ OUT relfilenode oid,
+ OUT forknum int2,
+ OUT blocknum int8,
+ OUT bufferid int4)
+RETURNS SETOF record
+AS 'MODULE_PATHNAME', 'pg_buffercache_lookup_table_entries'
+LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Create a view for convenient access.
+CREATE VIEW pg_buffercache_lookup_table AS
+ SELECT * FROM pg_buffercache_lookup_table_entries();
+
+-- Don't want these to be available to public.
+REVOKE ALL ON FUNCTION pg_buffercache_lookup_table_entries() FROM PUBLIC;
+REVOKE ALL ON pg_buffercache_lookup_table FROM PUBLIC;
+
+-- Grant access to monitoring role.
+GRANT EXECUTE ON FUNCTION pg_buffercache_lookup_table_entries() TO pg_read_all_stats;
+GRANT SELECT ON pg_buffercache_lookup_table TO pg_read_all_stats;
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 510455998aa..9512f1efa2f 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -77,6 +77,7 @@ PG_FUNCTION_INFO_V1(pg_buffercache_evict_all);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_relation);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_all);
+PG_FUNCTION_INFO_V1(pg_buffercache_lookup_table_entries);
/* Only need to touch memory once per backend process lifetime */
@@ -922,3 +923,19 @@ pg_buffercache_mark_dirty_all(PG_FUNCTION_ARGS)
PG_RETURN_DATUM(result);
}
+
+/*
+ * Return lookup table content as a set of records.
+ */
+Datum
+pg_buffercache_lookup_table_entries(PG_FUNCTION_ARGS)
+{
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* Fill the tuplestore */
+ BufTableGetContents(rsinfo->setResult, rsinfo->setDesc);
+
+ return (Datum) 0;
+}
diff --git a/contrib/pg_buffercache/sql/pg_buffercache.sql b/contrib/pg_buffercache/sql/pg_buffercache.sql
index be89b5f5a3a..d0726b2b2f9 100644
--- a/contrib/pg_buffercache/sql/pg_buffercache.sql
+++ b/contrib/pg_buffercache/sql/pg_buffercache.sql
@@ -18,6 +18,18 @@ from pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -26,6 +38,8 @@ SELECT * FROM pg_buffercache_os_pages;
SELECT * FROM pg_buffercache_pages() AS p (wrong int);
SELECT * FROM pg_buffercache_summary();
SELECT * FROM pg_buffercache_usage_counts();
+SELECT * FROM pg_buffercache_lookup_table_entries();
+SELECT * FROM pg_buffercache_lookup_table;
RESET role;
-- Check that pg_monitor is allowed to query view / function
@@ -42,6 +56,12 @@ FROM pg_buffercache_pages() AS p
LIMIT 1;
RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+RESET role;
+
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 5ea19d68622..6b905498337 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -71,6 +71,11 @@
<entry>backend memory contexts</entry>
</row>
+ <row>
+ <entry><link linkend="view-pg-buffer-lookup-table"><structname>pg_buffer_lookup_table</structname></link></entry>
+ <entry>shared buffer lookup table</entry>
+ </row>
+
<row>
<entry><link linkend="view-pg-config"><structname>pg_config</structname></link></entry>
<entry>compile-time configuration parameters</entry>
@@ -929,6 +934,90 @@ AND c1.path[c2.level] = c2.path[c2.level];
</para>
</sect1>
+ <sect1 id="view-pg-buffer-lookup-table">
+ <title><structname>pg_buffer_lookup_table</structname></title>
+ <indexterm>
+ <primary>pg_buffer_lookup_table</primary>
+ </indexterm>
+ <para>
+ The <structname>pg_buffer_lookup_table</structname> view exposes the current
+ contents of the shared buffer lookup table. Each row represents an entry in
+ the lookup table mapping a relation page to the ID of buffer in which it is
+ cached. The shared buffer lookup table is locked for a short duration while
+ reading so as to ensure consistency. This may affect performance if this view
+ is queried very frequently.
+ </para>
+ <table id="pg-buffer-lookup-table-view" xreflabel="pg_buffer_lookup_table">
+ <title><structname>pg_buffer_lookup_table</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>tablespace</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the tablespace containing the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>database</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database containing the relation (zero for shared relations)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relfilenode</structfield> <type>oid</type>
+ </para>
+ <para>
+ relfilenode identifying the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>forknum</structfield> <type>int2</type>
+ </para>
+ <para>
+ Fork number within the relation (see <xref linkend="storage-file-layout"/>)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>blocknum</structfield> <type>int8</type>
+ </para>
+ <para>
+ Block number within the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>bufferid</structfield> <type>int4</type>
+ </para>
+ <para>
+ ID of the buffer caching the page
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ <para>
+ Access to this view is restricted to members of the
+ <literal>pg_read_all_stats</literal> role by default.
+ </para>
+ </sect1>
+
<sect1 id="view-pg-config">
<title><structname>pg_config</structname></title>
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 5c8ccbee13f..c82b71deaa6 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -21,7 +21,12 @@
*/
#include "postgres.h"
+#include "fmgr.h"
+#include "funcapi.h"
#include "storage/buf_internals.h"
+#include "utils/rel.h"
+#include "utils/builtins.h"
+#include "storage/lwlock.h"
#include "storage/subsystems.h"
/* entry for buffer lookup hashtable */
@@ -167,3 +172,56 @@ BufTableDelete(BufferTag *tagPtr, uint32 hashcode)
if (!result) /* shouldn't happen */
elog(ERROR, "shared buffer hash table corrupted");
}
+
+/*
+ * BufTableGetContents
+ * Fill the given tuplestore with contents of the shared buffer lookup table
+ *
+ * This function is used by pg_buffercache extension to expose buffer lookup
+ * table contents via SQL. The caller is responsible for setting up the
+ * tuplestore and result set info.
+ */
+void
+BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc)
+{
+/* Expected number of attributes of the buffer lookup table entry. */
+#define BUFTABLE_CONTENTS_COLS 6
+
+ HASH_SEQ_STATUS hstat;
+ BufferLookupEnt *ent;
+ Datum values[BUFTABLE_CONTENTS_COLS];
+ bool nulls[BUFTABLE_CONTENTS_COLS];
+ int i;
+
+ memset(nulls, 0, sizeof(nulls));
+
+ Assert(tupdesc->natts == BUFTABLE_CONTENTS_COLS);
+
+ /*
+ * Lock all buffer mapping partitions to ensure a consistent view of the
+ * hash table during the scan. Must grab LWLocks in partition-number order
+ * to avoid LWLock deadlock.
+ */
+ for (i = 0; i < NUM_BUFFER_PARTITIONS; i++)
+ LWLockAcquire(BufMappingPartitionLockByIndex(i), LW_SHARED);
+
+ hash_seq_init(&hstat, SharedBufHash);
+ while ((ent = (BufferLookupEnt *) hash_seq_search(&hstat)) != NULL)
+ {
+ values[0] = ObjectIdGetDatum(ent->key.spcOid);
+ values[1] = ObjectIdGetDatum(ent->key.dbOid);
+ values[2] = ObjectIdGetDatum(ent->key.relNumber);
+ values[3] = ObjectIdGetDatum(ent->key.forkNum);
+ values[4] = Int64GetDatum(ent->key.blockNum);
+ values[5] = Int32GetDatum(ent->id);
+
+ tuplestore_putvalues(tupstore, tupdesc, values, nulls);
+ }
+
+ /*
+ * Release all buffer mapping partition locks in the reverse order so as
+ * to avoid LWLock deadlock.
+ */
+ for (i = NUM_BUFFER_PARTITIONS - 1; i >= 0; i--)
+ LWLockRelease(BufMappingPartitionLockByIndex(i));
+}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index a7606c6e92b..678065b3de2 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -29,6 +29,7 @@
#include "storage/spin.h"
#include "utils/relcache.h"
#include "utils/resowner.h"
+#include "utils/tuplestore.h"
/*
* Buffer state is a single 64-bit variable where following data is combined.
@@ -599,6 +600,7 @@ extern uint32 BufTableHashCode(BufferTag *tagPtr);
extern int BufTableLookup(BufferTag *tagPtr, uint32 hashcode);
extern int BufTableInsert(BufferTag *tagPtr, uint32 hashcode, int buf_id);
extern void BufTableDelete(BufferTag *tagPtr, uint32 hashcode);
+extern void BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc);
/* localbuf.c */
extern bool PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount);
--
2.34.1
[text/x-patch] v20260724-0006-Resizable-shared-memory-structures.patch (129.2K, ../../CAExHW5tYUYY13b4S1fv-zdxNsEEJzTy1EW6mRUj+FFEpA4Q0=g@mail.gmail.com/5-v20260724-0006-Resizable-shared-memory-structures.patch)
download | inline diff:
From 9f6afda743bf6d574b7f51a432e787f9f0f6911f Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 17 Feb 2026 16:51:20 +0530
Subject: [PATCH v20260724 6/7] Resizable shared memory structures
Resizable shared memory structures can be requested by specifying a new
member ShmemStructOpts::maximum_size. At the startup or when the
structure is created, we reserve address space worth maximum_size in the
shared memory segment. It is expected that the subsystem which creates a
resizable structure would initialize only the memory worth its initial
size given by ShmemStructOpts::maximum_size when creating it. In an
mmap'ed memory, this should allocate memory worth only the initial size.
It should not allocate maximum_size worth of memory initially. As the
structure is resized using ShmemResizeStruct() memory is freed or
allocated in chunks of memory pages when shrinking and expanding the
structure respectively. Optional ShmemStructOpts::minimum_size
specification for resizable shared memory structures allows us to
enforce that a resizable structure cannot be shrunk below a certain
size. If minimum_size is not specified for a resizable structure, it the
minimum size defaults to 0. Two additional columns are added to
pg_shmem_allocations view to report minimum and maximum size
respectively.
The structure for which ShmemStructOpts::maximum_size is not specified
or is set to 0 is considered as a fixed size structure. Existing calls
to ShmemRequestStruct() requsting fixed size structures need no change.
For fixed-size structures, the minimum size and maximum size are set to
the initial size specified in ShmemRequestStructOpts::size. This makes
maximum size and minimum size reported in pg_shmem_allocations view
semantically consistent for both fixed-size and resizable structures. As
a side effect, a request which specifies same minimum_size and
maximum_size will be treated as a fixed size structure.
Minimum, maximum and initial sizes of main shared memory area
=============================================================
With the addition of resizable shared structures, the main shared memory
area has three sizes to be tracked:
- a. the minimum amount of shared memory that will always be needed
- b. the maximum amount of shared memory that may be allocated when all
the resizable structures have grown to their maximum size. It is also
is the size of the address space reserved for the main shared memory
area.
- c. the amount of memory allocated at the server startup i.e. initial
allocation.
mmap needs to use the maximum size to reserve enough address space to
accomodate all the shared structures. But a DBA may provision only the
memory required at the startup initially and increase or descrease the
memory provision as the demand changes.
These three sizes are reported as GUCs shared_memory_minimum_size,
shared_memory_maximum_size and shared_memory_initial_size respectively.
They replace the old GUC shared_memory_size which used to report both
size of the main shared memory area and amount of memory required in it.
Since these two things are not the same anymore, a single GUC is not
sufficient. These GUCs help DBAs to estimate memory to be provisioned
at the startup and during run time as the resizable structures change
their sizes.
Portability
===========
Resizable shared structures feature depends upon existence of function
madvise() and constants MADV_REMOVE and MADV_WRITE_POPULATE. On the
platforms which do not have these or the shared memory types (e.g.
Sys-V) which do not provide ability to free or allocate parts of shared
memory, we disable this feature. A run time GUC have_resizable_shmem
indicates whether a running server supports resizable structures or not.
The commit introduces a compile time flag HAVE_RESIZABLE_SHMEM which is
defined if MADV_REMOVE and MADV_WRITE_POPULATE exist. We don't check
existence of madvise separately, since existence of the constants
implies existence of the function. HAVE_RESIZABLE_SHMEM is not defined
in EXEC_BACKEND builds since that's largely used for Windows where the
APIs to free and allocate memory from and to a given address space are
not known to the author right now. Given that PostgreSQL is used widely
on Linux (with shared memory type mmap), providing this feature on Linux
benefits most of its users. Once we figure out the required Windows
APIs, we will support this feature on Windows as well.
Prohibiting access to the unused portion of a resizable structure
================================================================
We provide ShmemProtectStruct() to add protections on the address space
reserved for a resizable shared memory structure so that the address
space upto its current size is accessible whereas the part beyond that
is inaccessible. Since these protections are backend specific, the
subsystem using the resizable shared memory structures has to make sure
to call ShmemProtectStruct() after every resize before any backend tries
to access the portion of the structure between old and the new size.
This can coordinated using ProcSignalBarrier mechanism in a running
backend. When starting the server, Postmaster adds appropriate
protections when creating these structures. In an EXEC_BACKEND case,
when a backend starts, the protections are added according to the
current size of the backend (through attach_fn call). However, in
non-EXEC_BACKEND case, a new backend inherits stale protections from the
postmaster. It needs to apply the protections according to the current
sizes of these structures.
Following points need more discussion.
Discussion points
=================
adding initial protections in non-EXEC_BACKEND case
----------------------------------------------------------------------
When a backend starts, it needs to add shmem protections it in such a
way that a ProcSignalBarrier conveying the protection change is not
missed. Hence we do it along side InitLocalDataChecksumState(). This has
two problems 1. it delays backend startup process and 2. it loosely ties
resizable shared memory structure to ProcSignalBarrier mechanim. Do we
consider the complexity worth it? Is there any other way to do it
without these two drawbacks?
Further, we call ShmemReprotectResizableStructs() twice in the startup
sequence, once immediately after the place where
AttachSharedMemoryStructs() would be called in EXEC_BACKEND case and
then second time after ProcSignalInit(). The first one is needed so that
the processes can access the resizable structures safely. The second is
needed for the reasons mentioned there. But a change in the protection
during these calls will be missed by the backend and thus lead to
hazardous access. Can we avoid it? If ProcSignalInit() were to be called
nearer to the InitProcess(), it may make calling
ShmemReprotectResizeStructs() easier.
Protection in the resizing backend vs other backends
----------------------------------------------------
The patch provides an API to protect the address spaces not used because
of current size of the resizable structures in every backend. Since
every backend has to protect its own address space, a separate API is
required. But at the same time, in some cases the resize operation
itself needs to change the protection (e.g. when expanding the
structure). So there's some asymmetry between the moment when the
resizing backend adds protection and when other backends add it. That
doesn't affect end result since adding this protection is an idempotent
operation. But still there's some ugliness involved. This part also
needs more discussion.
HAVE_RESIZABLE_SHMEM and have_resizable_shmem
---------------------------------------------
The patch defines a build time macro HAVE_RESIZABLE_SHMEM to avoid
compilation failure on a build platform where the required system calls
or constants are not available. But we also need a run time check to see
whether the shared memory type being used allows resizing. Hence we
provide a GUC have_resizable_shmem which tells whether a running server
supports resizable shared memory or not. Most of the code which deals
with the shared memory is under HAVE_RESIZABLE_SHMEM, but there is some
related to resizable shared memory outside the macro as well. For
example, the members maximum_size, minimum_size are not under this
macro. This is mostly to allow writing readable code in both the modes
and still provide the same user visible views etc. This division of code
within and without macro needs more discussion.
TODOs
=====
There are some TODOs in the code still. We will address them as we
finalize that part of design/code. Comments/suggestions on those is
welcome.
Idea of using mmap to reserve address space in shared memory segments
and allocating memory within that address space on demand was proposed
by Dmitry Dolgov <9erthalion6@gmail.com>.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Reviewed-by: Matthias van de Meent <boekewurm+postgres@gmail.com>
---
configure.ac | 4 +
doc/src/sgml/config.sgml | 102 +++-
doc/src/sgml/ref/postgres-ref.sgml | 8 +-
doc/src/sgml/runtime.sgml | 9 +-
doc/src/sgml/system-views.sgml | 57 +-
doc/src/sgml/xfunc.sgml | 92 +++
meson.build | 16 +
src/backend/port/sysv_shmem.c | 259 +++++++-
src/backend/port/win32_shmem.c | 64 +-
src/backend/postmaster/auxprocess.c | 46 +-
src/backend/storage/ipc/ipci.c | 114 ++--
src/backend/storage/ipc/shmem.c | 555 ++++++++++++++++--
src/backend/storage/lmgr/proc.c | 16 +
src/backend/utils/init/postinit.c | 10 +
src/backend/utils/misc/guc_parameters.dat | 57 +-
src/backend/utils/misc/guc_tables.c | 15 +-
src/include/catalog/pg_proc.dat | 4 +-
src/include/pg_config.h.in | 8 +
src/include/pg_config_manual.h | 14 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_shmem.h | 6 +-
src/include/storage/shmem.h | 19 +
src/include/storage/shmem_internal.h | 2 +-
src/test/modules/test_shmem/meson.build | 3 +-
...mem_alloc.pl => 001_fixed_shmem_struct.pl} | 31 +
.../t/002_resizable_shmem_struct.pl | 382 ++++++++++++
.../modules/test_shmem/test_shmem--1.0.sql | 55 ++
src/test/modules/test_shmem/test_shmem.c | 504 +++++++++++++++-
src/test/regress/expected/rules.out | 7 +-
src/tools/pgindent/typedefs.list | 1 +
30 files changed, 2294 insertions(+), 168 deletions(-)
rename src/test/modules/test_shmem/t/{001_late_shmem_alloc.pl => 001_fixed_shmem_struct.pl} (58%)
create mode 100644 src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
diff --git a/configure.ac b/configure.ac
index a331749fcb5..bba12b8c8f4 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1912,6 +1912,10 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux-specific madvise constants needed for resizable shared memory. See similar checks in meson.build for explanation of why these checks are here.
+AC_CHECK_DECLS([MADV_POPULATE_WRITE], [], [], [#include <sys/mman.h>])
+AC_CHECK_DECLS([MADV_REMOVE], [], [], [#include <sys/mman.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index aa7b1bd75d2..5b6cc870093 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -12373,6 +12373,20 @@ dynamic_library_path = '/usr/local/lib/postgresql:$libdir'
</listitem>
</varlistentry>
+ <varlistentry id="guc-have-resizable-shmem" xreflabel="have_resizable_shmem">
+ <term><varname>have_resizable_shmem</varname> (<type>boolean</type>)
+ <indexterm>
+ <primary><varname>have_resizable_shmem</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports whether <productname>PostgreSQL</productname> supports <link
+ linkend="xfunc-shared-addin-resizable">Resizable shared memory structures</link>.
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry id="guc-huge-pages-status" xreflabel="huge_pages_status">
<term><varname>huge_pages_status</varname> (<type>enum</type>)
<indexterm>
@@ -12552,32 +12566,100 @@ dynamic_library_path = '/usr/local/lib/postgresql:$libdir'
</listitem>
</varlistentry>
- <varlistentry id="guc-shared-memory-size" xreflabel="shared_memory_size">
- <term><varname>shared_memory_size</varname> (<type>integer</type>)
+ <varlistentry id="guc-shared-memory-initial-size" xreflabel="shared_memory_initial_size">
+ <term><varname>shared_memory_initial_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_initial_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the amount of memory, rounded up to the nearest megabyte, allocated at the server startup in main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-minimum-size" xreflabel="shared_memory_minimum_size">
+ <term><varname>shared_memory_minimum_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_minimum_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the minimum amount of memory, rounded up to the nearest megabyte, required in the main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-maximum-size" xreflabel="shared_memory_maximum_size">
+ <term><varname>shared_memory_maximum_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_maximum_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the maximum size of the main shared memory area, rounded up
+ to the nearest megabyte. This is the amount of address space that
+ must be reserved for the main shared memory area. This is also the maximum amount of memory that may be allocated in the main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-initial-size-in-huge-pages" xreflabel="shared_memory_initial_size_in_huge_pages">
+ <term><varname>shared_memory_initial_size_in_huge_pages</varname> (<type>integer</type>)
<indexterm>
- <primary><varname>shared_memory_size</varname> configuration parameter</primary>
+ <primary><varname>shared_memory_initial_size_in_huge_pages</varname> configuration parameter</primary>
</indexterm>
</term>
<listitem>
<para>
- Reports the size of the main shared memory area, rounded up to the
- nearest megabyte.
+ Reports the number of huge pages needed in the main shared memory area
+ at the startup based on the specified <xref
+ linkend="guc-huge-page-size"/>. If huge pages are not supported, this
+ will be <literal>-1</literal>.
+ </para>
+ <para>
+ This setting is supported only on <productname>Linux</productname>. It
+ is always set to <literal>-1</literal> on other platforms. For more
+ details about using huge pages on <productname>Linux</productname>, see
+ <xref linkend="linux-huge-pages"/>.
</para>
</listitem>
</varlistentry>
- <varlistentry id="guc-shared-memory-size-in-huge-pages" xreflabel="shared_memory_size_in_huge_pages">
- <term><varname>shared_memory_size_in_huge_pages</varname> (<type>integer</type>)
+ <varlistentry id="guc-shared-memory-minimum-size-in-huge-pages" xreflabel="shared_memory_minimum_size_in_huge_pages">
+ <term><varname>shared_memory_minimum_size_in_huge_pages</varname> (<type>integer</type>)
<indexterm>
- <primary><varname>shared_memory_size_in_huge_pages</varname> configuration parameter</primary>
+ <primary><varname>shared_memory_minimum_size_in_huge_pages</varname> configuration parameter</primary>
</indexterm>
</term>
<listitem>
<para>
- Reports the number of huge pages that are needed for the main shared
- memory area based on the specified <xref linkend="guc-huge-page-size"/>.
+ Reports the minimum number of huge pages needed in the main shared memory area based on the specified <xref linkend="guc-huge-page-size"/>.
If huge pages are not supported, this will be <literal>-1</literal>.
</para>
+ <para>
+ This setting is supported only on <productname>Linux</productname>. It
+ is always set to <literal>-1</literal> on other platforms.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-maximum-size-in-huge-pages" xreflabel="shared_memory_maximum_size_in_huge_pages">
+ <term><varname>shared_memory_maximum_size_in_huge_pages</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_maximum_size_in_huge_pages</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the maximum number of huge pages needed in the
+ main shared memory area based on the specified
+ <xref linkend="guc-huge-page-size"/>. If huge pages are not
+ supported, this will be <literal>-1</literal>.
+ </para>
<para>
This setting is supported only on <productname>Linux</productname>. It
is always set to <literal>-1</literal> on other platforms. For more
diff --git a/doc/src/sgml/ref/postgres-ref.sgml b/doc/src/sgml/ref/postgres-ref.sgml
index b13a16a117f..3eefd629e06 100644
--- a/doc/src/sgml/ref/postgres-ref.sgml
+++ b/doc/src/sgml/ref/postgres-ref.sgml
@@ -143,8 +143,12 @@ PostgreSQL documentation
<para>
This can be used on a running server for most parameters. However,
the server must be shut down for some runtime-computed parameters
- (e.g., <xref linkend="guc-shared-memory-size"/>,
- <xref linkend="guc-shared-memory-size-in-huge-pages"/>, and
+ (e.g., <xref linkend="guc-shared-memory-initial-size"/>,
+ <xref linkend="guc-shared-memory-minimum-size"/>,
+ <xref linkend="guc-shared-memory-maximum-size"/>,
+ <xref linkend="guc-shared-memory-initial-size-in-huge-pages"/>,
+ <xref linkend="guc-shared-memory-minimum-size-in-huge-pages"/>,
+ <xref linkend="guc-shared-memory-maximum-size-in-huge-pages"/>, and
<xref linkend="guc-wal-segment-size"/>).
</para>
diff --git a/doc/src/sgml/runtime.sgml b/doc/src/sgml/runtime.sgml
index d9984910cc4..8ba4d233509 100644
--- a/doc/src/sgml/runtime.sgml
+++ b/doc/src/sgml/runtime.sgml
@@ -1450,11 +1450,12 @@ export PG_OOM_ADJUST_VALUE=0
<varname>CONFIG_HUGETLB_PAGE=y</varname>. You will also have to configure
the operating system to provide enough huge pages of the desired size.
The runtime-computed parameter
- <xref linkend="guc-shared-memory-size-in-huge-pages"/> reports the number
- of huge pages required. This parameter can be viewed before starting the
+ <xref linkend="guc-shared-memory-maximum-size-in-huge-pages"/> reports the
+ number of huge pages that must be available in the pool for postmaster
+ startup to succeed. This parameter can be viewed before starting the
server with a <command>postgres</command> command like:
<programlisting>
-$ <userinput>postgres -D $PGDATA -C shared_memory_size_in_huge_pages</userinput>
+$ <userinput>postgres -D $PGDATA -C shared_memory_maximum_size_in_huge_pages</userinput>
3170
$ <userinput>grep ^Hugepagesize /proc/meminfo</userinput>
Hugepagesize: 2048 kB
@@ -1465,7 +1466,7 @@ hugepages-1048576kB hugepages-2048kB
In this example the default is 2MB, but you can also explicitly request
either 2MB or 1GB with <xref linkend="guc-huge-page-size"/> to adapt
the number of pages calculated by
- <varname>shared_memory_size_in_huge_pages</varname>.
+ <varname>shared_memory_maximum_size_in_huge_pages</varname>.
While we need at least <literal>3170</literal> huge pages in this example,
a larger setting would be appropriate if other programs on the machine
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 6b905498337..7c4f4aeba14 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -4333,8 +4333,46 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
Size of the allocation in bytes including padding. For anonymous
allocations, no information about padding is available, so the
<literal>size</literal> and <literal>allocated_size</literal> columns
- will always be equal. Padding is not meaningful for free memory, so
- the columns will be equal in that case also.
+ will always be equal. Padding is not meaningful for free memory, so the
+ columns will be equal in that case also. For resizable allocations which
+ may span multiple memory pages, the padding includes the padding due to
+ page alignment.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>minimum_size</structfield> <type>int8</type>
+ </para>
+ <para>
+ Minimum size in bytes that the resizable allocation can shrink to. Equals
+ <structfield>size</structfield> for fixed-size allocations, anonymous
+ allocations, and free memory.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>maximum_size</structfield> <type>int8</type>
+ </para>
+ <para>
+ Maximum size in bytes that the resizable allocation can grow to. Equals
+ <structfield>size</structfield> for fixed-size allocations, anonymous
+ allocations, and free memory.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>reserved_space</structfield> <type>int8</type>
+ </para>
+ <para>
+ Address space reserved for the allocation in bytes. For resizable
+ structures, this is the total address space reserved to accommodate
+ growth up to <structfield>maximum_size</structfield>, and is greater
+ than or equal to <structfield>allocated_size</structfield>. For
+ fixed-size allocations, anonymous allocations, and free memory this
+ is same as <structfield>allocated_size</structfield>.
</para></entry>
</row>
</tbody>
@@ -4348,6 +4386,21 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
<literal>ShmemRequestHash()</literal>.
</para>
+ <para>
+ Resizable structures are allocations whose size can change while the server
+ is running. <xref linkend="guc-have-resizable-shmem"/> indicates whether the
+ running server supports resizable allocations. Each such allocation has a
+ minimum size give by <structfield>minimum_size</structfield> it can shrink to
+ and a maximum size given by <structfield>maximum_size</structfield> it can
+ grow to; enough address space is reserved up front to accommodate its
+ <structfield>maximum_size</structfield>. The aggregate initial, minimum and
+ maximum sizes of all shared memory allocations are reported by <xref
+ linkend="guc-shared-memory-initial-size"/>, <xref
+ linkend="guc-shared-memory-minimum-size"/> and <xref
+ linkend="guc-shared-memory-maximum-size"/> respectively (and their
+ <literal>_in_huge_pages</literal> counterparts).
+ </para>
+
<para>
By default, the <structname>pg_shmem_allocations</structname> view can be
read only by superusers or roles with privileges of the
diff --git a/doc/src/sgml/xfunc.sgml b/doc/src/sgml/xfunc.sgml
index 2b8a11e7ad0..2bfc664d7ca 100644
--- a/doc/src/sgml/xfunc.sgml
+++ b/doc/src/sgml/xfunc.sgml
@@ -3744,6 +3744,98 @@ my_shmem_init(void *arg)
</para>
</sect3>
+ <sect3 id="xfunc-shared-addin-resizable">
+ <title>Resizable shared memory structures</title>
+
+ <para>
+ A resizable memory structure can be requested using
+ <function>ShmemRequestStruct</function> by passing
+ <parameter>.maximum_size</parameter> along with
+ <parameter>.size</parameter>. <parameter>.maximum_size</parameter> is
+ maximum size upto which the structure can grow where as
+ <parameter>.size</parameter> is the initial size of the structure.
+ Optionally, <parameter>.minimum_size</parameter> can be set to the minimum
+ size that the structure can shrink to. While
+ contiguous address space worth <parameter>maximum_size</parameter> is
+ allocated to the structure, only memory worth <parameter>size</parameter>
+ bytes is allocated initially. The <function>init_fn</function> should only
+ initialize the <parameter>size</parameter> amount of memory. The actual
+ memory allocated to this structure at any point in time is given by <link
+ linkend="view-pg-shmem-allocations"><structname>pg_shmem_allocations</structname>.<structfield>allocated_size</structfield></link>
+ and the address space reserved for this structure is given by <link
+ linkend="view-pg-shmem-allocations"><structname>pg_shmem_allocations</structname>.<structfield>reserved_space</structfield></link>.
+ </para>
+
+ <para>
+ The structure can be resized using <function>ShmemResizeStruct</function> by
+ passing it the structure's name and the
+ new size which can be anywhere between <parameter>minimum_size</parameter>
+ and <parameter>maximum_size</parameter>. If the new size is smaller than the
+ current size of the structure, the memory between the new size and current
+ size is freed while keeping the contents of the memory upto new size intact.
+ If the new size is greater than the current size, memory is allocated upto
+ new size while keeping the current contents of the structure intact. The
+ starting address of the structure does not change because of resizing
+ operation.
+ </para>
+
+ <para>
+ <function>ShmemResizeStruct</function> returns <literal>true</literal> on
+ success. If the operating system cannot supply the additional memory
+ required to expand the structure, it emits a <literal>WARNING</literal>
+ naming the structure and the number of bytes that could not be allocated,
+ leaves the structure at its previous size, and returns
+ <literal>false</literal>.
+ </para>
+
+ <para>
+ <function>ShmemResizeStruct</function> does not coordinate with other
+ backends that may be accessing the shared structure. The caller is
+ responsible for ensuring that no backend is accessing the parts of the
+ structure between old size and the new size while
+ <function>ShmemResizeStruct</function> is executing. Similarly after the
+ function has finished, the caller must ensure that each backend knows the
+ new size before accessing the parts of the structure between the old size
+ and the new size. Accessing the range <literal>[current_size,
+ maximum_size)</literal> results in an undefined operating-system dependent
+ behaviour.
+ </para>
+
+ <para>
+ <function>ShmemProtectStruct</function> can be used to make the range beyond
+ the current size inaccessible in a backend's address space, so that stray
+ accesses cause a fault instead of silently succeeding. When a backend
+ starts, protections for all resizable structures are adjusted according to
+ their current sizes, so subsystems need not call
+ <function>ShmemProtectStruct</function> from their per-backend
+ initialization routines. During runtime after calling
+ <function>ShmemResizeStruct</function> from one backend,
+ <function>ShmemProtectStruct</function> should be called from all the
+ backends that may access the structure.
+ </para>
+
+ <para>
+ <productname>PostgreSQL</productname>'s
+ <function>ProcSignalBarrier</function> mechanism (see
+ <filename>src/include/storage/procsignal.h</filename>) may be used to
+ coordinate the resize across backends and adjusting access to the address
+ space.
+ </para>
+
+ <para>
+ This functionality is available only on the platforms which provide the APIs
+ necessary to reserve contiguous address space and to allocate or free memory
+ in that address space on demand. Macro <symbol>HAVE_RESIZABLE_SHMEM</symbol>
+ is defined on such platforms. It can be used to guard code related to
+ resizing a shared memory structure. The functionality is available on with
+ mmap'ed memory, so subsystems which use resizable structures may have to
+ addtionally disable resizable memory usage when <symbol>shared_memory_type</symbol> is not
+ <symbol>SHMEM_TYPE_MMAP</symbol>. A GUC <xref linkend="guc-have-resizable-shmem"/> is set to
+ <literal>on</literal> when this functionality is available in a running
+ server, <literal>off</literal> otherwise.
+ </para>
+ </sect3>
+
<sect3 id="xfunc-shared-addin-dynamic">
<title>Allocating Dynamic Shared Memory after Startup</title>
diff --git a/meson.build b/meson.build
index f4cde249242..4e0c736470c 100644
--- a/meson.build
+++ b/meson.build
@@ -2894,6 +2894,22 @@ decl_checks = [
['timingsafe_bcmp', 'string.h'],
]
+# Linux-specific madvise constants needed for resizable shared memory.
+# Usually we use AC_CHECK_DECLS to check for function declarations, but in this
+# case we are using it to detect existence of constants. These constants are
+# used to define HAVE_RESIZABLE_SHMEM which is used in storage/pg_shmem.h as
+# well as storage/shmem.h. The first abstracts the APIs to allocate shared
+# memory segments from the operating system whereas the second abstracts APIs to
+# allocate shared memory to various subsystems. Since they are related but
+# orthogonal to each other, including any one of them in the other file doesn't
+# make sense. pg_config_manual.h is the only place where HAVE_RESIZABLE_SHMEM
+# can be defined and made available to both without including sys/mman.h. But
+# for that we need constants that indicate the existence of following defines.
+decl_checks += [
+ ['MADV_POPULATE_WRITE', 'sys/mman.h'],
+ ['MADV_REMOVE', 'sys/mman.h'],
+]
+
# Need to check for function declarations for these functions, because
# checking for library symbols wouldn't handle deployment target
# restrictions on macOS
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 2e3886cf9fe..c052776e94c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -589,44 +589,113 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
return true;
}
+/*
+ * Get the page size being used by the shared memory.
+ *
+ * The function should be called only after the shared memory has been setup.
+ */
+size_t
+GetOSPageSize(void)
+{
+ size_t os_page_size;
+
+ Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
+
+ os_page_size = sysconf(_SC_PAGESIZE);
+
+ /* If huge pages are actually in use, use huge page size */
+ if (huge_pages_status == HUGE_PAGES_ON)
+ GetHugePageSize(&os_page_size, NULL);
+
+ return os_page_size;
+}
+
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * initial_size is the amount of memory required at the start of the server.
+ *
+ * *size is the size of the anonymous memory mapping to create. It will be
+ * updated to the actual size of the allocation, if it ends up allocating a
+ * segment that is larger than the requested. When *size > initial_size we want
+ * to reserve a large memory segment while allocating only initial_size amount of
+ * memory at the server start.
+ *
+ * When using huge pages, we make sure that there are enough huge pages
+ * configured to cover the initial_size worth of memory.
*/
static void *
-CreateAnonymousSegment(Size *size)
+CreateAnonymousSegment(Size initial_size, Size *size)
{
Size allocsize = *size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ bool resizable = (initial_size < allocsize);
int mmap_flags = MAP_SHARED | MAP_ANONYMOUS | MAP_HASSEMAPHORE;
+ Assert(initial_size <= allocsize);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
#else
if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
int huge_mmap_flags;
+ bool probe_ok = true;
+ /* Round up the request size to a suitable large value. */
GetHugePageSize(&hugepagesize, &huge_mmap_flags);
-
if (allocsize % hugepagesize != 0)
allocsize = add_size(allocsize, hugepagesize - (allocsize % hugepagesize));
+ if (initial_size % hugepagesize != 0)
+ initial_size = add_size(initial_size, hugepagesize - (initial_size % hugepagesize));
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | huge_mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ /*
+ * When the total amount of space requested is larger than the initial
+ * memory requirement, we do not allocate memory worth the entire
+ * requested space upfront. But we will need to make sure that there
+ * are enough huge pages configured to cover the initial memory
+ * requirement. Otherwise, we will choose huge pages map instead of
+ * falling back to normal pages and the server will not start. Hence
+ * we first try to map and allocate the initial memory requirement. If
+ * it fails we fall back to normal pages. If it succeeds, we unmap the
+ * memory and then try to map the entire requested space without
+ * allocating memory.
+ */
+ if (resizable)
+ {
+ void *probe;
+
+ probe = mmap(NULL, initial_size, PROT_READ | PROT_WRITE,
+ mmap_flags | huge_mmap_flags,
+ -1, 0);
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ probe_ok = false;
+ if (huge_pages == HUGE_PAGES_TRY)
+ elog(DEBUG1,
+ "huge-page probe mmap(%zu) failed, huge pages disabled: %m",
+ initial_size);
+ }
+ else if (munmap(probe, initial_size) != 0)
+ elog(ERROR, "munmap(%p, %zu) huge page probe failed: %m", probe, initial_size);
+ }
+
+ if (probe_ok)
+ {
+ if (resizable)
+ mmap_flags |= MAP_NORESERVE;
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ mmap_flags | huge_mmap_flags, -1, 0);
+ mmap_errno = errno;
+ if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
+ elog(DEBUG1,
+ "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ allocsize);
+ }
}
#endif
@@ -645,6 +714,10 @@ CreateAnonymousSegment(Size *size)
* to non-huge pages.
*/
allocsize = *size;
+
+ /* Use MAP_NORESERVE when we do not need all the memory up front. */
+ if (resizable)
+ mmap_flags |= MAP_NORESERVE;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
mmap_flags, -1, 0);
mmap_errno = errno;
@@ -693,13 +766,18 @@ AnonymousShmemDetach(int status, Datum arg)
* standard header. Also, register an on_shmem_exit callback to release
* the storage.
*
+ * initial_size is the amount of memory required at the start of the server.
+ * Used only when we want to reserve a large memory segment while allocating
+ * only initial_size amount of memory at the server start. That facility is only
+ * available when using anonymous shared memory segments.
+ *
* Dead Postgres segments pertinent to this DataDir are recycled if found, but
* we do not fail upon collision with foreign shmem segments. The idea here
* is to detect and re-use keys that may have been assigned by a crashed
* postmaster or backend.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(Size initial_size, Size size,
PGShmemHeader **shim)
{
IpcMemoryKey NextShmemSegID;
@@ -738,7 +816,7 @@ PGSharedMemoryCreate(Size size,
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
+ AnonymousShmem = CreateAnonymousSegment(initial_size, &size);
AnonymousShmemSize = size;
/* Register on-exit routine to unmap the anonymous segment */
@@ -754,6 +832,10 @@ PGSharedMemoryCreate(Size size,
/* huge pages are only available with mmap */
SetConfigOption("huge_pages_status", "off",
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ /* resizable shared memory is only available with mmap */
+ SetConfigOption("have_resizable_shmem", "off",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
}
/*
@@ -991,3 +1073,148 @@ PGSharedMemoryDetach(void)
AnonymousShmem = NULL;
}
}
+
+/*
+ * Make sure that the memory of given size from the given address is released.
+ *
+ * The address and size are expected to be page aligned.
+ *
+ * Only supported on platforms that support anonymous shared memory.
+ *
+ * On success returns true. false if system call fails.
+ */
+bool
+PGSharedMemoryEnsureFreed(void *addr, Size size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be freed"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(addr == (void *) TYPEALIGN(os_page_size, addr));
+ Assert(size == TYPEALIGN(os_page_size, size));
+ }
+#endif
+
+ if (madvise(addr, size, MADV_REMOVE) == -1)
+ {
+ ereport(WARNING, errmsg("could not free shared memory: %m"));
+ return false;
+ }
+
+ return true;
+#endif
+}
+
+/*
+ * Make sure that the memory of given size from the given address is allocated.
+ *
+ * The address and size are expected to be page aligned. The caller is
+ * responsible for ensuring that the range is already writable; this function
+ * only populates it.
+ *
+ * Only supported on platforms that support anonymous shared memory.
+ *
+ * On success returns true, false if system call fails.
+ */
+bool
+PGSharedMemoryEnsureAllocated(void *addr, Size size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be allocated at runtime"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(addr == (void *) TYPEALIGN(os_page_size, addr));
+ Assert(size == TYPEALIGN(os_page_size, size));
+ }
+#endif
+
+ if (madvise(addr, size, MADV_POPULATE_WRITE) == -1)
+ {
+ ereport(WARNING, errmsg("could not allocate shared memory: %m"));
+ return false;
+ }
+
+ return true;
+#endif
+}
+
+/*
+ * Set memory protection on the given region of shared memory.
+ *
+ * Makes [rw_start, rw_end) readable and writable, and [rw_end, prot_end)
+ * inaccessible.
+ *
+ * All addresses are expected to be page aligned.
+ *
+ * Only supported on platforms that support resizable shared memory.
+ *
+ * On success returns true, false if system call fails.
+ */
+bool
+PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be protected at runtime"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(rw_start == (void *) TYPEALIGN(os_page_size, rw_start));
+ Assert(rw_end == (void *) TYPEALIGN(os_page_size, rw_end));
+ Assert(prot_end == (void *) TYPEALIGN(os_page_size, prot_end));
+ }
+#endif
+ Assert(rw_end >= rw_start);
+
+ if (rw_end > rw_start)
+ {
+ if (mprotect(rw_start, (char *) rw_end - (char *) rw_start,
+ PROT_READ | PROT_WRITE) != 0)
+ {
+ ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ return false;
+ }
+ }
+
+ if (prot_end > rw_end)
+ {
+ if (mprotect(rw_end, (char *) prot_end - (char *) rw_end,
+ PROT_NONE) != 0)
+ {
+ ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ return false;
+ }
+ }
+
+ return true;
+#endif
+}
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 794e4fcb2ad..d266cdbe226 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -200,11 +200,13 @@ EnableLockPagesPrivilege(int elevel)
/*
* PGSharedMemoryCreate
*
- * Create a shared memory segment of the given size and initialize its
- * standard header.
+ * Create a shared memory segment of the given size and initialize its standard
+ * header. initial_size is only relevant when we want to create a larger memory
+ * segment with only a part of it being allocated initially. We don't support
+ * that on Windows, so we ignore it.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(Size initial_size, Size size,
PGShmemHeader **shim)
{
void *memAddress;
@@ -648,3 +650,59 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
}
return true;
}
+
+/*
+ * Get the page size used by the shared memory.
+ *
+ * The function should be called only after the shared memory has been setup.
+ */
+size_t
+GetOSPageSize(void)
+{
+ SYSTEM_INFO sysinfo;
+ size_t os_page_size;
+
+ Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
+
+ GetSystemInfo(&sysinfo);
+ os_page_size = sysinfo.dwPageSize;
+
+ /* If huge pages are actually in use, use huge page size */
+ if (huge_pages_status == HUGE_PAGES_ON)
+ GetHugePageSize(&os_page_size, NULL);
+
+ return os_page_size;
+}
+
+/*
+ * PGSharedMemoryEnsureFreed / PGSharedMemoryEnsureAllocated
+ *
+ * Not supported on Windows. These are only meaningful on platforms with
+ * resizable shared memory (mmap + madvise).
+ */
+bool
+PGSharedMemoryEnsureFreed(void *addr, Size size)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
+
+bool
+PGSharedMemoryEnsureAllocated(void *addr, Size size)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
+
+bool
+PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
diff --git a/src/backend/postmaster/auxprocess.c b/src/backend/postmaster/auxprocess.c
index ad4bf4bd2a8..07a3b5c5923 100644
--- a/src/backend/postmaster/auxprocess.c
+++ b/src/backend/postmaster/auxprocess.c
@@ -69,33 +69,43 @@ AuxiliaryProcessMainCommon(void)
BaseInit();
/*
- * Prevent consuming interrupts between setting ProcSignalInit and setting
- * the initial local data checksum value. If a barrier is emitted, and
- * absorbed, before local cached state is initialized the state transition
- * can be invalid.
- */
- HOLD_INTERRUPTS();
-
- ProcSignalInit(NULL, 0);
-
- /*
- * Initialize a local cache of the data_checksum_version, to be updated by
- * the procsignal-based barriers.
+ * Prevent consuming interrupts between ProcSignalInit() and the
+ * initialization of local states that are kept in sync with shared memory
+ * via procsignal-based barriers. If a barrier is emitted, and absorbed,
+ * before local cached state is initialized the state transition can be
+ * invalid.
*
- * This intentionally happens after initializing the procsignal, otherwise
- * we might miss a state change. This means we can get a barrier for the
- * state we've just initialized - but it can happen only once.
+ * These initializations intentionally happen after ProcSignalInit(),
+ * otherwise we might miss a state change. This means we may also receive
+ * a barrier for the state we've just initialized, but it can happen only
+ * once.
*
* The postmaster (which is what gets forked into the new child process)
- * does not handle barriers, therefore it may not have the current value
- * of LocalDataChecksumState value (it'll have the value read from the
- * control file, which may be arbitrarily old).
+ * does not handle barriers, therefore its local states may not reflect
+ * the current state of the shared memory.
*
* NB: Even if the postmaster handled barriers, the value might still be
* stale, as it might have changed after this process forked.
*/
+ HOLD_INTERRUPTS();
+
+ ProcSignalInit(NULL, 0);
+
+ /*
+ * LocalDataChecksumState inherited from Postmaster will have the value
+ * read from the control file, which may be arbitrarily old. Update it.
+ */
InitLocalDataChecksumState();
+ /*
+ * Refresh per-backend protections for resizable shmem structures, in case
+ * these structures have been resized since the startup. Structures are
+ * expected to be kept in sync by respective subsystems using
+ * procsignal-based barriers. But we modify their protections en-masse
+ * here, rather than relying on individual subsystems to do it.
+ */
+ ShmemReprotectResizableStructs();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index e149a738c8d..e5f20c4603e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -52,31 +52,50 @@ RequestAddinShmemSpace(Size size)
/*
* CalculateShmemSize
* Calculates the amount of shared memory needed.
+ *
+ * - `initial` is the amount of memory needed when the server startup.
+ * - `min` is the minimum amount of memory needed when all the resizable
+ * structures are shrunk to their respective minimum sizes.
+ * - `max` is the maximum amount of memory needed when all the resizable
+ * structures are expanded to their respective maximum sizes. It is also the
+ * size of address space that must be reserved for the shared memory segment.
+ *
+ * When no resizable structures are requested, all three totals are identical.
+ *
+ * We take some care to ensure that the total size request doesn't overflow
+ * size_t. If this gets through, we don't need to be so careful during the
+ * actual allocation phase.
*/
-Size
-CalculateShmemSize(void)
+void
+CalculateShmemSize(size_t *initial, size_t *min, size_t *max)
{
- Size size;
+ size_t initial_req;
+ size_t min_req;
+ size_t max_req;
+ size_t fixed_addins;
+
+ ShmemGetRequestedSize(&initial_req, &min_req, &max_req);
/*
* Size of the Postgres shared-memory block is estimated via moderately-
* accurate estimates for the big hogs, plus 100K for the stuff that's too
* small to bother with estimating.
*
- * We take some care to ensure that the total size request doesn't
- * overflow size_t. If this gets through, we don't need to be so careful
- * during the actual allocation phase.
+ * Also include additional requested shmem from preload libraries.
+ *
+ * These are not resizable, so they contribute equally to all three
+ * totals.
*/
- size = 100000;
- size = add_size(size, ShmemGetRequestedSize());
-
- /* include additional requested shmem from preload libraries */
- size = add_size(size, total_addin_request);
+ fixed_addins = add_size(100000, total_addin_request);
- /* might as well round it off to a multiple of a typical page size */
- size = add_size(size, 8192 - (size % 8192));
+ *initial = add_size(initial_req, fixed_addins);
+ *min = add_size(min_req, fixed_addins);
+ *max = add_size(max_req, fixed_addins);
- return size;
+ /* might as well round each off to a multiple of a typical page size */
+ *initial = add_size(*initial, 8192 - (*initial % 8192));
+ *min = add_size(*min, 8192 - (*min % 8192));
+ *max = add_size(*max, 8192 - (*max % 8192));
}
#ifdef EXEC_BACKEND
@@ -121,18 +140,24 @@ CreateSharedMemoryAndSemaphores(void)
{
PGShmemHeader *shim;
PGShmemHeader *seghdr;
- Size size;
+ size_t initial_size;
+ size_t min_size;
+ size_t max_size;
Assert(!IsUnderPostmaster);
/* Compute the size of the shared-memory block */
- size = CalculateShmemSize();
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+ CalculateShmemSize(&initial_size, &min_size, &max_size);
+ elog(DEBUG3, "invoking IpcMemoryCreate(initial size=%zu, minimum size=%zu, maximum size=%zu)",
+ initial_size, min_size, max_size);
/*
- * Create the shmem segment
+ * Create the shmem segment.
+ *
+ * Reserve enough shared address space to accommodate every requested
+ * structure grown to its maximum size.
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(initial_size, max_size, &shim);
/*
* Make sure that huge pages are never reported as "unknown" while the
@@ -189,32 +214,49 @@ void
InitializeShmemGUCs(void)
{
char buf[64];
- Size size_b;
- Size size_mb;
- Size hp_size;
+ size_t initial_b;
+ size_t min_b;
+ size_t max_b;
+ size_t hp_size;
+ size_t size_mb;
- /*
- * Calculate the shared memory size and round up to the nearest megabyte.
- */
- size_b = CalculateShmemSize();
- size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
+ CalculateShmemSize(&initial_b, &min_b, &max_b);
+
+ /* Round each size up to the nearest megabyte. */
+ size_mb = add_size(initial_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
- SetConfigOption("shared_memory_size", buf,
+ SetConfigOption("shared_memory_initial_size", buf,
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
- /*
- * Calculate the number of huge pages required.
- */
+ size_mb = add_size(min_b, (1024 * 1024) - 1) / (1024 * 1024);
+ sprintf(buf, "%zu", size_mb);
+ SetConfigOption("shared_memory_minimum_size", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ size_mb = add_size(max_b, (1024 * 1024) - 1) / (1024 * 1024);
+ sprintf(buf, "%zu", size_mb);
+ SetConfigOption("shared_memory_maximum_size", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ /* Calculate the number of huge pages required for each size. */
GetHugePageSize(&hp_size, NULL);
if (hp_size != 0)
{
- Size hp_required;
+ size_t hp_required;
+
+ hp_required = initial_b / hp_size + (initial_b % hp_size != 0);
+ sprintf(buf, "%zu", hp_required);
+ SetConfigOption("shared_memory_initial_size_in_huge_pages", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ hp_required = min_b / hp_size + (min_b % hp_size != 0);
+ sprintf(buf, "%zu", hp_required);
+ SetConfigOption("shared_memory_minimum_size_in_huge_pages", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
- hp_required = size_b / hp_size;
- if (size_b % hp_size != 0)
- hp_required = add_size(hp_required, 1);
+ hp_required = max_b / hp_size + (max_b % hp_size != 0);
sprintf(buf, "%zu", hp_required);
- SetConfigOption("shared_memory_size_in_huge_pages", buf,
+ SetConfigOption("shared_memory_maximum_size_in_huge_pages", buf,
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
}
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index a3d56cf55dd..fbc05f0dc13 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -19,11 +19,11 @@
* methods). The routines in this file are used for allocating and
* binding to shared memory data structures.
*
- * This module provides facilities to allocate fixed-size structures in shared
- * memory, for things like variables shared between all backend processes.
- * Each such structure has a string name to identify it, specified when it is
- * requested. shmem_hash.c provides a shared hash table implementation on top
- * of that.
+ * This module provides facilities to allocate fixed-size as well as resizable
+ * structures in shared memory, for things like variables shared between all
+ * backend processes. Each such structure has a string name to identify it,
+ * specified when it is requested. shmem_hash.c provides a shared hash table
+ * implementation on top of fixed-size structures.
*
* Shared memory areas should usually not be allocated after postmaster
* startup, although we do allow small allocations later for the benefit of
@@ -102,6 +102,23 @@
* (*options->ptr), and calls the attach_fn callback, if any, for additional
* per-backend setup.
*
+ * Resizable shared memory structures
+ * ----------------------------------
+ *
+ * In order to allocate resizable shared memory structures, set
+ * ShmemRequestStructOpts::maximum_size to the maximum size that the structure
+ * can grow to. The address space for the maximum size will be reserved at
+ * startup, but memory is allocated or freed as the structure grows or shrinks
+ * respectively. ShmemRequestStructOpts::size should be set to the initial size
+ * of the structure, which is the amount of memory allocated at the startup.
+ * Optionally, ShmemRequestStructOpts::minimum_size can be set to the minimum
+ * size that the structure can shrink to. After startup, the structure can be
+ * resized by calling ShmemResizeStruct(). ShmemResizeStruct() enforces that the
+ * new size is within [minimum_size, maximum_size].
+ *
+ * While resizable structures can be created after the startup, the memory
+ * available for them is quite limited.
+ *
* Legacy ShmemInitStruct()/ShmemInitHash() functions
* --------------------------------------------------
*
@@ -268,6 +285,10 @@ typedef struct
void *location; /* location in shared mem */
Size size; /* # bytes requested for the structure */
Size allocated_size; /* # bytes actually allocated */
+ Size minimum_size; /* the minimum size the structure can shrink
+ * to */
+ Size maximum_size; /* the maximum size the structure can grow to */
+ Size reserved_space; /* the total address space reserved */
} ShmemIndexEnt;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
@@ -276,6 +297,8 @@ static bool firstNumaTouch = true;
static void CallShmemCallbacksAfterStartup(const ShmemCallbacks *callbacks);
static void InitShmemIndexEntry(ShmemRequest *request);
static bool AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok);
+static Size EstimateAllocatedSize(ShmemIndexEnt *entry);
+static void ShmemProtectStructInternal(ShmemIndexEnt *entry);
Datum pg_numa_available(PG_FUNCTION_ARGS);
@@ -341,25 +364,63 @@ ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind)
if (options->name == NULL)
elog(ERROR, "shared memory request is missing 'name' option");
+#ifndef HAVE_RESIZABLE_SHMEM
+ if (options->maximum_size > 0)
+ elog(ERROR, "resizable shared memory is not supported on this platform");
+ if (options->minimum_size > 0)
+ elog(ERROR, "resizable shared memory is not supported on this platform");
+#else
+ if (options->maximum_size > 0 && shared_memory_type != SHMEM_TYPE_MMAP)
+ elog(ERROR, "resizable shared memory requires shared_memory_type = mmap");
+#endif
+
if (IsUnderPostmaster)
{
if (options->size <= 0 && options->size != SHMEM_ATTACH_UNKNOWN_SIZE)
elog(ERROR, "invalid size %zd for shared memory request for \"%s\"",
options->size, options->name);
+ if (options->minimum_size < 0 && options->minimum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ elog(ERROR, "invalid minimum_size %zd for shared memory request for \"%s\"",
+ options->minimum_size, options->name);
+ if (options->maximum_size < 0 && options->maximum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ elog(ERROR, "invalid maximum_size %zd for shared memory request for \"%s\"",
+ options->maximum_size, options->name);
}
else
{
- if (options->size == SHMEM_ATTACH_UNKNOWN_SIZE)
+ if (options->size == SHMEM_ATTACH_UNKNOWN_SIZE ||
+ options->minimum_size == SHMEM_ATTACH_UNKNOWN_SIZE ||
+ options->maximum_size == SHMEM_ATTACH_UNKNOWN_SIZE)
elog(ERROR, "SHMEM_ATTACH_UNKNOWN_SIZE cannot be used during startup");
if (options->size <= 0)
elog(ERROR, "invalid size %zd for shared memory request for \"%s\"",
options->size, options->name);
+ if (options->minimum_size < 0)
+ elog(ERROR, "invalid minimum_size %zd for shared memory request for \"%s\"",
+ options->minimum_size, options->name);
+ if (options->maximum_size < 0)
+ elog(ERROR, "invalid maximum_size %zd for shared memory request for \"%s\"",
+ options->maximum_size, options->name);
}
if (options->alignment != 0 && pg_nextpower2_size_t(options->alignment) != options->alignment)
elog(ERROR, "invalid alignment %zu for shared memory request for \"%s\"",
options->alignment, options->name);
+ if (options->minimum_size > 0 && options->size != SHMEM_ATTACH_UNKNOWN_SIZE &&
+ options->minimum_size > options->size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have minimum size (%zd) less than or equal to size (%zd)",
+ options->name, options->minimum_size, options->size);
+
+ if (options->maximum_size > 0 && options->size > options->maximum_size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have maximum size (%zd) greater than size (%zd)",
+ options->name, options->maximum_size, options->size);
+
+ if (options->minimum_size > 0 && options->maximum_size > 0 &&
+ options->minimum_size > options->maximum_size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have minimum size (%zd) less than or equal to maximum size (%zd)",
+ options->name, options->minimum_size, options->maximum_size);
+
/* Check that we're in the right state */
if (shmem_request_state != SRS_REQUESTING)
elog(ERROR, "ShmemRequestStruct can only be called from a shmem_request callback");
@@ -381,36 +442,70 @@ ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind)
}
/*
- * ShmemGetRequestedSize() --- estimate the total size of all registered shared
- * memory structures.
+ * ShmemGetRequestedSize() --- estimate total size of all registered shared
+ * memory structures.
+ *
+ * Returns three totals:
+ * - initial - the sum of initial sizes of the requested structures which is the
+ * total amount of memory required at the startup.
+ * - min - the total of minimum sizes of structures
+ * - max - the sum of maximum sizes of the structures, which is the address
+ * space that must be reserved.
+ *
+ * When there are no resizable structures or on the platforms that do not
+ * support resizable structures, all three totals are the same.
*
* This is called at postmaster startup, before the shared memory segment has
* been created.
*/
-size_t
-ShmemGetRequestedSize(void)
+void
+ShmemGetRequestedSize(size_t *initial, size_t *min, size_t *max)
{
- size_t size;
+ size_t initial_size;
+ size_t min_size;
+ size_t max_size;
/* memory needed for the ShmemIndex */
- size = hash_estimate_size(list_length(pending_shmem_requests) + SHMEM_INDEX_ADDITIONAL_SIZE,
- sizeof(ShmemIndexEnt));
- size = CACHELINEALIGN(size);
+ initial_size = hash_estimate_size(list_length(pending_shmem_requests) + SHMEM_INDEX_ADDITIONAL_SIZE,
+ sizeof(ShmemIndexEnt));
+ initial_size = CACHELINEALIGN(initial_size);
+ min_size = initial_size;
+ max_size = initial_size;
/* memory needed for all the requested areas */
foreach_ptr(ShmemRequest, request, pending_shmem_requests)
{
size_t alignment = request->options->alignment;
+ size_t req_min;
+ size_t req_max;
+ size_t req_initial = request->options->size;
+
+ if (request->options->maximum_size > 0)
+ {
+ req_min = request->options->minimum_size;
+ req_max = request->options->maximum_size;
+ }
+ else
+ {
+ req_min = req_initial;
+ req_max = req_initial;
+ }
/* pad the start address for alignment like ShmemAllocRaw() does */
if (alignment < PG_CACHE_LINE_SIZE)
alignment = PG_CACHE_LINE_SIZE;
- size = TYPEALIGN(alignment, size);
+ initial_size = TYPEALIGN(alignment, initial_size);
+ min_size = TYPEALIGN(alignment, min_size);
+ max_size = TYPEALIGN(alignment, max_size);
- size = add_size(size, request->options->size);
+ initial_size = add_size(initial_size, req_initial);
+ min_size = add_size(min_size, req_min);
+ max_size = add_size(max_size, req_max);
}
- return size;
+ *initial = initial_size;
+ *min = min_size;
+ *max = max_size;
}
/*
@@ -431,6 +526,23 @@ ShmemInitRequested(void)
* Initialize the ShmemIndex entries and perform basic initialization of
* all the requested memory areas. There are no concurrent processes yet,
* so no need for locking.
+ *
+ * TODO: If we have resizable structures, we will mmap with MAP_NORESERVE.
+ * In case there is not enough memory to cover the initial sizes of the
+ * shared structures, the mmap will succeed but the initialization will
+ * fail with SIGBUS. Instead we should allocate the memory worth the
+ * initial size of each structure using madvise(MADV_WRITE_POPOULATE). But
+ * instead of doing that for each structure, we should combine contiguous
+ * structures and do it once for the whole range. The algorithm to use is
+ * as follows 1. start from the first request 2. note the start address of
+ * the first structure in the range 3. walk upto a resizable structure or
+ * the last request, noting the initial end address of the last structure
+ * in the range 4. call madvise(MADV_WRITE_POPULATE) for the range (start
+ * address, end address) 5. Next structure becomes the first structure in
+ * the next range, repeat from step 2
+ *
+ * If there are no resizable structure, no need to do anything, as all the
+ * memory is reserved during mmap itself.
*/
foreach_ptr(ShmemRequest, request, pending_shmem_requests)
{
@@ -514,6 +626,7 @@ InitShmemIndexEntry(ShmemRequest *request)
ShmemIndexEnt *index_entry;
bool found;
size_t allocated_size;
+ size_t requested_size;
void *structPtr;
/* look it up in the shmem index */
@@ -531,10 +644,19 @@ InitShmemIndexEntry(ShmemRequest *request)
}
/*
- * We inserted the entry to the shared memory index. Allocate requested
- * amount of shared memory for it, and initialize the index entry.
+ * We inserted the entry to the shared memory index. Allocate requested
+ * amount of address space in the shared memory segment for it, and do
+ * basic initializion. The memory gets allocated during initialization as
+ * the corresponding memory pages are written to. Allocate enough space
+ * for a resizable structure to grow to its maximum size. It is expected
+ * that the initialization callback will use only as much memory as the
+ * initial size of the resizable structure. (Well, if it doesn't, more
+ * memory will be allocated initially than expected, no further harm is
+ * done.)
*/
- structPtr = ShmemAllocRaw(request->options->size,
+ requested_size = request->options->maximum_size > 0 ?
+ request->options->maximum_size : request->options->size;
+ structPtr = ShmemAllocRaw(requested_size,
request->options->alignment,
&allocated_size);
if (structPtr == NULL)
@@ -543,13 +665,36 @@ InitShmemIndexEntry(ShmemRequest *request)
hash_search(ShmemIndex, name, HASH_REMOVE, NULL);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("not enough shared memory for data structure"
+ errmsg("not enough shared memory space for data structure"
" \"%s\" (%zd bytes requested)",
- name, request->options->size)));
+ name, requested_size)));
}
index_entry->size = request->options->size;
index_entry->allocated_size = allocated_size;
index_entry->location = structPtr;
+ index_entry->reserved_space = allocated_size;
+ if (request->options->maximum_size > 0)
+ {
+ index_entry->minimum_size = request->options->minimum_size;
+ index_entry->maximum_size = request->options->maximum_size;
+
+ /* Adjust allocated size of a resizable structure. */
+ index_entry->allocated_size = EstimateAllocatedSize(index_entry);
+
+ /*
+ * Protect the unused part of the reserved address space for a
+ * resizable structure. If the structure has same minimum and maximum
+ * size, it is effectively a fixed-size structure without any unused
+ * space; no protection is required.
+ */
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ ShmemProtectStructInternal(index_entry);
+ }
+ else
+ {
+ index_entry->minimum_size = request->options->size;
+ index_entry->maximum_size = request->options->size;
+ }
/* Initialize depending on the kind of shmem area it is */
switch (request->kind)
@@ -594,7 +739,7 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
return false;
}
- /* Check that the size in the index matches the request */
+ /* Check that the sizes in the index match the request. */
if (index_entry->size != request->options->size &&
request->options->size != SHMEM_ATTACH_UNKNOWN_SIZE)
{
@@ -604,6 +749,43 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
name, index_entry->size, request->options->size)));
}
+ /*
+ * For resizable structures, also check that minimum_size and maximum_size
+ * match. For fixed-size structures, these are derived (set to size) in
+ * the index entry and not meaningful in the request.
+ */
+ if (request->options->maximum_size != 0)
+ {
+ if (index_entry->minimum_size != request->options->minimum_size &&
+ request->options->minimum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ {
+ ereport(ERROR,
+ errmsg("shared memory struct \"%s\" was created with"
+ " different minimum_size: existing %zu, requested %zu",
+ name, index_entry->minimum_size,
+ request->options->minimum_size));
+ }
+
+ if (index_entry->maximum_size != request->options->maximum_size &&
+ request->options->maximum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ {
+ ereport(ERROR,
+ errmsg("shared memory struct \"%s\" was created with"
+ " different maximum_size: existing %zu, requested %zu",
+ name, index_entry->maximum_size,
+ request->options->maximum_size));
+ }
+ }
+ else
+ {
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ elog(ERROR, "shared memory struct \"%s\" was created as resizable, but requested as fixed-size",
+ name);
+ }
+
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ ShmemProtectStructInternal(index_entry);
+
/*
* Re-establish the caller's pointer variable, or do other actions to
* attach depending on the kind of shmem area it is.
@@ -625,6 +807,291 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
return true;
}
+/*
+ * Estimate the actual memory allocated for a resizable structure.
+ *
+ * ... based on the assumption that the memory is allocated in pages.
+ *
+ * The memory pages covered by the current size of a resizable structure are
+ * considered to be allocated. The memory page where the maximal structure ends
+ * also hosts the next structure, unless the maximal structure ends on a page
+ * boundary. Hence that page is allocated because of the next structure. The
+ * memory pages between the page where the current structure ends and the page
+ * where the next structure starts remain unallocated. Thus the memory allocated
+ * for a resizable structure can be estimated as the total address space
+ * reserved for the structure minus the unallocated memory pages between the
+ * current end and the next structure.
+ */
+static Size
+EstimateAllocatedSize(ShmemIndexEnt *entry)
+{
+ Size page_size = GetOSPageSize();
+ char *align_end = (char *) TYPEALIGN(page_size, (char *) entry->location + entry->size);
+ char *floor_max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) entry->location + entry->maximum_size);
+
+ Assert(entry->maximum_size >= entry->size);
+ Assert(entry->reserved_space >= entry->maximum_size);
+
+ if (align_end < floor_max_end)
+ return entry->reserved_space - (floor_max_end - align_end);
+
+ return entry->reserved_space;
+}
+
+/*
+ * ShmemResizeStruct() --- resize a resizable shared memory structure.
+ *
+ * The new size must be within [minimum_size, maximum_size]. If the structure
+ * is being shrunk, the memory pages that are no longer needed are freed. If
+ * the structure is being expanded, the memory pages that are needed for the
+ * new size are allocated. See EstimateAllocatedSize() for explanation of which
+ * pages are allocated for a resizable structure.
+ *
+ * The caller must ensure that no other backend is accessing the part of the
+ * structure between the old size and the new size while this function is
+ * running. It should also ensure that all backends that may access the
+ * structure have observed the new size before they access the range between
+ * the old size and the new size.
+ *
+ * If we can not allocate memory pages when expanding the structure, this
+ * function will return false. On success it returns true. Instead
+ * of returning false, an error is raised if we can not free the memory pages
+ * when shrinking the structure, which should not happen in practice.
+ */
+bool
+ShmemResizeStruct(const char *name, Size new_size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ pg_unreachable();
+#else
+ ShmemIndexEnt *result;
+ bool found;
+ Size page_size = GetOSPageSize();
+ char *new_end;
+ bool success = true;
+
+ Assert(new_size > 0);
+
+ /*
+ * Resizable shared memory structures are only supported with mmap'ed
+ * memory.
+ */
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
+
+ /* look it up in the shmem index */
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+ result = (ShmemIndexEnt *) hash_search(ShmemIndex, name, HASH_FIND, &found);
+ if (!found)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shmem struct \"%s\" is not initialized", name));
+
+ Assert(result);
+
+ if (result->minimum_size == result->maximum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shared memory struct \"%s\" is not resizable", name));
+
+ if (new_size < result->minimum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("cannot shrink shared memory structure \"%s\" below minimum size"
+ "(requested %zu bytes, minimum %zu bytes)",
+ name, new_size, result->minimum_size));
+
+ if (result->maximum_size < new_size)
+ ereport(ERROR,
+ errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough address space is reserved for resizing structure \"%s\""
+ " (required %zu bytes, reserved %zu bytes)",
+ name, new_size, result->maximum_size));
+
+ /*
+ * A structure requires memory pages from the page containing the start of
+ * the structure and current end of the structure to be allocated. When
+ * expanding, we make sure that pages are allocated up to the new end.
+ * When shrinking, release memory pages beyond the new end, but not the
+ * page containing maximal end of the structure, as it may be used by the
+ * next structure.
+ *
+ * We do not consider the current end of the structure as it simplifies
+ * the calculations. Instead we rely on the underlying APIs not to touch
+ * the memory pages that will not be affected by the change in size.
+ */
+ new_end = (char *) TYPEALIGN(page_size, (char *) result->location + new_size);
+ if (new_size < result->size)
+ {
+ char *max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+
+ if (max_end > new_end)
+ {
+ if (!PGSharedMemoryEnsureFreed(new_end, max_end - new_end))
+ ereport(ERROR,
+ errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not free %zu bytes of shared memory from structure \"%s\"",
+ result->size - new_size, name));
+ }
+ }
+ else if (new_size > result->size)
+ {
+ char *struct_start = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location);
+
+ if (new_end > struct_start)
+ {
+ ShmemIndexEnt entry_copy = *result;
+
+ /*
+ * Allocating memory pages in the expanded range may require the
+ * corresponding address space to have read-write access.
+ */
+ entry_copy.size = new_size;
+ ShmemProtectStructInternal(&entry_copy);
+
+ if (!PGSharedMemoryEnsureAllocated(struct_start, new_end - struct_start))
+ {
+ ShmemProtectStructInternal(result);
+ ereport(WARNING,
+ errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("could not allocate %zu bytes of shared memory to structure \"%s\"",
+ new_size - result->size, name));
+ success = false;
+ }
+ }
+ }
+
+ /* Update shmem index entry. */
+ if (success)
+ {
+ result->size = new_size;
+ result->allocated_size = EstimateAllocatedSize(result);
+ }
+
+ LWLockRelease(ShmemIndexLock);
+
+ return success;
+#endif
+}
+
+/*
+ * ShmemProtectStruct() --- protect the unused portion of the given resizable
+ * structure.
+ *
+ * Makes the region beyond the current size up to maximum_size inaccessible, and
+ * ensures the region up to the current size is readable and writable. Depending
+ * upon the platform, the protection honours the page boundaries. So it may be
+ * more permissible than strictly needed.
+ *
+ * This function only affects the calling backend's address space. After each
+ * ShmemResizeStruct(), every backend that may access the structure should call
+ * this function before its next access. When backends are still in the middle
+ * of that round of ShmemProtectStruct() calls, ShmemResizeStruct() should not
+ * be called.
+ */
+void
+ShmemProtectStruct(const char *name)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ ShmemIndexEnt *result;
+ bool found;
+
+ LWLockAcquire(ShmemIndexLock, LW_SHARED);
+
+ result = (ShmemIndexEnt *) hash_search(ShmemIndex, name, HASH_FIND, &found);
+ if (!found)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shmem struct \"%s\" is not initialized", name));
+
+ if (result->minimum_size == result->maximum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shared memory struct \"%s\" is not resizable", name));
+
+ ShmemProtectStructInternal(result);
+
+ LWLockRelease(ShmemIndexLock);
+#endif
+}
+
+/*
+ * ShmemProtectStructInternal --- same as ShmemProtectStruct()
+ *
+ * ..., but called when the ShmemIndexEnt of the struct is available. The caller
+ * should hold ShmemIndexLock if required.
+ */
+static void
+ShmemProtectStructInternal(ShmemIndexEnt *entry)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ Size page_size = GetOSPageSize();
+ char *rw_start;
+ char *rw_end;
+ char *prot_end;
+
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
+ Assert(entry->minimum_size != entry->maximum_size);
+
+ /* Make at least [location, location+size) readable and writable */
+ rw_start = (char *) TYPEALIGN_DOWN(page_size, entry->location);
+ rw_end = (char *) TYPEALIGN(page_size,
+ (char *) entry->location + entry->size);
+
+ /*
+ * Make remaining portion inaccessible while making sure that the portion
+ * after maximum_size is not affected since it may be used by other
+ * structures.
+ */
+ prot_end = (char *) TYPEALIGN_DOWN(page_size,
+ (char *) entry->location + entry->maximum_size);
+
+ if (!PGSharedMemoryProtect(rw_start, rw_end, prot_end))
+ ereport(ERROR,
+ errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not protect shared memory structure \"%s\"", entry->key));
+#endif
+}
+
+/*
+ * ShmemReprotectResizableStructs() --- re-apply per-backend protection to every
+ * resizable shmem structure.
+ *
+ * On !EXEC_BACKEND platforms, a new backend inherits shmem mappings and their
+ * per-process protections from the postmaster. If any resizable structure has
+ * been resized since postmaster start, the inherited protections no longer
+ * match the current size. This is called early in per-backend startup so the
+ * new backend sees the protections according to the current sizes of the
+ * resizable structures.
+ */
+void
+ShmemReprotectResizableStructs(void)
+{
+#ifdef HAVE_RESIZABLE_SHMEM
+ HASH_SEQ_STATUS status;
+ ShmemIndexEnt *entry;
+
+ LWLockAcquire(ShmemIndexLock, LW_SHARED);
+ hash_seq_init(&status, ShmemIndex);
+ while ((entry = (ShmemIndexEnt *) hash_seq_search(&status)) != NULL)
+ {
+ if (entry->minimum_size != entry->maximum_size)
+ ShmemProtectStructInternal(entry);
+ }
+ LWLockRelease(ShmemIndexLock);
+#endif
+}
+
/*
* InitShmemAllocator() --- set up basic pointers to shared memory.
*
@@ -731,6 +1198,9 @@ InitShmemAllocator(PGShmemHeader *seghdr)
Assert(!found);
result->size = ShmemAllocator->index_size;
result->allocated_size = ShmemAllocator->index_size;
+ result->minimum_size = result->size;
+ result->maximum_size = result->size;
+ result->reserved_space = result->allocated_size;
result->location = ShmemAllocator->index;
}
}
@@ -1047,7 +1517,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Datum
pg_get_shmem_allocations(PG_FUNCTION_ARGS)
{
-#define PG_GET_SHMEM_SIZES_COLS 4
+#define PG_GET_SHMEM_SIZES_COLS 7
ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
HASH_SEQ_STATUS hstat;
ShmemIndexEnt *ent;
@@ -1069,7 +1539,20 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
- named_allocated += ent->allocated_size;
+ values[4] = Int64GetDatum(ent->minimum_size);
+ values[5] = Int64GetDatum(ent->maximum_size);
+ values[6] = Int64GetDatum(ent->reserved_space);
+
+ /*
+ * Anonymous areas are allocated in the area remaining after all named
+ * areas have been allocated. Thus the amount of shared memory
+ * allocated for anonymous areas can be calculated as the total amount
+ * of space allocated minus the amount of space allocated for named
+ * areas. The amount of free shared memory at the end of the segment
+ * can be calculated as the total size of the segment minus the total
+ * amount of space allocated.
+ */
+ named_allocated += ent->reserved_space;
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
values, nulls);
@@ -1080,6 +1563,9 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
nulls[1] = true;
values[2] = Int64GetDatum(ShmemAllocator->free_offset - named_allocated);
values[3] = values[2];
+ values[4] = values[2];
+ values[5] = values[2];
+ values[6] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
@@ -1088,6 +1574,9 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
nulls[1] = false;
values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemAllocator->free_offset);
values[3] = values[2];
+ values[4] = values[2];
+ values[5] = values[2];
+ values[6] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
LWLockRelease(ShmemIndexLock);
@@ -1274,23 +1763,9 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size
pg_get_shmem_pagesize(void)
{
- Size os_page_size;
-#ifdef WIN32
- SYSTEM_INFO sysinfo;
-
- GetSystemInfo(&sysinfo);
- os_page_size = sysinfo.dwPageSize;
-#else
- os_page_size = sysconf(_SC_PAGESIZE);
-#endif
-
Assert(IsUnderPostmaster);
- Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
-
- if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
- return os_page_size;
+ return GetOSPageSize();
}
Datum
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 9d6e69175a5..780cdb5e9ff 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -575,6 +575,14 @@ InitProcess(void)
if (IsUnderPostmaster)
AttachSharedMemoryStructs();
#endif
+
+ /*
+ * Update access to the address space occupied by the resizable shared
+ * structures. We do it here so that the structures can be accessed safely
+ * by this backend. But we will do this again after ProcSignalInit() for
+ * the reasons mentioned there.
+ */
+ ShmemReprotectResizableStructs();
}
/*
@@ -756,6 +764,14 @@ InitAuxiliaryProcess(void)
if (IsUnderPostmaster)
AttachSharedMemoryStructs();
#endif
+
+ /*
+ * Update access to the address space occupied by the resizable shared
+ * structures. We do it here so that the structures can be accessed safely
+ * by this backend. But we will do this again after ProcSignalInit() for
+ * the reasons mentioned there.
+ */
+ ShmemReprotectResizableStructs();
}
/*
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 3d8c9bdebd5..815b865aa38 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -787,6 +787,16 @@ InitPostgres(const char *in_dbname, Oid dboid,
*/
InitLocalDataChecksumState();
+ /*
+ * Refresh per-backend protections for resizable shmem structures. Usually
+ * the subsystems using resizable shared structures will use
+ * ProcSignalBarrier mechanism to coordinate resizing which would involve
+ * adjusting the protections as well. Like InitLocalDataChecksumState()
+ * above, this must run after ProcSignalInit so as not to miss a barrier
+ * for protection change.
+ */
+ ShmemReprotectResizableStructs();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index dccc4c82507..2cbec1981c5 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -1226,6 +1226,13 @@
max => '1000.0',
},
+{ name => 'have_resizable_shmem', type => 'bool', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows whether the running server supports resizable shared memory.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE',
+ variable => 'have_resizable_shmem_enabled',
+ boot_val => 'HAVE_RESIZABLE_SHMEM_ENABLED',
+},
+
{ name => 'hba_file', type => 'string', context => 'PGC_POSTMASTER', group => 'FILE_LOCATIONS',
short_desc => 'Sets the server\'s "hba" configuration file.',
flags => 'GUC_SUPERUSER_ONLY',
@@ -2722,20 +2729,58 @@
max => 'INT_MAX / 2',
},
-{ name => 'shared_memory_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
- short_desc => 'Shows the size of the server\'s main shared memory area (rounded up to the nearest MB).',
+{ name => 'shared_memory_initial_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the amount of memory allocated at server startup in the main shared memory area (rounded up to the nearest MB).',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_initial_size_mb',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_initial_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the number of huge pages needed in the main shared memory area at server startup.',
+ long_desc => '-1 means huge pages are not supported.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_initial_size_in_huge_pages',
+ boot_val => '-1',
+ min => '-1',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_maximum_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the size of the main shared memory area and maximum memory that can be allocated in that area (rounded up to the nearest MB).',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_maximum_size_mb',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_maximum_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the maximum number of huge pages needed in the main shared memory area.',
+ long_desc => '-1 means huge pages are not supported.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_maximum_size_in_huge_pages',
+ boot_val => '-1',
+ min => '-1',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_minimum_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the minimum amount of memory required in the main shared memory area (rounded up to the nearest MB).',
flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
- variable => 'shared_memory_size_mb',
+ variable => 'shared_memory_minimum_size_mb',
boot_val => '0',
min => '0',
max => 'INT_MAX',
},
-{ name => 'shared_memory_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
- short_desc => 'Shows the number of huge pages needed for the main shared memory area.',
+{ name => 'shared_memory_minimum_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the minimum number of huge pages needed in the main shared memory area.',
long_desc => '-1 means huge pages are not supported.',
flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
- variable => 'shared_memory_size_in_huge_pages',
+ variable => 'shared_memory_minimum_size_in_huge_pages',
boot_val => '-1',
min => '-1',
max => 'INT_MAX',
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 1ec460b6a82..17d50ffde38 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -643,8 +643,12 @@ static int max_index_keys;
static int max_identifier_length;
static int block_size;
static int segment_size;
-static int shared_memory_size_mb;
-static int shared_memory_size_in_huge_pages;
+static int shared_memory_initial_size_mb;
+static int shared_memory_minimum_size_mb;
+static int shared_memory_maximum_size_mb;
+static int shared_memory_initial_size_in_huge_pages;
+static int shared_memory_minimum_size_in_huge_pages;
+static int shared_memory_maximum_size_in_huge_pages;
static int wal_block_size;
static int num_os_semaphores;
static int effective_wal_level = WAL_LEVEL_REPLICA;
@@ -664,6 +668,13 @@ static bool assert_enabled = DEFAULT_ASSERT_ENABLED;
#endif
static bool exec_backend_enabled = EXEC_BACKEND_ENABLED;
+#ifdef HAVE_RESIZABLE_SHMEM
+#define HAVE_RESIZABLE_SHMEM_ENABLED true
+#else
+#define HAVE_RESIZABLE_SHMEM_ENABLED false
+#endif
+static bool have_resizable_shmem_enabled = HAVE_RESIZABLE_SHMEM_ENABLED;
+
static char *recovery_target_timeline_string;
static char *recovery_target_string;
static char *recovery_target_xid_string;
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index f8a021987b5..712172760b3 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8692,8 +8692,8 @@
{ oid => '5052', descr => 'allocations from the main shared memory segment',
proname => 'pg_get_shmem_allocations', prorows => '50', proretset => 't',
provolatile => 'v', prorettype => 'record', proargtypes => '',
- proallargtypes => '{text,int8,int8,int8}', proargmodes => '{o,o,o,o}',
- proargnames => '{name,off,size,allocated_size}',
+ proallargtypes => '{text,int8,int8,int8,int8,int8,int8}', proargmodes => '{o,o,o,o,o,o,o}',
+ proargnames => '{name,off,size,allocated_size,minimum_size,maximum_size,reserved_space}',
prosrc => 'pg_get_shmem_allocations',
proacl => '{POSTGRES=X,pg_read_all_stats=X}' },
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 661c4a9b168..661c12e9bf3 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,14 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `MADV_POPULATE_WRITE', and to 0
+ if you don't. */
+#undef HAVE_DECL_MADV_POPULATE_WRITE
+
+/* Define to 1 if you have the declaration of `MADV_REMOVE', and to 0 if you
+ don't. */
+#undef HAVE_DECL_MADV_REMOVE
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/include/pg_config_manual.h b/src/include/pg_config_manual.h
index 521b49b8888..ab944babe2b 100644
--- a/src/include/pg_config_manual.h
+++ b/src/include/pg_config_manual.h
@@ -131,6 +131,20 @@
#define EXEC_BACKEND
#endif
+/*
+ * HAVE_RESIZABLE_SHMEM indicates whether resizable shared memory structures are
+ * supported. The implementation requires Linux-specific madvise constants
+ * (MADV_REMOVE and MADV_POPULATE_WRITE) and existence of mprotect() API.
+ *
+ * TODO: We may want to remove EXEC_BACKEND from the condition to test attaching
+ * and resizing resizable shared memory structures in EXEC_BACKEND mode. Windows
+ * will anyway won't have HAVE_RESIZABLE_SHMEM defined since it won't have
+ * MADV_REMOVE and MADV_POPULATE_WRITE.
+ */
+#if HAVE_DECL_MADV_REMOVE && HAVE_DECL_MADV_POPULATE_WRITE && !defined(EXEC_BACKEND)
+#define HAVE_RESIZABLE_SHMEM
+#endif
+
/*
* USE_POSIX_FADVISE controls whether Postgres will attempt to use the
* posix_fadvise() kernel call. Usually the automatic configure tests are
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index b205b00e7a1..46ae87fe863 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -78,7 +78,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
extern void RegisterBuiltinShmemCallbacks(void);
-extern Size CalculateShmemSize(void);
+extern void CalculateShmemSize(size_t *initial, size_t *min, size_t *max);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 10c7b065861..50474c2d579 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -85,10 +85,14 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(Size size,
+extern PGShmemHeader *PGSharedMemoryCreate(Size initial_size, Size max_size,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
+extern bool PGSharedMemoryEnsureFreed(void *addr, Size size);
+extern bool PGSharedMemoryEnsureAllocated(void *addr, Size size);
+extern bool PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern size_t GetOSPageSize(void);
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 43b636868c7..947debcede4 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -57,6 +57,22 @@ typedef struct ShmemStructOpts
*/
size_t alignment;
+ /*
+ * Minimum size this structure can shrink to. Should be set to 0 for
+ * fixed-size structures.
+ */
+ ssize_t minimum_size;
+
+ /*
+ * Maximum size this structure can grow upto in future. The memory is not
+ * allocated right away but the corresponding address space is reserved so
+ * that memory can be mapped to it when the structure grows. Typically
+ * should be used for large resizable structures which need several pages
+ * worth of contiguous memory. Should be set to 0 for fixed-size
+ * structures.
+ */
+ ssize_t maximum_size;
+
/*
* When the shmem area is initialized or attached to, pointer to it is
* stored in *ptr. It usually points to a global variable, used to access
@@ -168,6 +184,9 @@ typedef struct ShmemCallbacks
extern void RegisterShmemCallbacks(const ShmemCallbacks *callbacks);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemResizeStruct(const char *name, Size new_size);
+extern void ShmemProtectStruct(const char *name);
+extern void ShmemReprotectResizableStructs(void);
/*
* These macros provide syntactic sugar for calling the underlying functions
diff --git a/src/include/storage/shmem_internal.h b/src/include/storage/shmem_internal.h
index 8746b614fa3..6c8c81812a8 100644
--- a/src/include/storage/shmem_internal.h
+++ b/src/include/storage/shmem_internal.h
@@ -36,7 +36,7 @@ extern void ResetShmemAllocator(void);
extern void ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind);
-extern size_t ShmemGetRequestedSize(void);
+extern void ShmemGetRequestedSize(size_t *initial, size_t *min, size_t *max);
extern void ShmemInitRequested(void);
#ifdef EXEC_BACKEND
extern void ShmemAttachRequested(void);
diff --git a/src/test/modules/test_shmem/meson.build b/src/test/modules/test_shmem/meson.build
index fb4bf328b8f..1f795aae7eb 100644
--- a/src/test/modules/test_shmem/meson.build
+++ b/src/test/modules/test_shmem/meson.build
@@ -27,7 +27,8 @@ tests += {
'bd': meson.current_build_dir(),
'tap': {
'tests': [
- 't/001_late_shmem_alloc.pl',
+ 't/001_fixed_shmem_struct.pl',
+ 't/002_resizable_shmem_struct.pl',
],
},
}
diff --git a/src/test/modules/test_shmem/t/001_late_shmem_alloc.pl b/src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
similarity index 58%
rename from src/test/modules/test_shmem/t/001_late_shmem_alloc.pl
rename to src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
index 5cf07d071ec..d231821e0b4 100644
--- a/src/test/modules/test_shmem/t/001_late_shmem_alloc.pl
+++ b/src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
@@ -56,5 +56,36 @@ else
);
}
+###
+# Test that a fixed-size shared memory structure cannot be resized.
+# Only relevant on platforms that support resizable shmem.
+###
+my $have_resizable_shmem =
+ $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+if ($have_resizable_shmem)
+{
+ # Try expanding the fixed-size structure
+ my ($ret, $stdout, $stderr) =
+ $node->psql("postgres", "SELECT test_shmem_resize_fixed(1000);");
+ isnt($ret, 0, "expanding a fixed-size structure fails");
+ like($stderr, qr/is not resizable/, "expand error message mentions not resizable");
+
+ # Try shrinking the fixed-size structure
+ ($ret, $stdout, $stderr) =
+ $node->psql("postgres", "SELECT test_shmem_resize_fixed(1);");
+ isnt($ret, 0, "shrinking a fixed-size structure fails");
+ like($stderr, qr/is not resizable/, "shrink error message mentions not resizable");
+}
+
+###
+# Test that minimum_size and maximum_size equal size for a fixed-size structure
+# in pg_shmem_allocations.
+###
+is($node->safe_psql('postgres',
+ "SELECT minimum_size = size AND maximum_size = size FROM pg_shmem_allocations WHERE name = 'test_shmem area';"),
+ 't', "fixed-size structure has minimum_size = maximum_size = size");
+
$node->stop;
+
done_testing();
diff --git a/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl b/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
new file mode 100644
index 00000000000..4d2fd3c4282
--- /dev/null
+++ b/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
@@ -0,0 +1,382 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Test resizable shared memory functionality, both when loaded at startup via
+# shared_preload_libraries and when loaded after startup (late allocation).
+
+# Verify that enough shared memory is allocated to cover the resizable_shmem
+# structure at its current size but does not exceed the memory required by the
+# current sizes of all shared memory structures. We expect that the backend
+# where we run the query will have touched the entire resizable_shmem structure,
+# so that all the memory pages covering the resizable structure are mapped to
+# the backend's address space.
+#
+# Since we have configured the server so that resizable shared struture
+# dominates the main shared memory segment, the total memory allocated to other
+# shared memory structures does not result in false positive tests below.
+sub check_shmem_usage
+{
+ my ($session, $label, $node) = @_;
+
+ my $shmem_usage = $session->query_safe('SELECT test_shmem_usage();', verbose => 0);
+ my $total_alloc = $node->safe_psql('postgres', "SELECT sum(allocated_size) FROM pg_shmem_allocations;");
+ my $resizable_alloc = $node->safe_psql('postgres',
+ "SELECT allocated_size FROM pg_shmem_allocations WHERE name = 'resizable_shmem';");
+
+ diag "$label: shmem_usage=$shmem_usage, resizable_shmem allocated=$resizable_alloc, sum(allocated_size)=$total_alloc";
+ ok($shmem_usage <= $total_alloc,
+ "$label: allocated shared memory does not exceed total allocated size");
+ ok($shmem_usage >= $resizable_alloc,
+ "$label: shared memory usage covers the resizable_shmem allocation");
+}
+
+# Test a resize operation: resize, verify old data, write new data, verify
+# new data, and check shmem usage. Returns updated ($num_entries, $value).
+sub test_resize
+{
+ my ($node, $prefix, $old_num_entries, $old_value, $new_num_entries, $new_value, $label) = @_;
+
+ $label = "$prefix: $label";
+
+ my $session1 = $node->background_psql('postgres');
+ my $session2 = $node->background_psql('postgres');
+
+ $session1->query_safe("SELECT resizable_shmem_resize($new_num_entries);",
+ verbose => 0);
+
+ # Old data should still be intact in the (possibly smaller) area
+ my $readable_entries = ($new_num_entries < $old_num_entries) ? $new_num_entries : $old_num_entries;
+ is($session1->query_safe("SELECT resizable_shmem_read($readable_entries, $old_value);",
+ verbose => 0),
+ 't', "old data readable after $label");
+
+ $session2->query_safe("SELECT resizable_shmem_write($new_value);",
+ verbose => 0);
+ is($session1->query_safe("SELECT resizable_shmem_read($new_num_entries, $new_value);",
+ verbose => 0),
+ 't', "new data readable after $label");
+
+ check_shmem_usage($session1, "$label (session 1)", $node);
+ check_shmem_usage($session2, "$label (session 2)", $node);
+
+ $session1->quit;
+ $session2->quit;
+
+ return ($new_num_entries, $new_value);
+}
+
+# Verify that reads or writes past the current size, but within the reserved
+# maximum, fault when they reach the protected region.
+sub test_fault_beyond_size
+{
+ my ($node, $initial_entries, $prefix) = @_;
+
+ # Enable restart_after_crash to test postmaster driven restart with
+ # resizable shared memory.
+ $node->safe_psql('postgres',
+ 'ALTER SYSTEM SET restart_after_crash = on;');
+ $node->reload;
+
+ for my $mode ('write', 'read')
+ {
+ my ($ret, $stdout, $stderr) = $node->psql('postgres',
+ "SELECT resizable_shmem_access_beyond_size('$mode');");
+ ok($ret != 0, "$prefix: $mode past current size crashes the backend");
+ like($stderr,
+ qr/server closed the connection unexpectedly|connection to server was lost/,
+ "$prefix: $mode crash reports lost connection");
+
+ $node->poll_query_until('postgres', 'SELECT 1', '1')
+ or die "server did not come back after $mode crash";
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT resizable_shmem_read($initial_entries, 0);"),
+ 't', "$prefix: read succeeds after crash recovery");
+
+ $node->safe_psql('postgres', 'ALTER SYSTEM RESET restart_after_crash;');
+ $node->reload;
+}
+
+# Run the full suite of resizable shared memory tests on the given node.
+sub run_resizable_tests
+{
+ my ($node, $initial_entries, $max_entries, $prefix) = @_;
+ my $have_resizable_shmem = $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+ my $num_entries = $initial_entries;
+
+ # Basic read/write should work on all platforms
+ my $value = 100;
+ $node->safe_psql('postgres', "SELECT resizable_shmem_write($value);");
+ is($node->safe_psql('postgres', "SELECT resizable_shmem_read($num_entries, $value);"),
+ 't', "$prefix: data read after write successful");
+
+ if ($have_resizable_shmem)
+ {
+ # Initial structure state
+ my $session1 = $node->background_psql('postgres');
+ my $session2 = $node->background_psql('postgres');
+
+ $value = 100;
+ # Write and read the initial set of entries.
+ $session1->query_safe("SELECT resizable_shmem_write($value);", verbose => 0);
+ is($session2->query_safe("SELECT resizable_shmem_read($num_entries, $value);",
+ verbose => 0),
+ 't', "$prefix: data read after write successful");
+ check_shmem_usage($session1, "$prefix: initial write (session 1)", $node);
+ check_shmem_usage($session2, "$prefix: initial write (session 2)", $node);
+ $session1->quit;
+ $session2->quit;
+
+ # Verify no other structure is resizable
+ is($node->safe_psql('postgres', "SELECT count(*) FROM pg_shmem_allocations WHERE name <> 'resizable_shmem' AND maximum_size <> minimum_size;"),
+ '0', "$prefix: no other resizable structures");
+
+ # Resize to maximum
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ $max_entries, 500, 'resize to maximum');
+
+ # Shrink to 75% of max
+ my $shrink_entries = int($max_entries * 3 / 4);
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ $shrink_entries, 999, 'shrinking');
+
+ # Resize to the same size (no-op)
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ $num_entries, 1999, 'no-op resize');
+
+ # Shrink to minimum i.e. zero entries and grow back
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ 0, 0, 'shrink to minimum');
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ $initial_entries, 2999,
+ 'grow back from minimum');
+
+ # Test resize failure (attempt to resize beyond max - should fail)
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', "SELECT resizable_shmem_resize(" . ($max_entries * 2) . ");");
+ ok($ret != 0 || $stderr =~ /ERROR/, "$prefix: Resize beyond maximum fails");
+
+ # Resize to a size below minimum_size must fail.
+ ($ret, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT resizable_shmem_resize(-1);');
+ ok($ret != 0, "$prefix: resize below minimum_size fails");
+ like($stderr,
+ qr/cannot shrink shared memory structure "resizable_shmem" below minimum size/,
+ "$prefix: resize-below-minimum error comes from ShmemResizeStruct");
+
+ # The fault test relies on a hole being present between the current end
+ # of the structure and its maximal end. Skip when the structure does not
+ # span multiple pages.
+ my $spans_pages = $node->safe_psql('postgres', qq{
+ SELECT (maximum_size - minimum_size) >= test_shmem_pagesize()
+ FROM pg_shmem_allocations WHERE name = 'resizable_shmem';
+ });
+ if ($spans_pages ne 't')
+ {
+ diag "$prefix: skipping fault-beyond-size test: resizable_shmem does not span multiple shmem pages";
+ }
+ else
+ {
+ test_fault_beyond_size($node, $initial_entries, $prefix);
+ }
+ }
+ else
+ {
+ # On unsupported platforms, resizing should fail with a clear error
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', "SELECT resizable_shmem_resize($num_entries);");
+ ok($ret != 0, "$prefix: resize fails on unsupported platform");
+ like($stderr, qr/not supported/, "$prefix: resize error mentions not supported");
+ }
+}
+
+# Check the runtime-computed shared_memory_{initial,minimum,maximum}_size GUC
+# invariants. min <= initial <= max must always hold. When a resizable
+# structure has been registered on a server that supports resizable shared
+# memory structures, min must additionally be strictly less than max;
+# otherwise all three GUCs must be equal.
+sub check_shmem_size_gucs
+{
+ my ($node, $label) = @_;
+ my $pgdata = $node->data_dir;
+
+ my $get = sub {
+ my ($guc) = @_;
+ my ($stdout, $stderr) = run_command([ 'postgres', '-D' => $pgdata, '-C' => $guc ]);
+
+ return $stdout;
+ };
+
+ my $have_resizable_shmem = $get->('have_resizable_shmem');
+ my $ini = 0 + $get->('shared_memory_initial_size');
+ my $min = 0 + $get->('shared_memory_minimum_size');
+ my $max = 0 + $get->('shared_memory_maximum_size');
+ my $have_resizable_struct = ($get->('resizable_shmem.max_entries') ne '');
+
+ ok($min <= $ini && $ini <= $max, "$label: shared_memory size GUCs in expected order");
+
+ if ($have_resizable_struct && $have_resizable_shmem eq 'on')
+ {
+ ok($min < $max, "$label: min < max with resizable structures");
+ }
+ else
+ {
+ ok($min == $ini && $ini == $max,
+ "$label: all shared_memory size GUCs equal when no resizable structures");
+ }
+}
+
+# Log the runtime shared_memory_{initial,minimum,maximum}_size GUCs and huge
+# pages usage information for easier debugging.
+sub diag_shmem_sizes
+{
+ my ($node, $label) = @_;
+
+ my $vals = $node->safe_psql('postgres', q{
+ SELECT format('initial=%s minimum=%s maximum=%s huge_pages_status=%s huge_page_size=%s shmem_page_size=%s',
+ current_setting('shared_memory_initial_size'),
+ current_setting('shared_memory_minimum_size'),
+ current_setting('shared_memory_maximum_size'),
+ current_setting('huge_pages_status'),
+ current_setting('huge_page_size'),
+ test_shmem_pagesize());
+ });
+ diag "$label: $vals";
+}
+
+### Set up a test node.
+#
+# Configure minimal shared memory so that the resizable_shmem structure dominates
+# and any unexpected increase is easy to detect.
+#
+# If we turn on huge pages and the machine where the test is running does not
+# have huge pages available, the test will fail midway because it will not be
+# able to allocate memory pages when expanding the resizable_shmem structure.
+# Hence we turn off huge pages for this test. The test outputs the GUCs
+# shared_memory_{initial,minimum,maximum}_size and information about huge pages.
+# By provisioning enough huge pages, and by changing huge_pages = try/on, the
+# test can be run with huge pages enabled.
+###
+my $node = PostgreSQL::Test::Cluster->new('resizable_shmem');
+$node->init;
+
+$node->append_conf('postgresql.conf', 'huge_pages = off');
+$node->append_conf('postgresql.conf', 'shared_buffers = 128kB');
+$node->append_conf('postgresql.conf', 'max_connections = 5');
+$node->append_conf('postgresql.conf', 'max_worker_processes = 0');
+$node->append_conf('postgresql.conf', 'max_wal_senders = 0');
+$node->append_conf('postgresql.conf', 'max_prepared_transactions = 0');
+$node->append_conf('postgresql.conf', 'max_locks_per_transaction = 10');
+$node->append_conf('postgresql.conf', 'max_pred_locks_per_transaction = 10');
+$node->append_conf('postgresql.conf', 'wal_buffers = 32kB');
+
+###
+# Test 1: Startup allocation via shared_preload_libraries
+###
+my $startup_initial = 25 * 1024 * 1024;
+my $startup_max = 100 * 1024 * 1024;
+
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = test_shmem');
+$node->append_conf('postgresql.conf', "resizable_shmem.initial_entries = $startup_initial");
+$node->append_conf('postgresql.conf', "resizable_shmem.max_entries = $startup_max");
+
+check_shmem_size_gucs($node, 'startup preload');
+
+$node->start;
+$node->safe_psql('postgres', 'CREATE EXTENSION test_shmem;');
+diag_shmem_sizes($node, 'startup');
+run_resizable_tests($node, $startup_initial, $startup_max, 'startup');
+
+my $have_resizable_shmem = $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+###
+# Test 2: Late allocation (loaded after startup, not in shared_preload_libraries).
+# Use much smaller sizes since only ~100KB of shared memory is available for
+# structures allocated after startup.
+###
+my $late_initial = 5 * 1024;
+my $late_max = 12 * 1024;
+
+$node->safe_psql('postgres', qq{
+ ALTER SYSTEM RESET shared_preload_libraries;
+ ALTER SYSTEM SET resizable_shmem.initial_entries = $late_initial;
+ ALTER SYSTEM SET resizable_shmem.max_entries = $late_max;
+});
+$node->safe_psql('postgres', 'DROP EXTENSION test_shmem;');
+$node->restart;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION test_shmem;');
+diag_shmem_sizes($node, 'late');
+run_resizable_tests($node, $late_initial, $late_max, 'late');
+
+###
+# Test sysv shared memory does not support resizable shmem. Only relevant on
+# platforms that support resizable shmem (HAVE_RESIZABLE_SHMEM), since the
+# module only sets maximum_size in that case.
+###
+if ($have_resizable_shmem)
+{
+ ###
+ # Test 3: Verify that CREATE EXTENSION fails with sysv shared memory
+ # when loaded after startup (not in shared_preload_libraries).
+ ###
+ $node->safe_psql('postgres', 'DROP EXTENSION test_shmem;');
+
+ # Remove settings that would cause the library to auto-load at startup:
+ # shared_preload_libraries and module-prefixed GUCs. ALTER SYSTEM RESET
+ # only affects postgresql.auto.conf, so we must use adjust_conf to remove
+ # from postgresql.conf.
+ $node->adjust_conf('postgresql.conf', 'shared_preload_libraries', undef);
+ $node->adjust_conf('postgresql.conf', 'resizable_shmem.initial_entries', undef);
+ $node->adjust_conf('postgresql.conf', 'resizable_shmem.max_entries', undef);
+ $node->adjust_conf('postgresql.auto.conf', 'shared_preload_libraries', undef);
+ $node->adjust_conf('postgresql.auto.conf', 'resizable_shmem.initial_entries', undef);
+ $node->adjust_conf('postgresql.auto.conf', 'resizable_shmem.max_entries', undef);
+ $node->safe_psql('postgres', qq{
+ ALTER SYSTEM SET shared_memory_type = 'sysv';
+ });
+
+ $node->stop;
+
+ check_shmem_size_gucs($node, 'sysv');
+
+ $node->start;
+
+ is($node->safe_psql('postgres', 'SHOW have_resizable_shmem;'),
+ 'off',
+ 'have_resizable_shmem reports off with shared_memory_type = sysv');
+
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', 'CREATE EXTENSION test_shmem;');
+ ok($ret != 0, 'CREATE EXTENSION fails with resizable shmem on sysv');
+ like($stderr, qr/resizable shared memory requires shared_memory_type = mmap/,
+ 'CREATE EXTENSION error mentions shared_memory_type = mmap requirement');
+
+ ###
+ # Test 4: Verify that resizable structures are also rejected with sysv
+ # shared memory when loaded at startup via shared_preload_libraries.
+ ###
+ $node->safe_psql('postgres', qq{
+ ALTER SYSTEM SET shared_preload_libraries = 'test_shmem';
+ ALTER SYSTEM SET resizable_shmem.initial_entries = $startup_initial;
+ ALTER SYSTEM SET resizable_shmem.max_entries = $startup_max;
+ });
+ $node->stop;
+
+ ok(!$node->start(fail_ok => 1),
+ 'server fails to start with resizable shmem on sysv');
+
+ my $log = slurp_file($node->logfile);
+ like($log, qr/resizable shared memory requires shared_memory_type = mmap/,
+ 'log mentions shared_memory_type = mmap requirement');
+}
+
+done_testing();
diff --git a/src/test/modules/test_shmem/test_shmem--1.0.sql b/src/test/modules/test_shmem/test_shmem--1.0.sql
index 2d01fd9256c..eb695604b1d 100644
--- a/src/test/modules/test_shmem/test_shmem--1.0.sql
+++ b/src/test/modules/test_shmem/test_shmem--1.0.sql
@@ -4,6 +4,61 @@
\echo Use "CREATE EXTENSION test_shmem" to load this file. \quit
+-- ===================================================================
+-- Fixed-size shared memory structure
+-- ===================================================================
+
CREATE FUNCTION get_test_shmem_attach_count()
RETURNS pg_catalog.int4 STRICT
AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION test_shmem_resize_fixed(pg_catalog.int4)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+-- ===================================================================
+-- Resizable shared memory structure
+-- ===================================================================
+
+-- Function to resize the test structure in the shared memory
+CREATE FUNCTION resizable_shmem_resize(new_entries integer)
+RETURNS bool
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to write data to all entries in the test structure in shared memory
+-- Writing all the entries makes sure that the memory is actually allocated and
+-- mapped to the process, so that we can later measure the memory usage.
+CREATE FUNCTION resizable_shmem_write(entry_value integer)
+RETURNS void
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to verify that specified number of initial entries have expected value.
+-- Reading all the entries makes sure that the memory is actually mapped to the
+-- process, so that we can later measure the memory usage.
+CREATE FUNCTION resizable_shmem_read(entry_count integer, entry_value integer)
+RETURNS boolean
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to report memory mapped against the main shared memory segment in
+-- the backend where this function runs.
+CREATE FUNCTION test_shmem_usage()
+RETURNS bigint
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to get the shared memory page size
+CREATE FUNCTION test_shmem_pagesize()
+RETURNS integer
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to crash the backend by walking entries past the current size up to
+-- the reserved maximum, reading or writing each one as decided by mode.
+CREATE FUNCTION resizable_shmem_access_beyond_size(mode text)
+RETURNS integer
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
diff --git a/src/test/modules/test_shmem/test_shmem.c b/src/test/modules/test_shmem/test_shmem.c
index 9bd4012b435..83f8ea9cc3c 100644
--- a/src/test/modules/test_shmem/test_shmem.c
+++ b/src/test/modules/test_shmem/test_shmem.c
@@ -1,11 +1,10 @@
/*-------------------------------------------------------------------------
*
* test_shmem.c
- * Helpers to test shmem allocation routines
+ * Helpers to test shmem management routines
*
- * Test basic memory allocation in an extension module. One notable feature
- * that is not exercised by any other module in the repository is the
- * allocating (non-DSM) shared memory after postmaster startup.
+ * Test fixed-size and resizable shared memory structures created during
+ * postmaster startup and after startup respectively.
*
* Copyright (c) 2020-2026, PostgreSQL Global Development Group
*
@@ -17,13 +16,26 @@
#include "postgres.h"
+#include <limits.h>
+
+#include "commands/extension.h"
#include "fmgr.h"
#include "miscadmin.h"
+#include "storage/fd.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
+#include "utils/builtins.h"
+#include "utils/guc.h"
PG_MODULE_MAGIC;
+
+/* ----------------------------------------------------------------
+ * Fixed-size shared memory structure
+ * ----------------------------------------------------------------
+ */
+
typedef struct TestShmemData
{
int value;
@@ -35,17 +47,6 @@ static TestShmemData *TestShmem;
static bool attached_or_initialized = false;
-static void test_shmem_request(void *arg);
-static void test_shmem_init(void *arg);
-static void test_shmem_attach(void *arg);
-
-static const ShmemCallbacks TestShmemCallbacks = {
- .flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP,
- .request_fn = test_shmem_request,
- .init_fn = test_shmem_init,
- .attach_fn = test_shmem_attach,
-};
-
static void
test_shmem_request(void *arg)
{
@@ -60,6 +61,17 @@ static void
test_shmem_init(void *arg)
{
elog(LOG, "init callback called");
+
+ /*
+ * Reset the per-process flag and the shared "initialized" marker during
+ * postmaster induced restart.
+ */
+ if (!IsUnderPostmaster)
+ {
+ attached_or_initialized = false;
+ TestShmem->initialized = false;
+ }
+
if (TestShmem->initialized)
elog(ERROR, "shmem area already initialized");
TestShmem->initialized = true;
@@ -82,12 +94,12 @@ test_shmem_attach(void *arg)
attached_or_initialized = true;
}
-void
-_PG_init(void)
-{
- elog(LOG, "test_shmem module's _PG_init called");
- RegisterShmemCallbacks(&TestShmemCallbacks);
-}
+static const ShmemCallbacks TestShmemCallbacks = {
+ .flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP,
+ .request_fn = test_shmem_request,
+ .init_fn = test_shmem_init,
+ .attach_fn = test_shmem_attach,
+};
PG_FUNCTION_INFO_V1(get_test_shmem_attach_count);
Datum
@@ -99,3 +111,453 @@ get_test_shmem_attach_count(PG_FUNCTION_ARGS)
elog(ERROR, "shmem area not yet initialized");
PG_RETURN_INT32(TestShmem->attach_count);
}
+
+/*
+ * Attempt to resize the fixed-size shared memory structure. This should
+ * fail because the structure was not allocated with a maximum_size.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_resize_fixed);
+Datum
+test_shmem_resize_fixed(PG_FUNCTION_ARGS)
+{
+ int32 new_size = PG_GETARG_INT32(0);
+
+ ShmemResizeStruct("test_shmem area", new_size);
+ PG_RETURN_VOID();
+}
+
+
+/* ----------------------------------------------------------------
+ * Resizable shared memory structure
+ * ----------------------------------------------------------------
+ */
+
+/*
+ * The test module may be loaded after postmaster startup in which case only
+ * 100K of shared memory is available for the extension. Keep the default
+ * initial and maximum sizes small enough to fit in that space.
+ */
+#define TEST_INITIAL_ENTRIES_DEFAULT 1
+#define TEST_MAX_ENTRIES_DEFAULT 1024
+
+#define TEST_ENTRY_SIZE sizeof(int32) /* Size of each entry */
+
+/*
+ * Resizable test data structure stored in shared memory.
+ *
+ * The test performs resizing, reads or writes, only one at a time and never
+ * concurrently. Hence, there is no need for locks in the test structure.
+ */
+typedef struct TestResizableShmemStruct
+{
+ /* Metadata */
+ int32 num_entries; /* Number of entries that can fit */
+
+ /* Data area - variable size */
+ int32 data[FLEXIBLE_ARRAY_MEMBER];
+} TestResizableShmemStruct;
+
+static TestResizableShmemStruct *resizable_shmem = NULL;
+
+/* GUC variables controlling the size of the test structure */
+static int test_initial_entries;
+static int test_max_entries;
+
+/* Whether to use SHMEM_ATTACH_UNKNOWN_SIZE when attaching to the shared memory */
+/* TODO: We may use opaque_arg to pass this value to the request function.*/
+static bool use_unknown_size = false;
+
+/*
+ * Request shared memory resources.
+ */
+static void
+resizable_shmem_request(void *arg)
+{
+ Size initial_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(test_initial_entries, TEST_ENTRY_SIZE));
+
+/*
+ * Create resizable structure on the platforms which support it. Otherwise create
+ * as a fixed-size structure. Other way would be to conditionally include
+ * .maximum_size in the call to ShmemRequestStruct().
+ */
+#ifdef HAVE_RESIZABLE_SHMEM
+ Size max_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(test_max_entries, TEST_ENTRY_SIZE));
+ Size min_size = offsetof(TestResizableShmemStruct, data);
+#else
+ Size max_size = 0;
+ Size min_size = 0;
+#endif
+
+ ShmemRequestStruct(.name = "resizable_shmem",
+ .size = use_unknown_size ? SHMEM_ATTACH_UNKNOWN_SIZE : initial_size,
+ .minimum_size = min_size,
+ .maximum_size = max_size,
+ .ptr = (void **) &resizable_shmem,
+ );
+}
+
+/*
+ * Initialize shared memory structure.
+ */
+static void
+resizable_shmem_shmem_init(void *arg)
+{
+ Assert(resizable_shmem != NULL);
+
+ resizable_shmem->num_entries = test_initial_entries;
+ memset(resizable_shmem->data, 0, mul_size(test_initial_entries, TEST_ENTRY_SIZE));
+}
+
+/*
+ * Attach to the already-allocated shared memory structure.
+ */
+static void
+resizable_shmem_shmem_attach(void *arg)
+{
+ Assert(resizable_shmem != NULL);
+}
+
+static ShmemCallbacks resizable_shmem_callbacks = {
+ .request_fn = resizable_shmem_request,
+ .init_fn = resizable_shmem_shmem_init,
+ .attach_fn = resizable_shmem_shmem_attach,
+};
+
+/*
+ * Resize the shared memory structure to accommodate the specified number of
+ * entries.
+ *
+ * Negative value for new_entries can be used to test resizing below the
+ * minimum size.
+ *
+ * Returns true if the resize was successful, false if ShmemResizeStruct()
+ * could not allocate the requested memory. On platforms that do not support
+ * resizable shared memory, ShmemResizeStruct() raises an error.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_resize);
+Datum
+resizable_shmem_resize(PG_FUNCTION_ARGS)
+{
+ int32 new_entries = PG_GETARG_INT32(0);
+ Size new_size;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (new_entries < 0)
+ new_size = 1;
+ else
+ new_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(new_entries, TEST_ENTRY_SIZE));
+ if (!ShmemResizeStruct("resizable_shmem", new_size))
+ PG_RETURN_BOOL(false);
+
+ ShmemProtectStruct("resizable_shmem");
+ resizable_shmem->num_entries = new_entries;
+
+ PG_RETURN_BOOL(true);
+}
+
+/*
+ * Write the given integer value to all entries in the data array.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_write);
+Datum
+resizable_shmem_write(PG_FUNCTION_ARGS)
+{
+ int32 entry_value = PG_GETARG_INT32(0);
+ int32 i;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (i = 0; i < resizable_shmem->num_entries; i++)
+ resizable_shmem->data[i] = entry_value;
+
+ PG_RETURN_VOID();
+}
+
+/*
+ * Check whether the first 'entry_count' entries all have the expected 'entry_value'.
+ * Returns true if all match, false otherwise.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_read);
+Datum
+resizable_shmem_read(PG_FUNCTION_ARGS)
+{
+ int32 entry_count = PG_GETARG_INT32(0);
+ int32 entry_value = PG_GETARG_INT32(1);
+ int32 i;
+
+ if (resizable_shmem == NULL)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (entry_count < 0 || entry_count > resizable_shmem->num_entries)
+ ereport(ERROR,
+ errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("entry_count %d is out of range (0..%d)", entry_count, resizable_shmem->num_entries));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (i = 0; i < entry_count; i++)
+ {
+ if (resizable_shmem->data[i] != entry_value)
+ PG_RETURN_BOOL(false);
+ }
+
+ PG_RETURN_BOOL(true);
+}
+
+/*
+ * Return the memory mapped against the main shared memory segment in this
+ * backend.
+ *
+ * The VMA containing our resizable_shmem pointer identifies the start of the
+ * main shared-memory segment.
+ *
+ * mprotect() calls issued when the resizable structure grows and shrinks can
+ * split the original mmap into several adjacent VMAs, so we sum the accounting
+ * fields across the base VMA and any VMAs contiguous with it.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_usage);
+Datum
+test_shmem_usage(PG_FUNCTION_ARGS)
+{
+ FILE *f;
+ char line[256];
+ uintptr_t target = (uintptr_t) resizable_shmem;
+ bool in_target_vma = false;
+ bool use_hugetlb = (huge_pages_status == HUGE_PAGES_ON);
+ unsigned long prev_end = 0;
+ int64 total_rss_kb = 0;
+ int64 total_swap_kb = 0;
+ int64 total_shared_hugetlb_kb = 0;
+ int64 val;
+ size_t result;
+
+ f = AllocateFile("/proc/self/smaps", "r");
+ if (f == NULL)
+ ereport(ERROR,
+ errcode_for_file_access(),
+ errmsg("could not open /proc/self/smaps: %m"));
+
+ while (fgets(line, sizeof(line), f) != NULL)
+ {
+ unsigned long start;
+ unsigned long end;
+
+ if (sscanf(line, "%lx-%lx", &start, &end) == 2)
+ {
+ if (in_target_vma)
+ {
+ /*
+ * Continue accumulating only across VMAs that are contiguous
+ * with the previous one; stop as soon as we hit a gap or a
+ * different mapping.
+ */
+ if (start != prev_end)
+ break;
+ }
+ else
+ in_target_vma = (target >= start && target < end);
+
+ prev_end = end;
+ }
+ else if (in_target_vma)
+ {
+ if (use_hugetlb)
+ {
+ if (sscanf(line, "Shared_Hugetlb: %ld kB", &val) == 1)
+ total_shared_hugetlb_kb += val;
+ }
+ else
+ {
+ if (sscanf(line, "Rss: %ld kB", &val) == 1)
+ total_rss_kb += val;
+ else if (sscanf(line, "Swap: %ld kB", &val) == 1)
+ total_swap_kb += val;
+ }
+ }
+ }
+
+ FreeFile(f);
+
+ if (use_hugetlb)
+ result = mul_size(total_shared_hugetlb_kb, 1024);
+ else
+ {
+ result = mul_size(total_rss_kb, 1024);
+ result = add_size(result, mul_size(total_swap_kb, 1024));
+ }
+
+ PG_RETURN_INT64(result);
+}
+
+/*
+ * Return the shared memory page size.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_pagesize);
+Datum
+test_shmem_pagesize(PG_FUNCTION_ARGS)
+{
+ PG_RETURN_INT32(pg_get_shmem_pagesize());
+}
+
+/*
+ * Walk the entries between the current size and the reserved maximum, accessing
+ * each one. Ideally, this function should (seg)fault the moment we try to access
+ * the entry outside the currently allocated size, but the memory allocation and
+ * protection mechanisms work on page basis. Hence it may only (seg)fault when a
+ * page boundary is crossed. The mode argument selects between "read" and
+ * "write" access.
+ *
+ * When the current end of the structure and end of maximal structure are on the
+ * same page, this function may not (seg)fault at all.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_access_beyond_size);
+Datum
+resizable_shmem_access_beyond_size(PG_FUNCTION_ARGS)
+{
+ text *mode_txt = PG_GETARG_TEXT_PP(0);
+ const char *mode = text_to_cstring(mode_txt);
+ bool do_write;
+ int32 sink = 0;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (strcmp(mode, "read") == 0)
+ do_write = false;
+ else if (strcmp(mode, "write") == 0)
+ do_write = true;
+ else
+ ereport(ERROR,
+ errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("mode must be \"read\" or \"write\""));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (int i = resizable_shmem->num_entries; i < test_max_entries; i++)
+ {
+ if (do_write)
+ resizable_shmem->data[i] = 0xdead;
+ else
+ sink = resizable_shmem->data[i];
+ }
+
+ /*
+ * Return the last read value so that compiler doesn't optimize away the
+ * assignment to sink.
+ */
+ PG_RETURN_INT32(sink);
+}
+
+
+/* ----------------------------------------------------------------
+ * Module initialization
+ * ----------------------------------------------------------------
+ */
+
+void
+_PG_init(void)
+{
+ int guc_context;
+
+ elog(LOG, "test_shmem module's _PG_init called");
+
+ RegisterShmemCallbacks(&TestShmemCallbacks);
+
+ /*
+ * Use PGC_POSTMASTER when loaded at startup so the values are fixed once
+ * the shared memory segment is created. When loaded after startup
+ * PGC_POSTMASTER is not allowed, so we use PGC_SIGHUP instead. Although
+ * we do not intend to change these values at config reload, PGC_SIGHUP is
+ * the least permissive context that allows defining the GUC after startup
+ * and still prevents it from being changed via SET.
+ */
+ if (process_shared_preload_libraries_in_progress)
+ guc_context = PGC_POSTMASTER;
+ else
+ {
+ guc_context = PGC_SIGHUP;
+ resizable_shmem_callbacks.flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP;
+ }
+
+ DefineCustomIntVariable("resizable_shmem.initial_entries",
+ "Initial number of entries in the test structure.",
+ NULL,
+ &test_initial_entries,
+ TEST_INITIAL_ENTRIES_DEFAULT,
+ 1,
+ INT_MAX,
+ guc_context,
+ 0,
+ NULL, NULL, NULL);
+
+ DefineCustomIntVariable("resizable_shmem.max_entries",
+ "Maximum number of entries in the test structure.",
+ NULL,
+ &test_max_entries,
+ TEST_MAX_ENTRIES_DEFAULT,
+ 1,
+ INT_MAX,
+ guc_context,
+ 0,
+ NULL, NULL, NULL);
+
+ /*
+ * When loaded after startup by a backend that is not creating the
+ * extension, the shared memory might have been resized to a size other
+ * than the initial size. Use SHMEM_ATTACH_UNKNOWN_SIZE to attach without
+ * knowing the exact size.
+ */
+ if (!process_shared_preload_libraries_in_progress && !creating_extension)
+ use_unknown_size = true;
+
+ RegisterShmemCallbacks(&resizable_shmem_callbacks);
+}
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 6a3341356da..f2c47ac1c28 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1770,8 +1770,11 @@ pg_shadow| SELECT pg_authid.rolname AS usename,
pg_shmem_allocations| SELECT name,
off,
size,
- allocated_size
- FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size);
+ allocated_size,
+ minimum_size,
+ maximum_size,
+ reserved_space
+ FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size, minimum_size, maximum_size, reserved_space);
pg_shmem_allocations_numa| SELECT name,
numa_node,
size
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 56c1f997f88..65f0fa4d5cb 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -3186,6 +3186,7 @@ TestDSMRegistryHashEntry
TestDSMRegistryStruct
TestDecodingData
TestDecodingTxnData
+TestResizableShmemStruct
TestShmemData
TestSpec
TestValueType
--
2.34.1
[text/x-patch] v20260724-0004-Pass-use_units-parameter-to-GucShowHook-fu.patch (15.6K, ../../CAExHW5tYUYY13b4S1fv-zdxNsEEJzTy1EW6mRUj+FFEpA4Q0=g@mail.gmail.com/6-v20260724-0004-Pass-use_units-parameter-to-GucShowHook-fu.patch)
download | inline diff:
From 7c01e2c39ca6f0d9edfe02a8f51cd2e4a3b452bc Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 25 Feb 2026 15:48:14 +0530
Subject: [PATCH v20260724 4/7] Pass use_units parameter to GucShowHook
functions
ShowGUCOption() prints the value of a GUC variable with units if the variable
has a unit and the caller passes use_units as true. However when a GUC variable
has a show_hook associated with it, the show_hook does not receive use_units
parameter. Hence a show_hook cannot determine whether to show the value with
units or not. This was not a problem until now because all the GUC variables
with show_hook were either didn't have any units associated with them or their
show_hook used a fixed unit to print the value of the variable.
shared_buffers is a GUC variable whose value is shows in different units based
on the GUC variable's value. With shared buffer pool resizing, we want to show
the pending size of the buffer pool, if any with appropriate units. We do this
using a show_hook, which needs use_units input. This commit adds use_units
parameter to GucShowHook and passes it from ShowGUCOption(). The actual commit
which uses this new parameter will be in a followup commit.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Reported by: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
---
src/backend/access/transam/xlog.c | 8 ++++----
src/backend/commands/variable.c | 14 +++++++-------
src/backend/executor/instrument.c | 2 +-
src/backend/libpq/pqcomm.c | 8 ++++----
src/backend/utils/misc/guc.c | 10 +++++-----
src/include/access/xlog.h | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 30 +++++++++++++++---------------
8 files changed, 38 insertions(+), 38 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 26e4929e167..a4057223cf7 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4975,7 +4975,7 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
/* guc hook */
const char *
-show_data_checksums(void)
+show_data_checksums(bool use_units)
{
return get_checksum_state_string(LocalDataChecksumState);
}
@@ -5211,7 +5211,7 @@ InitializeWalConsistencyChecking(void)
* GUC show_hook for archive_command
*/
const char *
-show_archive_command(void)
+show_archive_command(bool use_units)
{
if (XLogArchivingActive())
return XLogArchiveCommand;
@@ -5223,7 +5223,7 @@ show_archive_command(void)
* GUC show_hook for in_hot_standby
*/
const char *
-show_in_hot_standby(void)
+show_in_hot_standby(bool use_units)
{
/*
* We display the actual state based on shared memory, so that this GUC
@@ -5238,7 +5238,7 @@ show_in_hot_standby(void)
* GUC show_hook for effective_wal_level
*/
const char *
-show_effective_wal_level(void)
+show_effective_wal_level(bool use_units)
{
if (wal_level == WAL_LEVEL_MINIMAL)
return "minimal";
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 8afd252fc8c..491c8aa392d 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -389,7 +389,7 @@ assign_timezone(const char *newval, void *extra)
* show_timezone: GUC show_hook for timezone
*/
const char *
-show_timezone(void)
+show_timezone(bool use_units)
{
const char *tzn;
@@ -462,7 +462,7 @@ assign_log_timezone(const char *newval, void *extra)
* show_log_timezone: GUC show_hook for log_timezone
*/
const char *
-show_log_timezone(void)
+show_log_timezone(bool use_units)
{
const char *tzn;
@@ -674,7 +674,7 @@ assign_random_seed(double newval, void *extra)
}
const char *
-show_random_seed(void)
+show_random_seed(bool use_units)
{
return "unavailable";
}
@@ -1031,7 +1031,7 @@ assign_role(const char *newval, void *extra)
}
const char *
-show_role(void)
+show_role(bool use_units)
{
/*
* Check whether SET ROLE is active; if not return "none". This is a
@@ -1179,7 +1179,7 @@ assign_io_combine_limit(int newval, void *extra)
* GUC show_hook for data_directory_mode
*/
const char *
-show_data_directory_mode(void)
+show_data_directory_mode(bool use_units)
{
static char buf[12];
@@ -1191,7 +1191,7 @@ show_data_directory_mode(void)
* GUC show_hook for log_file_mode
*/
const char *
-show_log_file_mode(void)
+show_log_file_mode(bool use_units)
{
static char buf[12];
@@ -1203,7 +1203,7 @@ show_log_file_mode(void)
* GUC show_hook for unix_socket_permissions
*/
const char *
-show_unix_socket_permissions(void)
+show_unix_socket_permissions(bool use_units)
{
static char buf[12];
diff --git a/src/backend/executor/instrument.c b/src/backend/executor/instrument.c
index ffbcd572133..d4c8373336a 100644
--- a/src/backend/executor/instrument.c
+++ b/src/backend/executor/instrument.c
@@ -423,7 +423,7 @@ assign_timing_clock_source(int newval, void *extra)
}
const char *
-show_timing_clock_source(void)
+show_timing_clock_source(bool use_units)
{
switch (timing_clock_source)
{
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index aaae7214f13..b69bff03f09 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1972,7 +1972,7 @@ assign_tcp_keepalives_idle(int newval, void *extra)
* GUC show_hook for tcp_keepalives_idle
*/
const char *
-show_tcp_keepalives_idle(void)
+show_tcp_keepalives_idle(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -1995,7 +1995,7 @@ assign_tcp_keepalives_interval(int newval, void *extra)
* GUC show_hook for tcp_keepalives_interval
*/
const char *
-show_tcp_keepalives_interval(void)
+show_tcp_keepalives_interval(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -2018,7 +2018,7 @@ assign_tcp_keepalives_count(int newval, void *extra)
* GUC show_hook for tcp_keepalives_count
*/
const char *
-show_tcp_keepalives_count(void)
+show_tcp_keepalives_count(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -2041,7 +2041,7 @@ assign_tcp_user_timeout(int newval, void *extra)
* GUC show_hook for tcp_user_timeout
*/
const char *
-show_tcp_user_timeout(void)
+show_tcp_user_timeout(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 774bbc9be5f..1a5a168bc1a 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -5381,7 +5381,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_bool *conf = &record->_bool;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
val = *conf->variable ? "on" : "off";
}
@@ -5392,7 +5392,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_int *conf = &record->_int;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
{
/*
@@ -5421,7 +5421,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_real *conf = &record->_real;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
{
double result = *conf->variable;
@@ -5446,7 +5446,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_string *conf = &record->_string;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else if (*conf->variable && **conf->variable)
val = *conf->variable;
else
@@ -5459,7 +5459,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_enum *conf = &record->_enum;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
val = config_enum_lookup_by_value(record, *conf->variable);
}
diff --git a/src/include/access/xlog.h b/src/include/access/xlog.h
index 4dd98624204..ee15b4cbb16 100644
--- a/src/include/access/xlog.h
+++ b/src/include/access/xlog.h
@@ -255,7 +255,7 @@ extern bool DataChecksumsInProgressOn(void);
extern void SetDataChecksumsOnInProgress(void);
extern void SetDataChecksumsOn(void);
extern void SetDataChecksumsOff(void);
-extern const char *show_data_checksums(void);
+extern const char *show_data_checksums(bool use_units);
extern const char *get_checksum_state_string(uint32 state);
extern void InitLocalDataChecksumState(void);
extern void SetLocalDataChecksumState(uint32 data_checksum_version);
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 8057d7870ad..2a6e2ed18b3 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -192,7 +192,7 @@ typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
-typedef const char *(*GucShowHook) (void);
+typedef const char *(*GucShowHook) (bool use_units);
/*
* Miscellaneous
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 6a76f8d5ed6..df048517a0f 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -28,7 +28,7 @@
extern bool check_application_name(char **newval, void **extra,
GucSource source);
extern void assign_application_name(const char *newval, void *extra);
-extern const char *show_archive_command(void);
+extern const char *show_archive_command(bool use_units);
extern bool check_autovacuum_work_mem(int *newval, void **extra,
GucSource source);
extern bool check_vacuum_buffer_usage_limit(int *newval, void **extra,
@@ -46,7 +46,7 @@ extern void assign_client_encoding(const char *newval, void *extra);
extern bool check_cluster_name(char **newval, void **extra, GucSource source);
extern bool check_commit_ts_buffers(int *newval, void **extra,
GucSource source);
-extern const char *show_data_directory_mode(void);
+extern const char *show_data_directory_mode(bool use_units);
extern bool check_datestyle(char **newval, void **extra, GucSource source);
extern void assign_datestyle(const char *newval, void *extra);
extern bool check_debug_io_direct(char **newval, void **extra, GucSource source);
@@ -61,11 +61,11 @@ extern bool check_default_text_search_config(char **newval, void **extra, GucSou
extern void assign_default_text_search_config(const char *newval, void *extra);
extern bool check_default_with_oids(bool *newval, void **extra,
GucSource source);
-extern const char *show_effective_wal_level(void);
+extern const char *show_effective_wal_level(bool use_units);
extern bool check_huge_page_size(int *newval, void **extra, GucSource source);
extern void assign_io_method(int newval, void *extra);
extern bool check_io_max_concurrency(int *newval, void **extra, GucSource source);
-extern const char *show_in_hot_standby(void);
+extern const char *show_in_hot_standby(bool use_units);
extern bool check_locale_messages(char **newval, void **extra, GucSource source);
extern void assign_locale_messages(const char *newval, void *extra);
extern bool check_locale_monetary(char **newval, void **extra, GucSource source);
@@ -77,11 +77,11 @@ extern void assign_locale_time(const char *newval, void *extra);
extern bool check_log_destination(char **newval, void **extra,
GucSource source);
extern void assign_log_destination(const char *newval, void *extra);
-extern const char *show_log_file_mode(void);
+extern const char *show_log_file_mode(bool use_units);
extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
-extern const char *show_log_timezone(void);
+extern const char *show_log_timezone(bool use_units);
extern void assign_maintenance_io_concurrency(int newval, void *extra);
extern void assign_io_max_combine_limit(int newval, void *extra);
extern void assign_io_combine_limit(int newval, void *extra);
@@ -97,7 +97,7 @@ extern bool check_primary_slot_name(char **newval, void **extra,
GucSource source);
extern bool check_random_seed(double *newval, void **extra, GucSource source);
extern void assign_random_seed(double newval, void *extra);
-extern const char *show_random_seed(void);
+extern const char *show_random_seed(bool use_units);
extern bool check_recovery_prefetch(int *new_value, void **extra,
GucSource source);
extern void assign_recovery_prefetch(int new_value, void *extra);
@@ -118,7 +118,7 @@ extern bool check_recovery_target_xid(char **newval, void **extra,
extern void assign_recovery_target_xid(const char *newval, void *extra);
extern bool check_role(char **newval, void **extra, GucSource source);
extern void assign_role(const char *newval, void *extra);
-extern const char *show_role(void);
+extern const char *show_role(bool use_units);
extern bool check_restrict_nonsystem_relation_kind(char **newval, void **extra,
GucSource source);
extern void assign_restrict_nonsystem_relation_kind(const char *newval, void *extra);
@@ -143,32 +143,32 @@ extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
extern void assign_tcp_keepalives_count(int newval, void *extra);
-extern const char *show_tcp_keepalives_count(void);
+extern const char *show_tcp_keepalives_count(bool use_units);
extern void assign_tcp_keepalives_idle(int newval, void *extra);
-extern const char *show_tcp_keepalives_idle(void);
+extern const char *show_tcp_keepalives_idle(bool use_units);
extern void assign_tcp_keepalives_interval(int newval, void *extra);
-extern const char *show_tcp_keepalives_interval(void);
+extern const char *show_tcp_keepalives_interval(bool use_units);
extern void assign_tcp_user_timeout(int newval, void *extra);
-extern const char *show_tcp_user_timeout(void);
+extern const char *show_tcp_user_timeout(bool use_units);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
GucSource source);
extern void assign_temp_tablespaces(const char *newval, void *extra);
extern bool check_timezone(char **newval, void **extra, GucSource source);
extern void assign_timezone(const char *newval, void *extra);
-extern const char *show_timezone(void);
+extern const char *show_timezone(bool use_units);
extern bool check_timezone_abbreviations(char **newval, void **extra,
GucSource source);
extern void assign_timezone_abbreviations(const char *newval, void *extra);
extern void assign_timing_clock_source(int newval, void *extra);
extern bool check_timing_clock_source(int *newval, void **extra, GucSource source);
-extern const char *show_timing_clock_source(void);
+extern const char *show_timing_clock_source(bool use_units);
extern bool check_transaction_buffers(int *newval, void **extra, GucSource source);
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
extern void assign_transaction_timeout(int newval, void *extra);
-extern const char *show_unix_socket_permissions(void);
+extern const char *show_unix_socket_permissions(bool use_units);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
GucSource source);
--
2.34.1
[text/x-patch] v20260724-0002-Decouple-GUC-shared_buffers-and-size-of-th.patch (18.6K, ../../CAExHW5tYUYY13b4S1fv-zdxNsEEJzTy1EW6mRUj+FFEpA4Q0=g@mail.gmail.com/7-v20260724-0002-Decouple-GUC-shared_buffers-and-size-of-th.patch)
download | inline diff:
From cb99234f9d5f708eba068867aea156af12280be9 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 4 Jun 2026 15:10:56 +0530
Subject: [PATCH v20260724 2/7] Decouple GUC shared_buffers and size of the
buffer pool
A single variable NBuffers holds the value of the GUC 'shared_buffers' and the
size of the buffer pool in number of buffers. With the introduction of resizable
shared buffer pool feature, the value of 'shared_buffers' GUC and the size of
the buffer pool can be different. This commit prepares for the same by
decoupling the two. A new variable NBuffersGUC holds the value of the GUC
'shared_buffers'. The variable NBuffers continues to hold the size of the
buffer pool. It is set to the value of NBuffersGUC during initialization.
Because of this decoupling, the code which references NBuffers as the value of
the GUC 'shared_buffers' is clearly differentiated from the code which
references NBuffers as the size of the buffer pool.
Also change comments to use term "size of the buffer pool" or "number of
buffers" instead of referencing variable NBuffers.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
src/backend/access/heap/heapam.c | 15 ++++++++-------
src/backend/access/transam/slru.c | 2 +-
src/backend/access/transam/xlog.c | 8 ++++----
src/backend/optimizer/path/costsize.c | 8 ++++----
src/backend/postmaster/checkpointer.c | 11 ++++++-----
src/backend/storage/aio/aio_init.c | 2 +-
src/backend/storage/buffer/buf_init.c | 17 +++++++++++++----
src/backend/storage/buffer/buf_table.c | 14 ++++++++------
src/backend/storage/buffer/bufmgr.c | 13 +++++++------
src/backend/storage/buffer/freelist.c | 7 ++++---
src/backend/utils/init/globals.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 2 +-
src/include/miscadmin.h | 2 +-
src/include/storage/buf.h | 2 +-
src/include/storage/buf_internals.h | 4 ++--
src/include/storage/bufmgr.h | 3 ++-
16 files changed, 64 insertions(+), 48 deletions(-)
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 8b488cfd8f6..8eb84bd51d6 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -383,13 +383,14 @@ initscan(HeapScanDesc scan, ScanKey key, bool keep_startblock)
scan->rs_nblocks = RelationGetNumberOfBlocks(scan->rs_base.rs_rd);
/*
- * If the table is large relative to NBuffers, use a bulk-read access
- * strategy and enable synchronized scanning (see syncscan.c). Although
- * the thresholds for these features could be different, we make them the
- * same so that there are only two behaviors to tune rather than four.
- * (However, some callers need to be able to disable one or both of these
- * behaviors, independently of the size of the table; also there is a GUC
- * variable that can disable synchronized scanning.)
+ * If the table is large relative to the size of the buffer pool, use a
+ * bulk-read access strategy and enable synchronized scanning (see
+ * syncscan.c). Although the thresholds for these features could be
+ * different, we make them the same so that there are only two behaviors
+ * to tune rather than four. (However, some callers need to be able to
+ * disable one or both of these behaviors, independently of the size of
+ * the table; also there is a GUC variable that can disable synchronized
+ * scanning.)
*
* Note that table_block_parallelscan_initialize has a very similar test;
* if you change this, consider changing that one, too.
diff --git a/src/backend/access/transam/slru.c b/src/backend/access/transam/slru.c
index 47dd52d6749..b47ace37974 100644
--- a/src/backend/access/transam/slru.c
+++ b/src/backend/access/transam/slru.c
@@ -236,7 +236,7 @@ SimpleLruAutotuneBuffers(int divisor, int max)
{
return Min(max - (max % SLRU_BANK_SIZE),
Max(SLRU_BANK_SIZE,
- NBuffers / divisor - (NBuffers / divisor) % SLRU_BANK_SIZE));
+ NBuffersGUC / divisor - (NBuffersGUC / divisor) % SLRU_BANK_SIZE));
}
/*
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index f8b939853e9..26e4929e167 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -5017,14 +5017,14 @@ GetFakeLSNForUnloggedRel(void)
* and a minimum of 8 blocks (which was the default value prior to PostgreSQL
* 9.1, when auto-tuning was added).
*
- * This should not be called until NBuffers has received its final value.
+ * This should not be called until NBuffersGUC has received its final value.
*/
static int
XLOGChooseNumBuffers(void)
{
int xbuffers;
- xbuffers = NBuffers / 32;
+ xbuffers = NBuffersGUC / 32;
if (xbuffers > (wal_segment_size / XLOG_BLCKSZ))
xbuffers = (wal_segment_size / XLOG_BLCKSZ);
if (xbuffers < 8)
@@ -5297,8 +5297,8 @@ XLOGShmemRequest(void *arg)
/*
* If the value of wal_buffers is -1, use the preferred auto-tune value.
* This isn't an amazingly clean place to do this, but we must wait till
- * NBuffers has received its final value, and must do it before using the
- * value of XLOGbuffers to do anything important.
+ * NBuffersGUC has received its final value, and must do it before using
+ * the value of XLOGbuffers to do anything important.
*
* We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
* However, if the DBA explicitly set wal_buffers = -1 in the config file,
diff --git a/src/backend/optimizer/path/costsize.c b/src/backend/optimizer/path/costsize.c
index ac523ecf9a8..8be50631e9c 100644
--- a/src/backend/optimizer/path/costsize.c
+++ b/src/backend/optimizer/path/costsize.c
@@ -19,10 +19,10 @@
* is normally considerably less than random_page_cost. (However, if the
* database is fully cached in RAM, it is reasonable to set them equal.)
*
- * We also use a rough estimate "effective_cache_size" of the number of
- * disk pages in Postgres + OS-level disk cache. (We can't simply use
- * NBuffers for this purpose because that would ignore the effects of
- * the kernel's disk cache.)
+ * We also use a rough estimate "effective_cache_size" of the number of disk
+ * pages in Postgres + OS-level disk cache. (We can't simply use size of the
+ * buffer pool for this purpose because that would ignore the effects of the
+ * kernel's disk cache.)
*
* Obviously, taking constants for these values is an oversimplification,
* but it's tough enough to get any useful estimates even at this level of
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index f6351f1eb10..0db3c9e1a38 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -963,12 +963,13 @@ CheckpointerShmemRequest(void *arg)
Size size;
/*
- * The size of the requests[] array is arbitrarily set equal to NBuffers.
- * But there is a cap of MAX_CHECKPOINT_REQUESTS to prevent accumulating
- * too many checkpoint requests in the ring buffer.
+ * The size of the requests[] array is arbitrarily set equal to the
+ * initial size of buffer pool. But there is a cap of
+ * MAX_CHECKPOINT_REQUESTS to prevent accumulating too many checkpoint
+ * requests in the ring buffer.
*/
size = offsetof(CheckpointerShmemStruct, requests);
- size = add_size(size, mul_size(Min(NBuffers,
+ size = add_size(size, mul_size(Min(NBuffersGUC,
MAX_CHECKPOINT_REQUESTS),
sizeof(CheckpointerRequest)));
ShmemRequestStruct(.name = "Checkpointer Data",
@@ -985,7 +986,7 @@ static void
CheckpointerShmemInit(void *arg)
{
SpinLockInit(&CheckpointerShmem->ckpt_lck);
- CheckpointerShmem->max_requests = Min(NBuffers, MAX_CHECKPOINT_REQUESTS);
+ CheckpointerShmem->max_requests = Min(NBuffersGUC, MAX_CHECKPOINT_REQUESTS);
CheckpointerShmem->head = CheckpointerShmem->tail = 0;
ConditionVariableInit(&CheckpointerShmem->start_cv);
ConditionVariableInit(&CheckpointerShmem->done_cv);
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index de50e6a8a31..81825dd1b73 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -109,7 +109,7 @@ AioChooseMaxConcurrency(void)
/* Similar logic to LimitAdditionalPins() */
max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
- max_proportional_pins = NBuffers / max_backends;
+ max_proportional_pins = NBuffersGUC / max_backends;
max_proportional_pins = Max(max_proportional_pins, 1);
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 1407c930c56..9ddf6551fcd 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -77,21 +77,21 @@ static void
BufferManagerShmemRequest(void *arg)
{
ShmemRequestStruct(.name = "Buffer Descriptors",
- .size = NBuffers * sizeof(BufferDescPadded),
+ .size = NBuffersGUC * sizeof(BufferDescPadded),
/* Align descriptors to a cacheline boundary. */
.alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &BufferDescriptors,
);
ShmemRequestStruct(.name = "Buffer Blocks",
- .size = NBuffers * (Size) BLCKSZ,
+ .size = NBuffersGUC * (Size) BLCKSZ,
/* Align buffer pool on IO page size boundary. */
.alignment = PG_IO_ALIGN_SIZE,
.ptr = (void **) &BufferBlocks,
);
ShmemRequestStruct(.name = "Buffer IO Condition Variables",
- .size = NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ .size = NBuffersGUC * sizeof(ConditionVariableMinimallyPadded),
/* Align descriptors to a cacheline boundary. */
.alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &BufferIOCVArray,
@@ -105,7 +105,7 @@ BufferManagerShmemRequest(void *arg)
* painful.
*/
ShmemRequestStruct(.name = "Checkpoint BufferIds",
- .size = NBuffers * sizeof(CkptSortItem),
+ .size = NBuffersGUC * sizeof(CkptSortItem),
.ptr = (void **) &CkptBufferIds,
);
}
@@ -119,6 +119,12 @@ BufferManagerShmemRequest(void *arg)
static void
BufferManagerShmemInit(void *arg)
{
+ /*
+ * Set the size of the buffer pool, now that it's allocated and ready to
+ * be initialized.
+ */
+ NBuffers = NBuffersGUC;
+
/*
* Initialize all the buffer headers.
*/
@@ -147,6 +153,9 @@ BufferManagerShmemInit(void *arg)
static void
BufferManagerShmemAttach(void *arg)
{
+ /* Update the size of the buffer pool. */
+ NBuffers = NBuffersGUC;
+
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
&backend_flush_after);
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 347bf267d73..5c8ccbee13f 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -42,7 +42,7 @@ const ShmemCallbacks BufTableShmemCallbacks = {
/*
* Register shmem hash table for mapping buffers.
- * size is the desired hash table size (possibly more than NBuffers)
+ * size is the desired hash table size (possibly more than the size of the buffer pool).
*/
void
BufTableShmemRequest(void *arg)
@@ -54,12 +54,14 @@ BufTableShmemRequest(void *arg)
*
* Since we can't tolerate running out of lookup table entries, we must be
* sure to specify an adequate table size here. The maximum steady-state
- * usage is of course NBuffers entries, but BufferAlloc() tries to insert
- * a new entry before deleting the old. In principle this could be
- * happening in each partition concurrently, so we could need as many as
- * NBuffers + NUM_BUFFER_PARTITIONS entries.
+ * usage is of course as many entries as the number of buffers in the
+ * pool, but BufferAlloc() tries to insert a new entry before deleting the
+ * old. In principle this could be happening in each partition
+ * concurrently, so we could need as many as (number of buffers in the
+ * pool) + NUM_BUFFER_PARTITIONS entries. Since we are still requesting
+ * shared memory, use the GUC value instead of the actual size.
*/
- size = NBuffers + NUM_BUFFER_PARTITIONS;
+ size = NBuffersGUC + NUM_BUFFER_PARTITIONS;
ShmemRequestHash(.name = "Shared Buffer Lookup Table",
.nelems = size,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 3eea17115b5..db7f34d2274 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -223,6 +223,7 @@ int io_max_combine_limit = DEFAULT_IO_COMBINE_LIMIT;
int checkpoint_flush_after = DEFAULT_CHECKPOINT_FLUSH_AFTER;
int bgwriter_flush_after = DEFAULT_BGWRITER_FLUSH_AFTER;
int backend_flush_after = DEFAULT_BACKEND_FLUSH_AFTER;
+int NBuffers = 0; /* number of buffers in the buffer pool */
/* local state for LockBufferForCleanup */
static BufferDesc *PinCountWaitBuf = NULL;
@@ -239,11 +240,11 @@ static BufferDesc *PinCountWaitBuf = NULL;
* and, if so, in what mode.
*
*
- * To avoid - as we used to - requiring an array with NBuffers entries to keep
- * track of local buffers, we use a small sequentially searched array
- * (PrivateRefCountArrayKeys, with the corresponding data stored in
- * PrivateRefCountArray) and an overflow hash table (PrivateRefCountHash) to
- * keep track of backend local pins.
+ * To avoid - as we used to - requiring an array, with as many entries as the
+ * size of buffer pool, to keep track of local buffers, we use a small
+ * sequentially searched array (PrivateRefCountArrayKeys, with the corresponding
+ * data stored in PrivateRefCountArray) and an overflow hash table
+ * (PrivateRefCountHash) to keep track of backend local pins.
*
* Until no more than REFCOUNT_ARRAY_ENTRIES buffers are pinned at once, all
* refcounts are kept track of in the array; after that, new array entries
@@ -3642,7 +3643,7 @@ BufferSync(int flags)
set_bits, 0,
0);
- /* Check for barrier events in case NBuffers is large. */
+ /* Check for barrier events in case the buffer pool is large. */
if (ProcSignalBarrierPending)
ProcessProcSignalBarrier();
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index fdb5bad7910..4d5ee52ddc0 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -37,7 +37,8 @@ typedef struct
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo NBuffers.
+ * get an actual buffer, it needs to be used modulo size of the buffer
+ * pool.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -522,10 +523,10 @@ GetAccessStrategyWithSize(BufferAccessStrategyType btype, int ring_size_kb)
if (ring_buffers == 0)
return NULL;
- /* Cap to 1/8th of shared_buffers */
+ /* Cap to 1/8th of number of buffers in the buffer pool. */
ring_buffers = Min(NBuffers / 8, ring_buffers);
- /* NBuffers should never be less than 16, so this shouldn't happen */
+ /* Buffer pool should always have more than 16 buffers. */
Assert(ring_buffers > 0);
/* Allocate the object and initialize all elements to zeroes */
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index bbd28d14d99..ccf845e87b9 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -141,7 +141,7 @@ int max_parallel_maintenance_workers = 2;
* MaxBackends is computed by PostmasterMain after modules have had a chance to
* register background workers.
*/
-int NBuffers = 16384;
+int NBuffersGUC = 16384;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index adb72361ce0..dccc4c82507 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2716,7 +2716,7 @@
{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
- variable => 'NBuffers',
+ variable => 'NBuffersGUC',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 7170a4bff98..b9449e9aced 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -175,7 +175,7 @@ extern PGDLLIMPORT bool ExitOnAnyError;
extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
-extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersGUC;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/storage/buf.h b/src/include/storage/buf.h
index b21445522b1..a12d6b9082a 100644
--- a/src/include/storage/buf.h
+++ b/src/include/storage/buf.h
@@ -17,7 +17,7 @@
/*
* Buffer identifiers.
*
- * Zero is invalid, positive is the index of a shared buffer (1..NBuffers),
+ * Zero is invalid, positive is the index of a shared buffer (1..{size of shared buffer pool}),
* negative is the index of a local buffer (-1 .. -NLocBuffer).
*/
typedef int Buffer;
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index e4ff5619b79..a7606c6e92b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -135,8 +135,8 @@ StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2),
/*
* The maximum allowed value of usage_count represents a tradeoff between
- * accuracy and speed of the clock-sweep buffer management algorithm. A
- * large value (comparable to NBuffers) would approximate LRU semantics.
+ * accuracy and speed of the clock-sweep buffer management algorithm. A large
+ * value (comparable to the size of buffer pool) would approximate LRU semantics.
* But it can take as many as BM_MAX_USAGE_COUNT+1 complete cycles of the
* clock-sweep hand to find a free buffer, so in practice we don't want the
* value to be very large.
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 6837b35fc6d..f1f6e601f51 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -159,13 +159,14 @@ typedef struct ReadBuffersOperation ReadBuffersOperation;
typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
-extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersGUC;
/* in bufmgr.c */
extern PGDLLIMPORT bool zero_damaged_pages;
extern PGDLLIMPORT int bgwriter_lru_maxpages;
extern PGDLLIMPORT double bgwriter_lru_multiplier;
extern PGDLLIMPORT bool track_io_timing;
+extern PGDLLIMPORT int NBuffers;
#define DEFAULT_EFFECTIVE_IO_CONCURRENCY 16
#define DEFAULT_MAINTENANCE_IO_CONCURRENCY 16
--
2.34.1
[text/x-patch] v20260724-0001-Add-BgBufferSync-sanity-Asserts.patch (2.1K, ../../CAExHW5tYUYY13b4S1fv-zdxNsEEJzTy1EW6mRUj+FFEpA4Q0=g@mail.gmail.com/8-v20260724-0001-Add-BgBufferSync-sanity-Asserts.patch)
download | inline diff:
From f4f64eeb577cb469941b0a4dc0b9cf8d29147b36 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 21 Jul 2026 16:16:48 +0530
Subject: [PATCH v20260724 1/7] Add BgBufferSync sanity Asserts
Assert that BgBufferSync() is only ever invoked from the background
writer process, matching the actual caller in BackgroundWriterMain.
Also move the existing Assert(strategy_delta >= 0) to the end of the
enclosing block so that the surrounding elog(DEBUG2) messages get a
chance to reach the server log before the Assert fires, easing diagnosis
if the Assert ever fails.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
src/backend/storage/buffer/bufmgr.c | 15 +++++++++++++--
1 file changed, 13 insertions(+), 2 deletions(-)
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 3908529872a..3eea17115b5 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3894,6 +3894,8 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ Assert(AmBackgroundWriterProcess());
+
/*
* Find out where the clock-sweep currently is, and how many buffer
* allocations have happened since our last call.
@@ -3929,8 +3931,6 @@ BgBufferSync(WritebackContext *wb_context)
strategy_delta = strategy_buf_id - prev_strategy_buf_id;
strategy_delta += (long) passes_delta * NBuffers;
- Assert(strategy_delta >= 0);
-
if ((int32) (next_passes - strategy_passes) > 0)
{
/* we're one pass ahead of the strategy point */
@@ -3970,6 +3970,17 @@ BgBufferSync(WritebackContext *wb_context)
next_passes = strategy_passes;
bufs_to_lap = NBuffers;
}
+
+ /*
+ * We do not expect the current strategy point to be behind the
+ * previous one.
+ *
+ * If this Assert fails, we would have the debug messages printed in
+ * the server error log at appropriate debug level. Hence Asserting
+ * here provides minor convenience compared to Asserting right after
+ * calculating the difference.
+ */
+ Assert(strategy_delta >= 0);
}
else
{
base-commit: b77868f169adcdf31edbc80d8a875204ed7ba191
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-08-17 11:56 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 14:26 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2026-08-17 11:56 UTC (permalink / raw)
To: pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; +Cc: Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
On Fri, Jul 24, 2026 at 6:26 PM Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> Hi,
>
> On Fri, Feb 13, 2026 at 5:22 PM Ashutosh Bapat
> <ashutosh.bapat.oss@gmail.com> wrote:
> >
> > On Thu, Feb 12, 2026 at 7:43 PM Jakub Wartak
> > <jakub.wartak@enterprisedb.com> wrote:
> > >
> > >
> > > TBH, I haven't really looked at the code outside of that region, I'm just
> > > trespasser that was interested in memfd ;)
> >
> > Your trespassing has been very helpful. I have started a separate
> > thread to discuss resizable shared structures at [1]. Once the
> > implementation there is somewhat finalized, it will be good to try
> > your huge page tests again.
>
> Here's the next version of the patch implementing shared buffer pool
> resizing. The patch is based on the latest master. Here's the summary
> of changes since the last version:
>
> 1. The patch now uses the new shared memory infrastructure that was
> introduced in PG 19. Patch 0006 enhances that infrastructure to
> support resizable shared structures. I will also post the same patch
> to [1]. I am fine to discuss the patch in that thread or here. The
> APIs for registering and resizing the structures are documented in the
> programming interface documentation.
Palak has provided an incremental patch fixing CI failures. It needs
to be reviewed, hence not a part of this patchset.
>
> 2. Patch 0007 implements the shared buffer pool resizing using the
> resizable shared structures infrastructure. It has a lot of code
> improvements, including better documentation in comments, READMEs,
> user-facing documentation and more TAP tests. The
> storage/buffer/README has a section on buffer resizing. The buffer
> resizing is implemented in buf_resize.c, which also has detailed
> comments about the implementation. I suggest starting the review with
> the user documentation, README and buf_resize.c.
>
The attached patches have a major change in this patch: Stress tests.
I have added one stress test for every part of the code which scans
the buffer to stress exercise that portion of code again buffer pool
resizing. The patch also contains fixes for the crashes or bugs
revealed by the stress tests. Specifically the stress tests cover
synchronization between buffer pool resizing and
DropRelationBuffers(), DropRelationsAllBuffers(),
DropDatabaseBuffers(), CHECKPOINT, FlushRelationBuffers(),
FlushRelationsAllBuffers(), pg_prewarm, monitoring and diagnostic
functions in pg_buffercache. All these tests share common utility code
in test/buffermgr/StressUtility.pm. At the end of each stress run, it
carries out sanity checks to make sure that the database is not
corrupted, the shared buffer pool state is in a sane state etc.
FlushDatabaseBuffers() is not covered by any stress test since it is
only called during WAL replay of xl_dbase_create_file_copy_rec. The
primary CREATE DATABASE and ALTER DATABASE SET TABLESPACE paths use
RequestCheckpoint() instead. I could not find a way to reach
FlushDatabaseBuffers through normal SQL workload. But the function
should be able to cope with the resized buffer pool just like other
functions which scan the buffer pool.
These tests are run only when PG_TEST_EXTRA has bufmgr_stress in it
since these tests run longer (2 minutes each) and use many resources.
We may not want to accept all these stress tests necessarily. We may
want to pack all of them into a single test or just not accept any of
them. They are pretty useful to build confidence that the reisizing
protocol, shadow variables and barriers are working correctly and are
hazard free. I would like to keep these tests in the patchset as long
as possible and remove them just before the final commit to keep that
confidence as we change the code and protocol while responding to the
review comments. We may add more deterministic white box tests, like
001 and 002, using injection points for specific hazardous scenarios.
Following bugs/crashes were revealed by the stress tests and their
fixes (except one) are included in the patch.
1. BufferSync() is fixed to clean up a buffer-invalidated-by-resize
properly from the checkpointer datastructures.
2. Most of the loops scanning the buffer pool invoked CFI at the
beginning of the loop, which meant that the buffer being processed can
be invalidated right at the beginning of each iteration. Instead moved
CFI calls to the end of the loop.
3. Fixed pg_prewarm to not rely on NBuffers being static always,
instead it adapts to the new size after CFI. But possibly we could
change the function to process and write one buffer's tag at a time. I
think we need a separate discussion for this.
pg_buffercache_os_pages() still crashes when it hits a concurrent
resize. But the fix is already being written.
Changed pg_resize_shared_buffers to throw an error when the server
does not support resizable shared memory structures. Also changed all
the resize related tests to be skipped when the server does not
support resizable shared memory structures.
> 0002 - Decouples the use of the NBuffers variable as a GUC from its
> use as the size of the shared buffer pool. Prepares for the buffer
> pool size to differ from the GUC value, as required by the resizing
> feature.
In the attached 0002, I have fixed some places that were left out in
the previous version.
All other patches are the same as the corresponding patches in the
previous version.
There are still TODOs in the patch, which I will address in the next
few versions.
It will be good to review the resizing protocol and use of
ProcSignalBarrier to keep the process local buffer manager state
(shadow variables NBuffers, activeNBuffers and the address map
protection) in sync across all the backends. Once there's an agreement
over that, I can address more detailed TODOs but most importantly, I
will be able to implement the ability to roll back an interrupted
resize operation. That ability depends upon the protocol and the
barrier mechanism.
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-patch] v20260817-0003-Add-a-view-to-read-contents-of-shared-buff.patch (14.6K, ../../CAExHW5ts93Rnof7pjFFYrY9aTOmCo+xsdQazqeqN1BaKozTBvA@mail.gmail.com/2-v20260817-0003-Add-a-view-to-read-contents-of-shared-buff.patch)
download | inline diff:
From 4ad4cb2765972379433eb8b7b738c748522295d6 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 25 Aug 2025 19:23:50 +0530
Subject: [PATCH v20260817 3/7] Add a view to read contents of shared buffer
lookup table
The view exposes the contents of the shared buffer lookup table for
debugging, testing and investigation.
This helped me in debugging issues where the buffer descriptor array and
buffer lookup table were out of sync; either the buffer lookup table had
a mapping page->buffer which wasn't present in the buffer descriptor
array or a page in the buffer descriptor array didn't have corresponding
entry in the buffer lookup table. pg_buffercache doesn't help with those
kind of issues. Also doing that under the debugger in very painful.
I intend to keep this patch while the rest of the code matures. If it is
found useful as a debugging tool, we may consider make it committable
and commit it.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../expected/pg_buffercache.out | 39 ++++++++
.../pg_buffercache--1.5--1.6.sql | 24 +++++
contrib/pg_buffercache/pg_buffercache_pages.c | 17 ++++
contrib/pg_buffercache/sql/pg_buffercache.sql | 20 +++++
doc/src/sgml/system-views.sgml | 89 +++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 58 ++++++++++++
src/include/storage/buf_internals.h | 2 +
7 files changed, 249 insertions(+)
diff --git a/contrib/pg_buffercache/expected/pg_buffercache.out b/contrib/pg_buffercache/expected/pg_buffercache.out
index c52a8491ff9..1ce5a970334 100644
--- a/contrib/pg_buffercache/expected/pg_buffercache.out
+++ b/contrib/pg_buffercache/expected/pg_buffercache.out
@@ -33,6 +33,26 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
t
(1 row)
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -46,6 +66,10 @@ SELECT * FROM pg_buffercache_summary();
ERROR: permission denied for function pg_buffercache_summary
SELECT * FROM pg_buffercache_usage_counts();
ERROR: permission denied for function pg_buffercache_usage_counts
+SELECT * FROM pg_buffercache_lookup_table_entries();
+ERROR: permission denied for function pg_buffercache_lookup_table_entries
+SELECT * FROM pg_buffercache_lookup_table;
+ERROR: permission denied for view pg_buffercache_lookup_table
RESET role;
-- Check that pg_monitor is allowed to query view / function
SET ROLE pg_monitor;
@@ -81,6 +105,21 @@ FROM pg_buffercache_pages() AS p
LIMIT 1;
ERROR: function return row and query-specified return row do not match
DETAIL: Returned type boolean at ordinal position 7, but query expects text.
+RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
RESET role;
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
index 458f054a691..9bf58567878 100644
--- a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
+++ b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
@@ -44,3 +44,27 @@ CREATE FUNCTION pg_buffercache_evict_all(
OUT buffers_skipped int4)
AS 'MODULE_PATHNAME', 'pg_buffercache_evict_all'
LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Add the buffer lookup table function
+CREATE FUNCTION pg_buffercache_lookup_table_entries(
+ OUT tablespace oid,
+ OUT database oid,
+ OUT relfilenode oid,
+ OUT forknum int2,
+ OUT blocknum int8,
+ OUT bufferid int4)
+RETURNS SETOF record
+AS 'MODULE_PATHNAME', 'pg_buffercache_lookup_table_entries'
+LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Create a view for convenient access.
+CREATE VIEW pg_buffercache_lookup_table AS
+ SELECT * FROM pg_buffercache_lookup_table_entries();
+
+-- Don't want these to be available to public.
+REVOKE ALL ON FUNCTION pg_buffercache_lookup_table_entries() FROM PUBLIC;
+REVOKE ALL ON pg_buffercache_lookup_table FROM PUBLIC;
+
+-- Grant access to monitoring role.
+GRANT EXECUTE ON FUNCTION pg_buffercache_lookup_table_entries() TO pg_read_all_stats;
+GRANT SELECT ON pg_buffercache_lookup_table TO pg_read_all_stats;
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 510455998aa..9512f1efa2f 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -77,6 +77,7 @@ PG_FUNCTION_INFO_V1(pg_buffercache_evict_all);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_relation);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_all);
+PG_FUNCTION_INFO_V1(pg_buffercache_lookup_table_entries);
/* Only need to touch memory once per backend process lifetime */
@@ -922,3 +923,19 @@ pg_buffercache_mark_dirty_all(PG_FUNCTION_ARGS)
PG_RETURN_DATUM(result);
}
+
+/*
+ * Return lookup table content as a set of records.
+ */
+Datum
+pg_buffercache_lookup_table_entries(PG_FUNCTION_ARGS)
+{
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* Fill the tuplestore */
+ BufTableGetContents(rsinfo->setResult, rsinfo->setDesc);
+
+ return (Datum) 0;
+}
diff --git a/contrib/pg_buffercache/sql/pg_buffercache.sql b/contrib/pg_buffercache/sql/pg_buffercache.sql
index be89b5f5a3a..d0726b2b2f9 100644
--- a/contrib/pg_buffercache/sql/pg_buffercache.sql
+++ b/contrib/pg_buffercache/sql/pg_buffercache.sql
@@ -18,6 +18,18 @@ from pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -26,6 +38,8 @@ SELECT * FROM pg_buffercache_os_pages;
SELECT * FROM pg_buffercache_pages() AS p (wrong int);
SELECT * FROM pg_buffercache_summary();
SELECT * FROM pg_buffercache_usage_counts();
+SELECT * FROM pg_buffercache_lookup_table_entries();
+SELECT * FROM pg_buffercache_lookup_table;
RESET role;
-- Check that pg_monitor is allowed to query view / function
@@ -42,6 +56,12 @@ FROM pg_buffercache_pages() AS p
LIMIT 1;
RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+RESET role;
+
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 5ea19d68622..6b905498337 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -71,6 +71,11 @@
<entry>backend memory contexts</entry>
</row>
+ <row>
+ <entry><link linkend="view-pg-buffer-lookup-table"><structname>pg_buffer_lookup_table</structname></link></entry>
+ <entry>shared buffer lookup table</entry>
+ </row>
+
<row>
<entry><link linkend="view-pg-config"><structname>pg_config</structname></link></entry>
<entry>compile-time configuration parameters</entry>
@@ -929,6 +934,90 @@ AND c1.path[c2.level] = c2.path[c2.level];
</para>
</sect1>
+ <sect1 id="view-pg-buffer-lookup-table">
+ <title><structname>pg_buffer_lookup_table</structname></title>
+ <indexterm>
+ <primary>pg_buffer_lookup_table</primary>
+ </indexterm>
+ <para>
+ The <structname>pg_buffer_lookup_table</structname> view exposes the current
+ contents of the shared buffer lookup table. Each row represents an entry in
+ the lookup table mapping a relation page to the ID of buffer in which it is
+ cached. The shared buffer lookup table is locked for a short duration while
+ reading so as to ensure consistency. This may affect performance if this view
+ is queried very frequently.
+ </para>
+ <table id="pg-buffer-lookup-table-view" xreflabel="pg_buffer_lookup_table">
+ <title><structname>pg_buffer_lookup_table</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>tablespace</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the tablespace containing the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>database</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database containing the relation (zero for shared relations)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relfilenode</structfield> <type>oid</type>
+ </para>
+ <para>
+ relfilenode identifying the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>forknum</structfield> <type>int2</type>
+ </para>
+ <para>
+ Fork number within the relation (see <xref linkend="storage-file-layout"/>)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>blocknum</structfield> <type>int8</type>
+ </para>
+ <para>
+ Block number within the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>bufferid</structfield> <type>int4</type>
+ </para>
+ <para>
+ ID of the buffer caching the page
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ <para>
+ Access to this view is restricted to members of the
+ <literal>pg_read_all_stats</literal> role by default.
+ </para>
+ </sect1>
+
<sect1 id="view-pg-config">
<title><structname>pg_config</structname></title>
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 5c8ccbee13f..c82b71deaa6 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -21,7 +21,12 @@
*/
#include "postgres.h"
+#include "fmgr.h"
+#include "funcapi.h"
#include "storage/buf_internals.h"
+#include "utils/rel.h"
+#include "utils/builtins.h"
+#include "storage/lwlock.h"
#include "storage/subsystems.h"
/* entry for buffer lookup hashtable */
@@ -167,3 +172,56 @@ BufTableDelete(BufferTag *tagPtr, uint32 hashcode)
if (!result) /* shouldn't happen */
elog(ERROR, "shared buffer hash table corrupted");
}
+
+/*
+ * BufTableGetContents
+ * Fill the given tuplestore with contents of the shared buffer lookup table
+ *
+ * This function is used by pg_buffercache extension to expose buffer lookup
+ * table contents via SQL. The caller is responsible for setting up the
+ * tuplestore and result set info.
+ */
+void
+BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc)
+{
+/* Expected number of attributes of the buffer lookup table entry. */
+#define BUFTABLE_CONTENTS_COLS 6
+
+ HASH_SEQ_STATUS hstat;
+ BufferLookupEnt *ent;
+ Datum values[BUFTABLE_CONTENTS_COLS];
+ bool nulls[BUFTABLE_CONTENTS_COLS];
+ int i;
+
+ memset(nulls, 0, sizeof(nulls));
+
+ Assert(tupdesc->natts == BUFTABLE_CONTENTS_COLS);
+
+ /*
+ * Lock all buffer mapping partitions to ensure a consistent view of the
+ * hash table during the scan. Must grab LWLocks in partition-number order
+ * to avoid LWLock deadlock.
+ */
+ for (i = 0; i < NUM_BUFFER_PARTITIONS; i++)
+ LWLockAcquire(BufMappingPartitionLockByIndex(i), LW_SHARED);
+
+ hash_seq_init(&hstat, SharedBufHash);
+ while ((ent = (BufferLookupEnt *) hash_seq_search(&hstat)) != NULL)
+ {
+ values[0] = ObjectIdGetDatum(ent->key.spcOid);
+ values[1] = ObjectIdGetDatum(ent->key.dbOid);
+ values[2] = ObjectIdGetDatum(ent->key.relNumber);
+ values[3] = ObjectIdGetDatum(ent->key.forkNum);
+ values[4] = Int64GetDatum(ent->key.blockNum);
+ values[5] = Int32GetDatum(ent->id);
+
+ tuplestore_putvalues(tupstore, tupdesc, values, nulls);
+ }
+
+ /*
+ * Release all buffer mapping partition locks in the reverse order so as
+ * to avoid LWLock deadlock.
+ */
+ for (i = NUM_BUFFER_PARTITIONS - 1; i >= 0; i--)
+ LWLockRelease(BufMappingPartitionLockByIndex(i));
+}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index a7606c6e92b..678065b3de2 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -29,6 +29,7 @@
#include "storage/spin.h"
#include "utils/relcache.h"
#include "utils/resowner.h"
+#include "utils/tuplestore.h"
/*
* Buffer state is a single 64-bit variable where following data is combined.
@@ -599,6 +600,7 @@ extern uint32 BufTableHashCode(BufferTag *tagPtr);
extern int BufTableLookup(BufferTag *tagPtr, uint32 hashcode);
extern int BufTableInsert(BufferTag *tagPtr, uint32 hashcode, int buf_id);
extern void BufTableDelete(BufferTag *tagPtr, uint32 hashcode);
+extern void BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc);
/* localbuf.c */
extern bool PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount);
--
2.34.1
[text/x-patch] v20260817-0004-Pass-use_units-parameter-to-GucShowHook-fu.patch (15.6K, ../../CAExHW5ts93Rnof7pjFFYrY9aTOmCo+xsdQazqeqN1BaKozTBvA@mail.gmail.com/3-v20260817-0004-Pass-use_units-parameter-to-GucShowHook-fu.patch)
download | inline diff:
From d63d5a6f99957f5b0d72db9e1fc7e9fa574c3a83 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 25 Feb 2026 15:48:14 +0530
Subject: [PATCH v20260817 4/7] Pass use_units parameter to GucShowHook
functions
ShowGUCOption() prints the value of a GUC variable with units if the variable
has a unit and the caller passes use_units as true. However when a GUC variable
has a show_hook associated with it, the show_hook does not receive use_units
parameter. Hence a show_hook cannot determine whether to show the value with
units or not. This was not a problem until now because all the GUC variables
with show_hook were either didn't have any units associated with them or their
show_hook used a fixed unit to print the value of the variable.
shared_buffers is a GUC variable whose value is shows in different units based
on the GUC variable's value. With shared buffer pool resizing, we want to show
the pending size of the buffer pool, if any with appropriate units. We do this
using a show_hook, which needs use_units input. This commit adds use_units
parameter to GucShowHook and passes it from ShowGUCOption(). The actual commit
which uses this new parameter will be in a followup commit.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Reported by: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
---
src/backend/access/transam/xlog.c | 8 ++++----
src/backend/commands/variable.c | 14 +++++++-------
src/backend/executor/instrument.c | 2 +-
src/backend/libpq/pqcomm.c | 8 ++++----
src/backend/utils/misc/guc.c | 10 +++++-----
src/include/access/xlog.h | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 30 +++++++++++++++---------------
8 files changed, 38 insertions(+), 38 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 10650d01b9c..c69371b792a 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4975,7 +4975,7 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
/* guc hook */
const char *
-show_data_checksums(void)
+show_data_checksums(bool use_units)
{
return get_checksum_state_string(LocalDataChecksumState);
}
@@ -5211,7 +5211,7 @@ InitializeWalConsistencyChecking(void)
* GUC show_hook for archive_command
*/
const char *
-show_archive_command(void)
+show_archive_command(bool use_units)
{
if (XLogArchivingActive())
return XLogArchiveCommand;
@@ -5223,7 +5223,7 @@ show_archive_command(void)
* GUC show_hook for in_hot_standby
*/
const char *
-show_in_hot_standby(void)
+show_in_hot_standby(bool use_units)
{
/*
* We display the actual state based on shared memory, so that this GUC
@@ -5238,7 +5238,7 @@ show_in_hot_standby(void)
* GUC show_hook for effective_wal_level
*/
const char *
-show_effective_wal_level(void)
+show_effective_wal_level(bool use_units)
{
if (wal_level == WAL_LEVEL_MINIMAL)
return "minimal";
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 8afd252fc8c..491c8aa392d 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -389,7 +389,7 @@ assign_timezone(const char *newval, void *extra)
* show_timezone: GUC show_hook for timezone
*/
const char *
-show_timezone(void)
+show_timezone(bool use_units)
{
const char *tzn;
@@ -462,7 +462,7 @@ assign_log_timezone(const char *newval, void *extra)
* show_log_timezone: GUC show_hook for log_timezone
*/
const char *
-show_log_timezone(void)
+show_log_timezone(bool use_units)
{
const char *tzn;
@@ -674,7 +674,7 @@ assign_random_seed(double newval, void *extra)
}
const char *
-show_random_seed(void)
+show_random_seed(bool use_units)
{
return "unavailable";
}
@@ -1031,7 +1031,7 @@ assign_role(const char *newval, void *extra)
}
const char *
-show_role(void)
+show_role(bool use_units)
{
/*
* Check whether SET ROLE is active; if not return "none". This is a
@@ -1179,7 +1179,7 @@ assign_io_combine_limit(int newval, void *extra)
* GUC show_hook for data_directory_mode
*/
const char *
-show_data_directory_mode(void)
+show_data_directory_mode(bool use_units)
{
static char buf[12];
@@ -1191,7 +1191,7 @@ show_data_directory_mode(void)
* GUC show_hook for log_file_mode
*/
const char *
-show_log_file_mode(void)
+show_log_file_mode(bool use_units)
{
static char buf[12];
@@ -1203,7 +1203,7 @@ show_log_file_mode(void)
* GUC show_hook for unix_socket_permissions
*/
const char *
-show_unix_socket_permissions(void)
+show_unix_socket_permissions(bool use_units)
{
static char buf[12];
diff --git a/src/backend/executor/instrument.c b/src/backend/executor/instrument.c
index ffbcd572133..d4c8373336a 100644
--- a/src/backend/executor/instrument.c
+++ b/src/backend/executor/instrument.c
@@ -423,7 +423,7 @@ assign_timing_clock_source(int newval, void *extra)
}
const char *
-show_timing_clock_source(void)
+show_timing_clock_source(bool use_units)
{
switch (timing_clock_source)
{
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index 3704d121003..9f08da25b35 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1969,7 +1969,7 @@ assign_tcp_keepalives_idle(int newval, void *extra)
* GUC show_hook for tcp_keepalives_idle
*/
const char *
-show_tcp_keepalives_idle(void)
+show_tcp_keepalives_idle(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -1992,7 +1992,7 @@ assign_tcp_keepalives_interval(int newval, void *extra)
* GUC show_hook for tcp_keepalives_interval
*/
const char *
-show_tcp_keepalives_interval(void)
+show_tcp_keepalives_interval(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -2015,7 +2015,7 @@ assign_tcp_keepalives_count(int newval, void *extra)
* GUC show_hook for tcp_keepalives_count
*/
const char *
-show_tcp_keepalives_count(void)
+show_tcp_keepalives_count(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -2038,7 +2038,7 @@ assign_tcp_user_timeout(int newval, void *extra)
* GUC show_hook for tcp_user_timeout
*/
const char *
-show_tcp_user_timeout(void)
+show_tcp_user_timeout(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 774bbc9be5f..1a5a168bc1a 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -5381,7 +5381,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_bool *conf = &record->_bool;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
val = *conf->variable ? "on" : "off";
}
@@ -5392,7 +5392,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_int *conf = &record->_int;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
{
/*
@@ -5421,7 +5421,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_real *conf = &record->_real;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
{
double result = *conf->variable;
@@ -5446,7 +5446,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_string *conf = &record->_string;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else if (*conf->variable && **conf->variable)
val = *conf->variable;
else
@@ -5459,7 +5459,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_enum *conf = &record->_enum;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
val = config_enum_lookup_by_value(record, *conf->variable);
}
diff --git a/src/include/access/xlog.h b/src/include/access/xlog.h
index 338d68d7424..7b23261dac2 100644
--- a/src/include/access/xlog.h
+++ b/src/include/access/xlog.h
@@ -268,7 +268,7 @@ extern bool DataChecksumsInProgressOn(void);
extern void SetDataChecksumsOnInProgress(void);
extern void SetDataChecksumsOn(void);
extern void SetDataChecksumsOff(void);
-extern const char *show_data_checksums(void);
+extern const char *show_data_checksums(bool use_units);
extern const char *get_checksum_state_string(uint32 state);
extern void InitLocalDataChecksumState(void);
extern void SetLocalDataChecksumState(uint32 data_checksum_version);
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 8057d7870ad..2a6e2ed18b3 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -192,7 +192,7 @@ typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
-typedef const char *(*GucShowHook) (void);
+typedef const char *(*GucShowHook) (bool use_units);
/*
* Miscellaneous
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 6a76f8d5ed6..df048517a0f 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -28,7 +28,7 @@
extern bool check_application_name(char **newval, void **extra,
GucSource source);
extern void assign_application_name(const char *newval, void *extra);
-extern const char *show_archive_command(void);
+extern const char *show_archive_command(bool use_units);
extern bool check_autovacuum_work_mem(int *newval, void **extra,
GucSource source);
extern bool check_vacuum_buffer_usage_limit(int *newval, void **extra,
@@ -46,7 +46,7 @@ extern void assign_client_encoding(const char *newval, void *extra);
extern bool check_cluster_name(char **newval, void **extra, GucSource source);
extern bool check_commit_ts_buffers(int *newval, void **extra,
GucSource source);
-extern const char *show_data_directory_mode(void);
+extern const char *show_data_directory_mode(bool use_units);
extern bool check_datestyle(char **newval, void **extra, GucSource source);
extern void assign_datestyle(const char *newval, void *extra);
extern bool check_debug_io_direct(char **newval, void **extra, GucSource source);
@@ -61,11 +61,11 @@ extern bool check_default_text_search_config(char **newval, void **extra, GucSou
extern void assign_default_text_search_config(const char *newval, void *extra);
extern bool check_default_with_oids(bool *newval, void **extra,
GucSource source);
-extern const char *show_effective_wal_level(void);
+extern const char *show_effective_wal_level(bool use_units);
extern bool check_huge_page_size(int *newval, void **extra, GucSource source);
extern void assign_io_method(int newval, void *extra);
extern bool check_io_max_concurrency(int *newval, void **extra, GucSource source);
-extern const char *show_in_hot_standby(void);
+extern const char *show_in_hot_standby(bool use_units);
extern bool check_locale_messages(char **newval, void **extra, GucSource source);
extern void assign_locale_messages(const char *newval, void *extra);
extern bool check_locale_monetary(char **newval, void **extra, GucSource source);
@@ -77,11 +77,11 @@ extern void assign_locale_time(const char *newval, void *extra);
extern bool check_log_destination(char **newval, void **extra,
GucSource source);
extern void assign_log_destination(const char *newval, void *extra);
-extern const char *show_log_file_mode(void);
+extern const char *show_log_file_mode(bool use_units);
extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
-extern const char *show_log_timezone(void);
+extern const char *show_log_timezone(bool use_units);
extern void assign_maintenance_io_concurrency(int newval, void *extra);
extern void assign_io_max_combine_limit(int newval, void *extra);
extern void assign_io_combine_limit(int newval, void *extra);
@@ -97,7 +97,7 @@ extern bool check_primary_slot_name(char **newval, void **extra,
GucSource source);
extern bool check_random_seed(double *newval, void **extra, GucSource source);
extern void assign_random_seed(double newval, void *extra);
-extern const char *show_random_seed(void);
+extern const char *show_random_seed(bool use_units);
extern bool check_recovery_prefetch(int *new_value, void **extra,
GucSource source);
extern void assign_recovery_prefetch(int new_value, void *extra);
@@ -118,7 +118,7 @@ extern bool check_recovery_target_xid(char **newval, void **extra,
extern void assign_recovery_target_xid(const char *newval, void *extra);
extern bool check_role(char **newval, void **extra, GucSource source);
extern void assign_role(const char *newval, void *extra);
-extern const char *show_role(void);
+extern const char *show_role(bool use_units);
extern bool check_restrict_nonsystem_relation_kind(char **newval, void **extra,
GucSource source);
extern void assign_restrict_nonsystem_relation_kind(const char *newval, void *extra);
@@ -143,32 +143,32 @@ extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
extern void assign_tcp_keepalives_count(int newval, void *extra);
-extern const char *show_tcp_keepalives_count(void);
+extern const char *show_tcp_keepalives_count(bool use_units);
extern void assign_tcp_keepalives_idle(int newval, void *extra);
-extern const char *show_tcp_keepalives_idle(void);
+extern const char *show_tcp_keepalives_idle(bool use_units);
extern void assign_tcp_keepalives_interval(int newval, void *extra);
-extern const char *show_tcp_keepalives_interval(void);
+extern const char *show_tcp_keepalives_interval(bool use_units);
extern void assign_tcp_user_timeout(int newval, void *extra);
-extern const char *show_tcp_user_timeout(void);
+extern const char *show_tcp_user_timeout(bool use_units);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
GucSource source);
extern void assign_temp_tablespaces(const char *newval, void *extra);
extern bool check_timezone(char **newval, void **extra, GucSource source);
extern void assign_timezone(const char *newval, void *extra);
-extern const char *show_timezone(void);
+extern const char *show_timezone(bool use_units);
extern bool check_timezone_abbreviations(char **newval, void **extra,
GucSource source);
extern void assign_timezone_abbreviations(const char *newval, void *extra);
extern void assign_timing_clock_source(int newval, void *extra);
extern bool check_timing_clock_source(int *newval, void **extra, GucSource source);
-extern const char *show_timing_clock_source(void);
+extern const char *show_timing_clock_source(bool use_units);
extern bool check_transaction_buffers(int *newval, void **extra, GucSource source);
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
extern void assign_transaction_timeout(int newval, void *extra);
-extern const char *show_unix_socket_permissions(void);
+extern const char *show_unix_socket_permissions(bool use_units);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
GucSource source);
--
2.34.1
[text/x-patch] v20260817-0005-PID-of-the-backend-process-backing-the-Bac.patch (3.0K, ../../CAExHW5ts93Rnof7pjFFYrY9aTOmCo+xsdQazqeqN1BaKozTBvA@mail.gmail.com/4-v20260817-0005-PID-of-the-backend-process-backing-the-Bac.patch)
download | inline diff:
From cd361022f8f3ea1821ba9a838d3261ac8eaec120 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 23 Jun 2026 09:19:06 +0530
Subject: [PATCH v20260817 5/7] PID of the backend process backing the
BackgroundPsql session
Many tests which use background psql sessions also fetch the pid of the backend
process as a separate step. They end up maintaing a separate variable for the
pid and pass it to the subroutines along with the BackgroundPsql object when
needed. Further the PID can not be fetched if the backend is blocked in a query
or an injection point or if the backend has died. This can be avoided by
proactively fetching the PID of the backend process when the BackgroundPsql is
started and storing it in the BackgroundPsql object. The PID can then be fetched
from the BackgroundPsql object when needed.
The pid is reset when the BackgroundPsql session is finished. Many of the tests
which use BackgroundPsql explicitly call {run}->finish to finish the session
when the backend is gone. They will have stale PID in the BackgroundPsql object.
The commit introduces a finish method in BackgroundPsql which will finish the
session and reset the PID.
Note to the reviewer:
This commit adds code to set the PID proactively and also the finish method.
These will be used in a future buffer resize test. But I have not modified the
existing tests to use the new capabilities. If we find this change useful, we
can modify the existing tests to use the new capabilities before committing this
change.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 29 ++++++++++++++++++-
1 file changed, 28 insertions(+), 1 deletion(-)
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index d7797225451..699334320d9 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -173,6 +173,14 @@ sub wait_connect
$self->{stderr} = '';
die "psql startup timed out" if $self->{timeout}->is_expired;
+
+ # Many tests which use background psql sessions also fetch the pid of the
+ # backend process, so we capture it here. The Callers that need the pid
+ # after a blocking query or after the backend has died can read it from
+ # $self->{backend_pid}.
+ my $pid = $self->query('SELECT pg_backend_pid()', verbose => 0);
+ chomp $pid;
+ $self->{backend_pid} = $pid;
}
=pod
@@ -190,6 +198,25 @@ sub quit
$self->{stdin} .= "\\q\n";
+ return $self->finish;
+}
+
+=pod
+
+=item $session->finish
+
+Reap the underlying IPC::Run handle without sending \q. Intended for
+sessions whose psql process has already exited (e.g. after the server
+terminated the backend or the client connection was killed).
+
+=cut
+
+sub finish
+{
+ my ($self) = @_;
+
+ $self->{backend_pid} = undef;
+
return $self->{run}->finish;
}
@@ -212,7 +239,7 @@ sub reconnect_and_clear
{
$self->{stdin} .= "\\q\n";
}
- $self->{run}->finish;
+ $self->finish;
# restart
$self->{run}->run();
--
2.34.1
[text/x-patch] v20260817-0006-Resizable-shared-memory-structures.patch (129.7K, ../../CAExHW5ts93Rnof7pjFFYrY9aTOmCo+xsdQazqeqN1BaKozTBvA@mail.gmail.com/5-v20260817-0006-Resizable-shared-memory-structures.patch)
download | inline diff:
From 6dd1408128979e89c1b2510e00fc841cd46e88bf Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 17 Feb 2026 16:51:20 +0530
Subject: [PATCH v20260817 6/7] Resizable shared memory structures
Resizable shared memory structures can be requested by specifying a new
member ShmemStructOpts::maximum_size. At the startup or when the
structure is created, we reserve address space worth maximum_size in the
shared memory segment. It is expected that the subsystem which creates a
resizable structure would initialize only the memory worth its initial
size given by ShmemStructOpts::maximum_size when creating it. In an
mmap'ed memory, this should allocate memory worth only the initial size.
It should not allocate maximum_size worth of memory initially. As the
structure is resized using ShmemResizeStruct() memory is freed or
allocated in chunks of memory pages when shrinking and expanding the
structure respectively. Optional ShmemStructOpts::minimum_size
specification for resizable shared memory structures allows us to
enforce that a resizable structure cannot be shrunk below a certain
size. If minimum_size is not specified for a resizable structure, it the
minimum size defaults to 0. Two additional columns are added to
pg_shmem_allocations view to report minimum and maximum size
respectively.
The structure for which ShmemStructOpts::maximum_size is not specified
or is set to 0 is considered as a fixed size structure. Existing calls
to ShmemRequestStruct() requsting fixed size structures need no change.
For fixed-size structures, the minimum size and maximum size are set to
the initial size specified in ShmemRequestStructOpts::size. This makes
maximum size and minimum size reported in pg_shmem_allocations view
semantically consistent for both fixed-size and resizable structures. As
a side effect, a request which specifies same minimum_size and
maximum_size will be treated as a fixed size structure.
Minimum, maximum and initial sizes of main shared memory area
=============================================================
With the addition of resizable shared structures, the main shared memory
area has three sizes to be tracked:
- a. the minimum amount of shared memory that will always be needed
- b. the maximum amount of shared memory that may be allocated when all
the resizable structures have grown to their maximum size. It is also
is the size of the address space reserved for the main shared memory
area.
- c. the amount of memory allocated at the server startup i.e. initial
allocation.
mmap needs to use the maximum size to reserve enough address space to
accomodate all the shared structures. But a DBA may provision only the
memory required at the startup initially and increase or descrease the
memory provision as the demand changes.
These three sizes are reported as GUCs shared_memory_minimum_size,
shared_memory_maximum_size and shared_memory_initial_size respectively.
They replace the old GUC shared_memory_size which used to report both
size of the main shared memory area and amount of memory required in it.
Since these two things are not the same anymore, a single GUC is not
sufficient. These GUCs help DBAs to estimate memory to be provisioned
at the startup and during run time as the resizable structures change
their sizes.
Portability
===========
Resizable shared structures feature depends upon existence of function
madvise() and constants MADV_REMOVE and MADV_WRITE_POPULATE. On the
platforms which do not have these or the shared memory types (e.g.
Sys-V) which do not provide ability to free or allocate parts of shared
memory, we disable this feature. A run time GUC have_resizable_shmem
indicates whether a running server supports resizable structures or not.
The commit introduces a compile time flag HAVE_RESIZABLE_SHMEM which is
defined if MADV_REMOVE and MADV_WRITE_POPULATE exist. We don't check
existence of madvise separately, since existence of the constants
implies existence of the function. HAVE_RESIZABLE_SHMEM is not defined
in EXEC_BACKEND builds since that's largely used for Windows where the
APIs to free and allocate memory from and to a given address space are
not known to the author right now. Given that PostgreSQL is used widely
on Linux (with shared memory type mmap), providing this feature on Linux
benefits most of its users. Once we figure out the required Windows
APIs, we will support this feature on Windows as well.
Prohibiting access to the unused portion of a resizable structure
================================================================
We provide ShmemProtectStruct() to add protections on the address space
reserved for a resizable shared memory structure so that the address
space upto its current size is accessible whereas the part beyond that
is inaccessible. Since these protections are backend specific, the
subsystem using the resizable shared memory structures has to make sure
to call ShmemProtectStruct() after every resize before any backend tries
to access the portion of the structure between old and the new size.
This can coordinated using ProcSignalBarrier mechanism in a running
backend. When starting the server, Postmaster adds appropriate
protections when creating these structures. In an EXEC_BACKEND case,
when a backend starts, the protections are added according to the
current size of the backend (through attach_fn call). However, in
non-EXEC_BACKEND case, a new backend inherits stale protections from the
postmaster. It needs to apply the protections according to the current
sizes of these structures.
Following points need more discussion.
Discussion points
=================
adding initial protections in non-EXEC_BACKEND case
----------------------------------------------------------------------
When a backend starts, it needs to add shmem protections it in such a
way that a ProcSignalBarrier conveying the protection change is not
missed. Hence we do it along side InitLocalDataChecksumState(). This has
two problems 1. it delays backend startup process and 2. it loosely ties
resizable shared memory structure to ProcSignalBarrier mechanim. Do we
consider the complexity worth it? Is there any other way to do it
without these two drawbacks?
Further, we call ShmemReprotectResizableStructs() twice in the startup
sequence, once immediately after the place where
AttachSharedMemoryStructs() would be called in EXEC_BACKEND case and
then second time after ProcSignalInit(). The first one is needed so that
the processes can access the resizable structures safely. The second is
needed for the reasons mentioned there. But a change in the protection
during these calls will be missed by the backend and thus lead to
hazardous access. Can we avoid it? If ProcSignalInit() were to be called
nearer to the InitProcess(), it may make calling
ShmemReprotectResizeStructs() easier.
Protection in the resizing backend vs other backends
----------------------------------------------------
The patch provides an API to protect the address spaces not used because
of current size of the resizable structures in every backend. Since
every backend has to protect its own address space, a separate API is
required. But at the same time, in some cases the resize operation
itself needs to change the protection (e.g. when expanding the
structure). So there's some asymmetry between the moment when the
resizing backend adds protection and when other backends add it. That
doesn't affect end result since adding this protection is an idempotent
operation. But still there's some ugliness involved. This part also
needs more discussion.
HAVE_RESIZABLE_SHMEM and have_resizable_shmem
---------------------------------------------
The patch defines a build time macro HAVE_RESIZABLE_SHMEM to avoid
compilation failure on a build platform where the required system calls
or constants are not available. But we also need a run time check to see
whether the shared memory type being used allows resizing. Hence we
provide a GUC have_resizable_shmem which tells whether a running server
supports resizable shared memory or not. Most of the code which deals
with the shared memory is under HAVE_RESIZABLE_SHMEM, but there is some
related to resizable shared memory outside the macro as well. For
example, the members maximum_size, minimum_size are not under this
macro. This is mostly to allow writing readable code in both the modes
and still provide the same user visible views etc. This division of code
within and without macro needs more discussion.
TODOs
=====
There are some TODOs in the code still. We will address them as we
finalize that part of design/code. Comments/suggestions on those is
welcome.
Idea of using mmap to reserve address space in shared memory segments
and allocating memory within that address space on demand was proposed
by Dmitry Dolgov <9erthalion6@gmail.com>.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Reviewed-by: Matthias van de Meent <boekewurm+postgres@gmail.com>
---
configure.ac | 4 +
doc/src/sgml/config.sgml | 102 +++-
doc/src/sgml/ref/postgres-ref.sgml | 8 +-
doc/src/sgml/runtime.sgml | 9 +-
doc/src/sgml/system-views.sgml | 57 +-
doc/src/sgml/xfunc.sgml | 92 +++
meson.build | 16 +
src/backend/port/sysv_shmem.c | 259 +++++++-
src/backend/port/win32_shmem.c | 64 +-
src/backend/postmaster/auxprocess.c | 46 +-
src/backend/storage/ipc/ipci.c | 114 ++--
src/backend/storage/ipc/shmem.c | 555 ++++++++++++++++--
src/backend/storage/lmgr/proc.c | 16 +
src/backend/utils/init/postinit.c | 10 +
src/backend/utils/misc/guc_parameters.dat | 57 +-
src/backend/utils/misc/guc_tables.c | 15 +-
src/include/catalog/pg_proc.dat | 4 +-
src/include/pg_config.h.in | 8 +
src/include/pg_config_manual.h | 14 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_shmem.h | 6 +-
src/include/storage/shmem.h | 19 +
src/include/storage/shmem_internal.h | 2 +-
src/test/modules/test_shmem/meson.build | 3 +-
...mem_alloc.pl => 001_fixed_shmem_struct.pl} | 40 ++
.../t/002_resizable_shmem_struct.pl | 450 ++++++++++++++
.../modules/test_shmem/test_shmem--1.0.sql | 55 ++
src/test/modules/test_shmem/test_shmem.c | 504 +++++++++++++++-
src/test/regress/expected/rules.out | 7 +-
src/tools/pgindent/typedefs.list | 1 +
30 files changed, 2371 insertions(+), 168 deletions(-)
rename src/test/modules/test_shmem/t/{001_late_shmem_alloc.pl => 001_fixed_shmem_struct.pl} (58%)
create mode 100644 src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
diff --git a/configure.ac b/configure.ac
index a331749fcb5..bba12b8c8f4 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1912,6 +1912,10 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux-specific madvise constants needed for resizable shared memory. See similar checks in meson.build for explanation of why these checks are here.
+AC_CHECK_DECLS([MADV_POPULATE_WRITE], [], [], [#include <sys/mman.h>])
+AC_CHECK_DECLS([MADV_REMOVE], [], [], [#include <sys/mman.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 596fd45a3db..f06e0f2e1a5 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -12456,6 +12456,20 @@ dynamic_library_path = '/usr/local/lib/postgresql:$libdir'
</listitem>
</varlistentry>
+ <varlistentry id="guc-have-resizable-shmem" xreflabel="have_resizable_shmem">
+ <term><varname>have_resizable_shmem</varname> (<type>boolean</type>)
+ <indexterm>
+ <primary><varname>have_resizable_shmem</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports whether <productname>PostgreSQL</productname> supports <link
+ linkend="xfunc-shared-addin-resizable">Resizable shared memory structures</link>.
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry id="guc-huge-pages-status" xreflabel="huge_pages_status">
<term><varname>huge_pages_status</varname> (<type>enum</type>)
<indexterm>
@@ -12635,32 +12649,100 @@ dynamic_library_path = '/usr/local/lib/postgresql:$libdir'
</listitem>
</varlistentry>
- <varlistentry id="guc-shared-memory-size" xreflabel="shared_memory_size">
- <term><varname>shared_memory_size</varname> (<type>integer</type>)
+ <varlistentry id="guc-shared-memory-initial-size" xreflabel="shared_memory_initial_size">
+ <term><varname>shared_memory_initial_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_initial_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the amount of memory, rounded up to the nearest megabyte, allocated at the server startup in main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-minimum-size" xreflabel="shared_memory_minimum_size">
+ <term><varname>shared_memory_minimum_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_minimum_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the minimum amount of memory, rounded up to the nearest megabyte, required in the main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-maximum-size" xreflabel="shared_memory_maximum_size">
+ <term><varname>shared_memory_maximum_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_maximum_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the maximum size of the main shared memory area, rounded up
+ to the nearest megabyte. This is the amount of address space that
+ must be reserved for the main shared memory area. This is also the maximum amount of memory that may be allocated in the main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-initial-size-in-huge-pages" xreflabel="shared_memory_initial_size_in_huge_pages">
+ <term><varname>shared_memory_initial_size_in_huge_pages</varname> (<type>integer</type>)
<indexterm>
- <primary><varname>shared_memory_size</varname> configuration parameter</primary>
+ <primary><varname>shared_memory_initial_size_in_huge_pages</varname> configuration parameter</primary>
</indexterm>
</term>
<listitem>
<para>
- Reports the size of the main shared memory area, rounded up to the
- nearest megabyte.
+ Reports the number of huge pages needed in the main shared memory area
+ at the startup based on the specified <xref
+ linkend="guc-huge-page-size"/>. If huge pages are not supported, this
+ will be <literal>-1</literal>.
+ </para>
+ <para>
+ This setting is supported only on <productname>Linux</productname>. It
+ is always set to <literal>-1</literal> on other platforms. For more
+ details about using huge pages on <productname>Linux</productname>, see
+ <xref linkend="linux-huge-pages"/>.
</para>
</listitem>
</varlistentry>
- <varlistentry id="guc-shared-memory-size-in-huge-pages" xreflabel="shared_memory_size_in_huge_pages">
- <term><varname>shared_memory_size_in_huge_pages</varname> (<type>integer</type>)
+ <varlistentry id="guc-shared-memory-minimum-size-in-huge-pages" xreflabel="shared_memory_minimum_size_in_huge_pages">
+ <term><varname>shared_memory_minimum_size_in_huge_pages</varname> (<type>integer</type>)
<indexterm>
- <primary><varname>shared_memory_size_in_huge_pages</varname> configuration parameter</primary>
+ <primary><varname>shared_memory_minimum_size_in_huge_pages</varname> configuration parameter</primary>
</indexterm>
</term>
<listitem>
<para>
- Reports the number of huge pages that are needed for the main shared
- memory area based on the specified <xref linkend="guc-huge-page-size"/>.
+ Reports the minimum number of huge pages needed in the main shared memory area based on the specified <xref linkend="guc-huge-page-size"/>.
If huge pages are not supported, this will be <literal>-1</literal>.
</para>
+ <para>
+ This setting is supported only on <productname>Linux</productname>. It
+ is always set to <literal>-1</literal> on other platforms.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-maximum-size-in-huge-pages" xreflabel="shared_memory_maximum_size_in_huge_pages">
+ <term><varname>shared_memory_maximum_size_in_huge_pages</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_maximum_size_in_huge_pages</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the maximum number of huge pages needed in the
+ main shared memory area based on the specified
+ <xref linkend="guc-huge-page-size"/>. If huge pages are not
+ supported, this will be <literal>-1</literal>.
+ </para>
<para>
This setting is supported only on <productname>Linux</productname>. It
is always set to <literal>-1</literal> on other platforms. For more
diff --git a/doc/src/sgml/ref/postgres-ref.sgml b/doc/src/sgml/ref/postgres-ref.sgml
index b13a16a117f..3eefd629e06 100644
--- a/doc/src/sgml/ref/postgres-ref.sgml
+++ b/doc/src/sgml/ref/postgres-ref.sgml
@@ -143,8 +143,12 @@ PostgreSQL documentation
<para>
This can be used on a running server for most parameters. However,
the server must be shut down for some runtime-computed parameters
- (e.g., <xref linkend="guc-shared-memory-size"/>,
- <xref linkend="guc-shared-memory-size-in-huge-pages"/>, and
+ (e.g., <xref linkend="guc-shared-memory-initial-size"/>,
+ <xref linkend="guc-shared-memory-minimum-size"/>,
+ <xref linkend="guc-shared-memory-maximum-size"/>,
+ <xref linkend="guc-shared-memory-initial-size-in-huge-pages"/>,
+ <xref linkend="guc-shared-memory-minimum-size-in-huge-pages"/>,
+ <xref linkend="guc-shared-memory-maximum-size-in-huge-pages"/>, and
<xref linkend="guc-wal-segment-size"/>).
</para>
diff --git a/doc/src/sgml/runtime.sgml b/doc/src/sgml/runtime.sgml
index d9984910cc4..8ba4d233509 100644
--- a/doc/src/sgml/runtime.sgml
+++ b/doc/src/sgml/runtime.sgml
@@ -1450,11 +1450,12 @@ export PG_OOM_ADJUST_VALUE=0
<varname>CONFIG_HUGETLB_PAGE=y</varname>. You will also have to configure
the operating system to provide enough huge pages of the desired size.
The runtime-computed parameter
- <xref linkend="guc-shared-memory-size-in-huge-pages"/> reports the number
- of huge pages required. This parameter can be viewed before starting the
+ <xref linkend="guc-shared-memory-maximum-size-in-huge-pages"/> reports the
+ number of huge pages that must be available in the pool for postmaster
+ startup to succeed. This parameter can be viewed before starting the
server with a <command>postgres</command> command like:
<programlisting>
-$ <userinput>postgres -D $PGDATA -C shared_memory_size_in_huge_pages</userinput>
+$ <userinput>postgres -D $PGDATA -C shared_memory_maximum_size_in_huge_pages</userinput>
3170
$ <userinput>grep ^Hugepagesize /proc/meminfo</userinput>
Hugepagesize: 2048 kB
@@ -1465,7 +1466,7 @@ hugepages-1048576kB hugepages-2048kB
In this example the default is 2MB, but you can also explicitly request
either 2MB or 1GB with <xref linkend="guc-huge-page-size"/> to adapt
the number of pages calculated by
- <varname>shared_memory_size_in_huge_pages</varname>.
+ <varname>shared_memory_maximum_size_in_huge_pages</varname>.
While we need at least <literal>3170</literal> huge pages in this example,
a larger setting would be appropriate if other programs on the machine
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 6b905498337..7c4f4aeba14 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -4333,8 +4333,46 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
Size of the allocation in bytes including padding. For anonymous
allocations, no information about padding is available, so the
<literal>size</literal> and <literal>allocated_size</literal> columns
- will always be equal. Padding is not meaningful for free memory, so
- the columns will be equal in that case also.
+ will always be equal. Padding is not meaningful for free memory, so the
+ columns will be equal in that case also. For resizable allocations which
+ may span multiple memory pages, the padding includes the padding due to
+ page alignment.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>minimum_size</structfield> <type>int8</type>
+ </para>
+ <para>
+ Minimum size in bytes that the resizable allocation can shrink to. Equals
+ <structfield>size</structfield> for fixed-size allocations, anonymous
+ allocations, and free memory.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>maximum_size</structfield> <type>int8</type>
+ </para>
+ <para>
+ Maximum size in bytes that the resizable allocation can grow to. Equals
+ <structfield>size</structfield> for fixed-size allocations, anonymous
+ allocations, and free memory.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>reserved_space</structfield> <type>int8</type>
+ </para>
+ <para>
+ Address space reserved for the allocation in bytes. For resizable
+ structures, this is the total address space reserved to accommodate
+ growth up to <structfield>maximum_size</structfield>, and is greater
+ than or equal to <structfield>allocated_size</structfield>. For
+ fixed-size allocations, anonymous allocations, and free memory this
+ is same as <structfield>allocated_size</structfield>.
</para></entry>
</row>
</tbody>
@@ -4348,6 +4386,21 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
<literal>ShmemRequestHash()</literal>.
</para>
+ <para>
+ Resizable structures are allocations whose size can change while the server
+ is running. <xref linkend="guc-have-resizable-shmem"/> indicates whether the
+ running server supports resizable allocations. Each such allocation has a
+ minimum size give by <structfield>minimum_size</structfield> it can shrink to
+ and a maximum size given by <structfield>maximum_size</structfield> it can
+ grow to; enough address space is reserved up front to accommodate its
+ <structfield>maximum_size</structfield>. The aggregate initial, minimum and
+ maximum sizes of all shared memory allocations are reported by <xref
+ linkend="guc-shared-memory-initial-size"/>, <xref
+ linkend="guc-shared-memory-minimum-size"/> and <xref
+ linkend="guc-shared-memory-maximum-size"/> respectively (and their
+ <literal>_in_huge_pages</literal> counterparts).
+ </para>
+
<para>
By default, the <structname>pg_shmem_allocations</structname> view can be
read only by superusers or roles with privileges of the
diff --git a/doc/src/sgml/xfunc.sgml b/doc/src/sgml/xfunc.sgml
index 050e1e50bec..a17b4401a9c 100644
--- a/doc/src/sgml/xfunc.sgml
+++ b/doc/src/sgml/xfunc.sgml
@@ -3744,6 +3744,98 @@ my_shmem_init(void *arg)
</para>
</sect3>
+ <sect3 id="xfunc-shared-addin-resizable">
+ <title>Resizable shared memory structures</title>
+
+ <para>
+ A resizable memory structure can be requested using
+ <function>ShmemRequestStruct</function> by passing
+ <parameter>.maximum_size</parameter> along with
+ <parameter>.size</parameter>. <parameter>.maximum_size</parameter> is
+ maximum size upto which the structure can grow where as
+ <parameter>.size</parameter> is the initial size of the structure.
+ Optionally, <parameter>.minimum_size</parameter> can be set to the minimum
+ size that the structure can shrink to. While
+ contiguous address space worth <parameter>maximum_size</parameter> is
+ allocated to the structure, only memory worth <parameter>size</parameter>
+ bytes is allocated initially. The <function>init_fn</function> should only
+ initialize the <parameter>size</parameter> amount of memory. The actual
+ memory allocated to this structure at any point in time is given by <link
+ linkend="view-pg-shmem-allocations"><structname>pg_shmem_allocations</structname>.<structfield>allocated_size</structfield></link>
+ and the address space reserved for this structure is given by <link
+ linkend="view-pg-shmem-allocations"><structname>pg_shmem_allocations</structname>.<structfield>reserved_space</structfield></link>.
+ </para>
+
+ <para>
+ The structure can be resized using <function>ShmemResizeStruct</function> by
+ passing it the structure's name and the
+ new size which can be anywhere between <parameter>minimum_size</parameter>
+ and <parameter>maximum_size</parameter>. If the new size is smaller than the
+ current size of the structure, the memory between the new size and current
+ size is freed while keeping the contents of the memory upto new size intact.
+ If the new size is greater than the current size, memory is allocated upto
+ new size while keeping the current contents of the structure intact. The
+ starting address of the structure does not change because of resizing
+ operation.
+ </para>
+
+ <para>
+ <function>ShmemResizeStruct</function> returns <literal>true</literal> on
+ success. If the operating system cannot supply the additional memory
+ required to expand the structure, it emits a <literal>WARNING</literal>
+ naming the structure and the number of bytes that could not be allocated,
+ leaves the structure at its previous size, and returns
+ <literal>false</literal>.
+ </para>
+
+ <para>
+ <function>ShmemResizeStruct</function> does not coordinate with other
+ backends that may be accessing the shared structure. The caller is
+ responsible for ensuring that no backend is accessing the parts of the
+ structure between old size and the new size while
+ <function>ShmemResizeStruct</function> is executing. Similarly after the
+ function has finished, the caller must ensure that each backend knows the
+ new size before accessing the parts of the structure between the old size
+ and the new size. Accessing the range <literal>[current_size,
+ maximum_size)</literal> results in an undefined operating-system dependent
+ behaviour.
+ </para>
+
+ <para>
+ <function>ShmemProtectStruct</function> can be used to make the range beyond
+ the current size inaccessible in a backend's address space, so that stray
+ accesses cause a fault instead of silently succeeding. When a backend
+ starts, protections for all resizable structures are adjusted according to
+ their current sizes, so subsystems need not call
+ <function>ShmemProtectStruct</function> from their per-backend
+ initialization routines. During runtime after calling
+ <function>ShmemResizeStruct</function> from one backend,
+ <function>ShmemProtectStruct</function> should be called from all the
+ backends that may access the structure.
+ </para>
+
+ <para>
+ <productname>PostgreSQL</productname>'s
+ <function>ProcSignalBarrier</function> mechanism (see
+ <filename>src/include/storage/procsignal.h</filename>) may be used to
+ coordinate the resize across backends and adjusting access to the address
+ space.
+ </para>
+
+ <para>
+ This functionality is available only on the platforms which provide the APIs
+ necessary to reserve contiguous address space and to allocate or free memory
+ in that address space on demand. Macro <symbol>HAVE_RESIZABLE_SHMEM</symbol>
+ is defined on such platforms. It can be used to guard code related to
+ resizing a shared memory structure. The functionality is available on with
+ mmap'ed memory, so subsystems which use resizable structures may have to
+ addtionally disable resizable memory usage when <symbol>shared_memory_type</symbol> is not
+ <symbol>SHMEM_TYPE_MMAP</symbol>. A GUC <xref linkend="guc-have-resizable-shmem"/> is set to
+ <literal>on</literal> when this functionality is available in a running
+ server, <literal>off</literal> otherwise.
+ </para>
+ </sect3>
+
<sect3 id="xfunc-shared-addin-dynamic">
<title>Allocating Dynamic Shared Memory after Startup</title>
diff --git a/meson.build b/meson.build
index f4cde249242..4e0c736470c 100644
--- a/meson.build
+++ b/meson.build
@@ -2894,6 +2894,22 @@ decl_checks = [
['timingsafe_bcmp', 'string.h'],
]
+# Linux-specific madvise constants needed for resizable shared memory.
+# Usually we use AC_CHECK_DECLS to check for function declarations, but in this
+# case we are using it to detect existence of constants. These constants are
+# used to define HAVE_RESIZABLE_SHMEM which is used in storage/pg_shmem.h as
+# well as storage/shmem.h. The first abstracts the APIs to allocate shared
+# memory segments from the operating system whereas the second abstracts APIs to
+# allocate shared memory to various subsystems. Since they are related but
+# orthogonal to each other, including any one of them in the other file doesn't
+# make sense. pg_config_manual.h is the only place where HAVE_RESIZABLE_SHMEM
+# can be defined and made available to both without including sys/mman.h. But
+# for that we need constants that indicate the existence of following defines.
+decl_checks += [
+ ['MADV_POPULATE_WRITE', 'sys/mman.h'],
+ ['MADV_REMOVE', 'sys/mman.h'],
+]
+
# Need to check for function declarations for these functions, because
# checking for library symbols wouldn't handle deployment target
# restrictions on macOS
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 2e3886cf9fe..c052776e94c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -589,44 +589,113 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
return true;
}
+/*
+ * Get the page size being used by the shared memory.
+ *
+ * The function should be called only after the shared memory has been setup.
+ */
+size_t
+GetOSPageSize(void)
+{
+ size_t os_page_size;
+
+ Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
+
+ os_page_size = sysconf(_SC_PAGESIZE);
+
+ /* If huge pages are actually in use, use huge page size */
+ if (huge_pages_status == HUGE_PAGES_ON)
+ GetHugePageSize(&os_page_size, NULL);
+
+ return os_page_size;
+}
+
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * initial_size is the amount of memory required at the start of the server.
+ *
+ * *size is the size of the anonymous memory mapping to create. It will be
+ * updated to the actual size of the allocation, if it ends up allocating a
+ * segment that is larger than the requested. When *size > initial_size we want
+ * to reserve a large memory segment while allocating only initial_size amount of
+ * memory at the server start.
+ *
+ * When using huge pages, we make sure that there are enough huge pages
+ * configured to cover the initial_size worth of memory.
*/
static void *
-CreateAnonymousSegment(Size *size)
+CreateAnonymousSegment(Size initial_size, Size *size)
{
Size allocsize = *size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ bool resizable = (initial_size < allocsize);
int mmap_flags = MAP_SHARED | MAP_ANONYMOUS | MAP_HASSEMAPHORE;
+ Assert(initial_size <= allocsize);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
#else
if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
int huge_mmap_flags;
+ bool probe_ok = true;
+ /* Round up the request size to a suitable large value. */
GetHugePageSize(&hugepagesize, &huge_mmap_flags);
-
if (allocsize % hugepagesize != 0)
allocsize = add_size(allocsize, hugepagesize - (allocsize % hugepagesize));
+ if (initial_size % hugepagesize != 0)
+ initial_size = add_size(initial_size, hugepagesize - (initial_size % hugepagesize));
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | huge_mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ /*
+ * When the total amount of space requested is larger than the initial
+ * memory requirement, we do not allocate memory worth the entire
+ * requested space upfront. But we will need to make sure that there
+ * are enough huge pages configured to cover the initial memory
+ * requirement. Otherwise, we will choose huge pages map instead of
+ * falling back to normal pages and the server will not start. Hence
+ * we first try to map and allocate the initial memory requirement. If
+ * it fails we fall back to normal pages. If it succeeds, we unmap the
+ * memory and then try to map the entire requested space without
+ * allocating memory.
+ */
+ if (resizable)
+ {
+ void *probe;
+
+ probe = mmap(NULL, initial_size, PROT_READ | PROT_WRITE,
+ mmap_flags | huge_mmap_flags,
+ -1, 0);
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ probe_ok = false;
+ if (huge_pages == HUGE_PAGES_TRY)
+ elog(DEBUG1,
+ "huge-page probe mmap(%zu) failed, huge pages disabled: %m",
+ initial_size);
+ }
+ else if (munmap(probe, initial_size) != 0)
+ elog(ERROR, "munmap(%p, %zu) huge page probe failed: %m", probe, initial_size);
+ }
+
+ if (probe_ok)
+ {
+ if (resizable)
+ mmap_flags |= MAP_NORESERVE;
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ mmap_flags | huge_mmap_flags, -1, 0);
+ mmap_errno = errno;
+ if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
+ elog(DEBUG1,
+ "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ allocsize);
+ }
}
#endif
@@ -645,6 +714,10 @@ CreateAnonymousSegment(Size *size)
* to non-huge pages.
*/
allocsize = *size;
+
+ /* Use MAP_NORESERVE when we do not need all the memory up front. */
+ if (resizable)
+ mmap_flags |= MAP_NORESERVE;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
mmap_flags, -1, 0);
mmap_errno = errno;
@@ -693,13 +766,18 @@ AnonymousShmemDetach(int status, Datum arg)
* standard header. Also, register an on_shmem_exit callback to release
* the storage.
*
+ * initial_size is the amount of memory required at the start of the server.
+ * Used only when we want to reserve a large memory segment while allocating
+ * only initial_size amount of memory at the server start. That facility is only
+ * available when using anonymous shared memory segments.
+ *
* Dead Postgres segments pertinent to this DataDir are recycled if found, but
* we do not fail upon collision with foreign shmem segments. The idea here
* is to detect and re-use keys that may have been assigned by a crashed
* postmaster or backend.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(Size initial_size, Size size,
PGShmemHeader **shim)
{
IpcMemoryKey NextShmemSegID;
@@ -738,7 +816,7 @@ PGSharedMemoryCreate(Size size,
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
+ AnonymousShmem = CreateAnonymousSegment(initial_size, &size);
AnonymousShmemSize = size;
/* Register on-exit routine to unmap the anonymous segment */
@@ -754,6 +832,10 @@ PGSharedMemoryCreate(Size size,
/* huge pages are only available with mmap */
SetConfigOption("huge_pages_status", "off",
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ /* resizable shared memory is only available with mmap */
+ SetConfigOption("have_resizable_shmem", "off",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
}
/*
@@ -991,3 +1073,148 @@ PGSharedMemoryDetach(void)
AnonymousShmem = NULL;
}
}
+
+/*
+ * Make sure that the memory of given size from the given address is released.
+ *
+ * The address and size are expected to be page aligned.
+ *
+ * Only supported on platforms that support anonymous shared memory.
+ *
+ * On success returns true. false if system call fails.
+ */
+bool
+PGSharedMemoryEnsureFreed(void *addr, Size size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be freed"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(addr == (void *) TYPEALIGN(os_page_size, addr));
+ Assert(size == TYPEALIGN(os_page_size, size));
+ }
+#endif
+
+ if (madvise(addr, size, MADV_REMOVE) == -1)
+ {
+ ereport(WARNING, errmsg("could not free shared memory: %m"));
+ return false;
+ }
+
+ return true;
+#endif
+}
+
+/*
+ * Make sure that the memory of given size from the given address is allocated.
+ *
+ * The address and size are expected to be page aligned. The caller is
+ * responsible for ensuring that the range is already writable; this function
+ * only populates it.
+ *
+ * Only supported on platforms that support anonymous shared memory.
+ *
+ * On success returns true, false if system call fails.
+ */
+bool
+PGSharedMemoryEnsureAllocated(void *addr, Size size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be allocated at runtime"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(addr == (void *) TYPEALIGN(os_page_size, addr));
+ Assert(size == TYPEALIGN(os_page_size, size));
+ }
+#endif
+
+ if (madvise(addr, size, MADV_POPULATE_WRITE) == -1)
+ {
+ ereport(WARNING, errmsg("could not allocate shared memory: %m"));
+ return false;
+ }
+
+ return true;
+#endif
+}
+
+/*
+ * Set memory protection on the given region of shared memory.
+ *
+ * Makes [rw_start, rw_end) readable and writable, and [rw_end, prot_end)
+ * inaccessible.
+ *
+ * All addresses are expected to be page aligned.
+ *
+ * Only supported on platforms that support resizable shared memory.
+ *
+ * On success returns true, false if system call fails.
+ */
+bool
+PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be protected at runtime"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(rw_start == (void *) TYPEALIGN(os_page_size, rw_start));
+ Assert(rw_end == (void *) TYPEALIGN(os_page_size, rw_end));
+ Assert(prot_end == (void *) TYPEALIGN(os_page_size, prot_end));
+ }
+#endif
+ Assert(rw_end >= rw_start);
+
+ if (rw_end > rw_start)
+ {
+ if (mprotect(rw_start, (char *) rw_end - (char *) rw_start,
+ PROT_READ | PROT_WRITE) != 0)
+ {
+ ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ return false;
+ }
+ }
+
+ if (prot_end > rw_end)
+ {
+ if (mprotect(rw_end, (char *) prot_end - (char *) rw_end,
+ PROT_NONE) != 0)
+ {
+ ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ return false;
+ }
+ }
+
+ return true;
+#endif
+}
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 794e4fcb2ad..d266cdbe226 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -200,11 +200,13 @@ EnableLockPagesPrivilege(int elevel)
/*
* PGSharedMemoryCreate
*
- * Create a shared memory segment of the given size and initialize its
- * standard header.
+ * Create a shared memory segment of the given size and initialize its standard
+ * header. initial_size is only relevant when we want to create a larger memory
+ * segment with only a part of it being allocated initially. We don't support
+ * that on Windows, so we ignore it.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(Size initial_size, Size size,
PGShmemHeader **shim)
{
void *memAddress;
@@ -648,3 +650,59 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
}
return true;
}
+
+/*
+ * Get the page size used by the shared memory.
+ *
+ * The function should be called only after the shared memory has been setup.
+ */
+size_t
+GetOSPageSize(void)
+{
+ SYSTEM_INFO sysinfo;
+ size_t os_page_size;
+
+ Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
+
+ GetSystemInfo(&sysinfo);
+ os_page_size = sysinfo.dwPageSize;
+
+ /* If huge pages are actually in use, use huge page size */
+ if (huge_pages_status == HUGE_PAGES_ON)
+ GetHugePageSize(&os_page_size, NULL);
+
+ return os_page_size;
+}
+
+/*
+ * PGSharedMemoryEnsureFreed / PGSharedMemoryEnsureAllocated
+ *
+ * Not supported on Windows. These are only meaningful on platforms with
+ * resizable shared memory (mmap + madvise).
+ */
+bool
+PGSharedMemoryEnsureFreed(void *addr, Size size)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
+
+bool
+PGSharedMemoryEnsureAllocated(void *addr, Size size)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
+
+bool
+PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
diff --git a/src/backend/postmaster/auxprocess.c b/src/backend/postmaster/auxprocess.c
index ad4bf4bd2a8..07a3b5c5923 100644
--- a/src/backend/postmaster/auxprocess.c
+++ b/src/backend/postmaster/auxprocess.c
@@ -69,33 +69,43 @@ AuxiliaryProcessMainCommon(void)
BaseInit();
/*
- * Prevent consuming interrupts between setting ProcSignalInit and setting
- * the initial local data checksum value. If a barrier is emitted, and
- * absorbed, before local cached state is initialized the state transition
- * can be invalid.
- */
- HOLD_INTERRUPTS();
-
- ProcSignalInit(NULL, 0);
-
- /*
- * Initialize a local cache of the data_checksum_version, to be updated by
- * the procsignal-based barriers.
+ * Prevent consuming interrupts between ProcSignalInit() and the
+ * initialization of local states that are kept in sync with shared memory
+ * via procsignal-based barriers. If a barrier is emitted, and absorbed,
+ * before local cached state is initialized the state transition can be
+ * invalid.
*
- * This intentionally happens after initializing the procsignal, otherwise
- * we might miss a state change. This means we can get a barrier for the
- * state we've just initialized - but it can happen only once.
+ * These initializations intentionally happen after ProcSignalInit(),
+ * otherwise we might miss a state change. This means we may also receive
+ * a barrier for the state we've just initialized, but it can happen only
+ * once.
*
* The postmaster (which is what gets forked into the new child process)
- * does not handle barriers, therefore it may not have the current value
- * of LocalDataChecksumState value (it'll have the value read from the
- * control file, which may be arbitrarily old).
+ * does not handle barriers, therefore its local states may not reflect
+ * the current state of the shared memory.
*
* NB: Even if the postmaster handled barriers, the value might still be
* stale, as it might have changed after this process forked.
*/
+ HOLD_INTERRUPTS();
+
+ ProcSignalInit(NULL, 0);
+
+ /*
+ * LocalDataChecksumState inherited from Postmaster will have the value
+ * read from the control file, which may be arbitrarily old. Update it.
+ */
InitLocalDataChecksumState();
+ /*
+ * Refresh per-backend protections for resizable shmem structures, in case
+ * these structures have been resized since the startup. Structures are
+ * expected to be kept in sync by respective subsystems using
+ * procsignal-based barriers. But we modify their protections en-masse
+ * here, rather than relying on individual subsystems to do it.
+ */
+ ShmemReprotectResizableStructs();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index e149a738c8d..e5f20c4603e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -52,31 +52,50 @@ RequestAddinShmemSpace(Size size)
/*
* CalculateShmemSize
* Calculates the amount of shared memory needed.
+ *
+ * - `initial` is the amount of memory needed when the server startup.
+ * - `min` is the minimum amount of memory needed when all the resizable
+ * structures are shrunk to their respective minimum sizes.
+ * - `max` is the maximum amount of memory needed when all the resizable
+ * structures are expanded to their respective maximum sizes. It is also the
+ * size of address space that must be reserved for the shared memory segment.
+ *
+ * When no resizable structures are requested, all three totals are identical.
+ *
+ * We take some care to ensure that the total size request doesn't overflow
+ * size_t. If this gets through, we don't need to be so careful during the
+ * actual allocation phase.
*/
-Size
-CalculateShmemSize(void)
+void
+CalculateShmemSize(size_t *initial, size_t *min, size_t *max)
{
- Size size;
+ size_t initial_req;
+ size_t min_req;
+ size_t max_req;
+ size_t fixed_addins;
+
+ ShmemGetRequestedSize(&initial_req, &min_req, &max_req);
/*
* Size of the Postgres shared-memory block is estimated via moderately-
* accurate estimates for the big hogs, plus 100K for the stuff that's too
* small to bother with estimating.
*
- * We take some care to ensure that the total size request doesn't
- * overflow size_t. If this gets through, we don't need to be so careful
- * during the actual allocation phase.
+ * Also include additional requested shmem from preload libraries.
+ *
+ * These are not resizable, so they contribute equally to all three
+ * totals.
*/
- size = 100000;
- size = add_size(size, ShmemGetRequestedSize());
-
- /* include additional requested shmem from preload libraries */
- size = add_size(size, total_addin_request);
+ fixed_addins = add_size(100000, total_addin_request);
- /* might as well round it off to a multiple of a typical page size */
- size = add_size(size, 8192 - (size % 8192));
+ *initial = add_size(initial_req, fixed_addins);
+ *min = add_size(min_req, fixed_addins);
+ *max = add_size(max_req, fixed_addins);
- return size;
+ /* might as well round each off to a multiple of a typical page size */
+ *initial = add_size(*initial, 8192 - (*initial % 8192));
+ *min = add_size(*min, 8192 - (*min % 8192));
+ *max = add_size(*max, 8192 - (*max % 8192));
}
#ifdef EXEC_BACKEND
@@ -121,18 +140,24 @@ CreateSharedMemoryAndSemaphores(void)
{
PGShmemHeader *shim;
PGShmemHeader *seghdr;
- Size size;
+ size_t initial_size;
+ size_t min_size;
+ size_t max_size;
Assert(!IsUnderPostmaster);
/* Compute the size of the shared-memory block */
- size = CalculateShmemSize();
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+ CalculateShmemSize(&initial_size, &min_size, &max_size);
+ elog(DEBUG3, "invoking IpcMemoryCreate(initial size=%zu, minimum size=%zu, maximum size=%zu)",
+ initial_size, min_size, max_size);
/*
- * Create the shmem segment
+ * Create the shmem segment.
+ *
+ * Reserve enough shared address space to accommodate every requested
+ * structure grown to its maximum size.
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(initial_size, max_size, &shim);
/*
* Make sure that huge pages are never reported as "unknown" while the
@@ -189,32 +214,49 @@ void
InitializeShmemGUCs(void)
{
char buf[64];
- Size size_b;
- Size size_mb;
- Size hp_size;
+ size_t initial_b;
+ size_t min_b;
+ size_t max_b;
+ size_t hp_size;
+ size_t size_mb;
- /*
- * Calculate the shared memory size and round up to the nearest megabyte.
- */
- size_b = CalculateShmemSize();
- size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
+ CalculateShmemSize(&initial_b, &min_b, &max_b);
+
+ /* Round each size up to the nearest megabyte. */
+ size_mb = add_size(initial_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
- SetConfigOption("shared_memory_size", buf,
+ SetConfigOption("shared_memory_initial_size", buf,
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
- /*
- * Calculate the number of huge pages required.
- */
+ size_mb = add_size(min_b, (1024 * 1024) - 1) / (1024 * 1024);
+ sprintf(buf, "%zu", size_mb);
+ SetConfigOption("shared_memory_minimum_size", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ size_mb = add_size(max_b, (1024 * 1024) - 1) / (1024 * 1024);
+ sprintf(buf, "%zu", size_mb);
+ SetConfigOption("shared_memory_maximum_size", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ /* Calculate the number of huge pages required for each size. */
GetHugePageSize(&hp_size, NULL);
if (hp_size != 0)
{
- Size hp_required;
+ size_t hp_required;
+
+ hp_required = initial_b / hp_size + (initial_b % hp_size != 0);
+ sprintf(buf, "%zu", hp_required);
+ SetConfigOption("shared_memory_initial_size_in_huge_pages", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ hp_required = min_b / hp_size + (min_b % hp_size != 0);
+ sprintf(buf, "%zu", hp_required);
+ SetConfigOption("shared_memory_minimum_size_in_huge_pages", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
- hp_required = size_b / hp_size;
- if (size_b % hp_size != 0)
- hp_required = add_size(hp_required, 1);
+ hp_required = max_b / hp_size + (max_b % hp_size != 0);
sprintf(buf, "%zu", hp_required);
- SetConfigOption("shared_memory_size_in_huge_pages", buf,
+ SetConfigOption("shared_memory_maximum_size_in_huge_pages", buf,
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
}
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index 228871d2525..8700dd58799 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -19,11 +19,11 @@
* methods). The routines in this file are used for allocating and
* binding to shared memory data structures.
*
- * This module provides facilities to allocate fixed-size structures in shared
- * memory, for things like variables shared between all backend processes.
- * Each such structure has a string name to identify it, specified when it is
- * requested. shmem_hash.c provides a shared hash table implementation on top
- * of that.
+ * This module provides facilities to allocate fixed-size as well as resizable
+ * structures in shared memory, for things like variables shared between all
+ * backend processes. Each such structure has a string name to identify it,
+ * specified when it is requested. shmem_hash.c provides a shared hash table
+ * implementation on top of fixed-size structures.
*
* Shared memory areas should usually not be allocated after postmaster
* startup, although we do allow small allocations later for the benefit of
@@ -102,6 +102,23 @@
* (*options->ptr), and calls the attach_fn callback, if any, for additional
* per-backend setup.
*
+ * Resizable shared memory structures
+ * ----------------------------------
+ *
+ * In order to allocate resizable shared memory structures, set
+ * ShmemRequestStructOpts::maximum_size to the maximum size that the structure
+ * can grow to. The address space for the maximum size will be reserved at
+ * startup, but memory is allocated or freed as the structure grows or shrinks
+ * respectively. ShmemRequestStructOpts::size should be set to the initial size
+ * of the structure, which is the amount of memory allocated at the startup.
+ * Optionally, ShmemRequestStructOpts::minimum_size can be set to the minimum
+ * size that the structure can shrink to. After startup, the structure can be
+ * resized by calling ShmemResizeStruct(). ShmemResizeStruct() enforces that the
+ * new size is within [minimum_size, maximum_size].
+ *
+ * While resizable structures can be created after the startup, the memory
+ * available for them is quite limited.
+ *
* Legacy ShmemInitStruct()/ShmemInitHash() functions
* --------------------------------------------------
*
@@ -268,6 +285,10 @@ typedef struct
void *location; /* location in shared mem */
Size size; /* # bytes requested for the structure */
Size allocated_size; /* # bytes actually allocated */
+ Size minimum_size; /* the minimum size the structure can shrink
+ * to */
+ Size maximum_size; /* the maximum size the structure can grow to */
+ Size reserved_space; /* the total address space reserved */
} ShmemIndexEnt;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
@@ -276,6 +297,8 @@ static bool firstNumaTouch = true;
static void CallShmemCallbacksAfterStartup(const ShmemCallbacks *callbacks);
static void InitShmemIndexEntry(ShmemRequest *request);
static bool AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok);
+static Size EstimateAllocatedSize(ShmemIndexEnt *entry);
+static void ShmemProtectStructInternal(ShmemIndexEnt *entry);
Datum pg_numa_available(PG_FUNCTION_ARGS);
@@ -341,25 +364,63 @@ ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind)
if (options->name == NULL)
elog(ERROR, "shared memory request is missing 'name' option");
+#ifndef HAVE_RESIZABLE_SHMEM
+ if (options->maximum_size > 0)
+ elog(ERROR, "resizable shared memory is not supported on this platform");
+ if (options->minimum_size > 0)
+ elog(ERROR, "resizable shared memory is not supported on this platform");
+#else
+ if (options->maximum_size > 0 && shared_memory_type != SHMEM_TYPE_MMAP)
+ elog(ERROR, "resizable shared memory requires shared_memory_type = mmap");
+#endif
+
if (IsUnderPostmaster)
{
if (options->size <= 0 && options->size != SHMEM_ATTACH_UNKNOWN_SIZE)
elog(ERROR, "invalid size %zd for shared memory request for \"%s\"",
options->size, options->name);
+ if (options->minimum_size < 0 && options->minimum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ elog(ERROR, "invalid minimum_size %zd for shared memory request for \"%s\"",
+ options->minimum_size, options->name);
+ if (options->maximum_size < 0 && options->maximum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ elog(ERROR, "invalid maximum_size %zd for shared memory request for \"%s\"",
+ options->maximum_size, options->name);
}
else
{
- if (options->size == SHMEM_ATTACH_UNKNOWN_SIZE)
+ if (options->size == SHMEM_ATTACH_UNKNOWN_SIZE ||
+ options->minimum_size == SHMEM_ATTACH_UNKNOWN_SIZE ||
+ options->maximum_size == SHMEM_ATTACH_UNKNOWN_SIZE)
elog(ERROR, "SHMEM_ATTACH_UNKNOWN_SIZE cannot be used during startup");
if (options->size <= 0)
elog(ERROR, "invalid size %zd for shared memory request for \"%s\"",
options->size, options->name);
+ if (options->minimum_size < 0)
+ elog(ERROR, "invalid minimum_size %zd for shared memory request for \"%s\"",
+ options->minimum_size, options->name);
+ if (options->maximum_size < 0)
+ elog(ERROR, "invalid maximum_size %zd for shared memory request for \"%s\"",
+ options->maximum_size, options->name);
}
if (options->alignment != 0 && pg_nextpower2_size_t(options->alignment) != options->alignment)
elog(ERROR, "invalid alignment %zu for shared memory request for \"%s\"",
options->alignment, options->name);
+ if (options->minimum_size > 0 && options->size != SHMEM_ATTACH_UNKNOWN_SIZE &&
+ options->minimum_size > options->size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have minimum size (%zd) less than or equal to size (%zd)",
+ options->name, options->minimum_size, options->size);
+
+ if (options->maximum_size > 0 && options->size > options->maximum_size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have maximum size (%zd) greater than size (%zd)",
+ options->name, options->maximum_size, options->size);
+
+ if (options->minimum_size > 0 && options->maximum_size > 0 &&
+ options->minimum_size > options->maximum_size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have minimum size (%zd) less than or equal to maximum size (%zd)",
+ options->name, options->minimum_size, options->maximum_size);
+
/* Check that we're in the right state */
if (shmem_request_state != SRS_REQUESTING)
elog(ERROR, "ShmemRequestStruct can only be called from a shmem_request callback");
@@ -381,36 +442,70 @@ ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind)
}
/*
- * ShmemGetRequestedSize() --- estimate the total size of all registered shared
- * memory structures.
+ * ShmemGetRequestedSize() --- estimate total size of all registered shared
+ * memory structures.
+ *
+ * Returns three totals:
+ * - initial - the sum of initial sizes of the requested structures which is the
+ * total amount of memory required at the startup.
+ * - min - the total of minimum sizes of structures
+ * - max - the sum of maximum sizes of the structures, which is the address
+ * space that must be reserved.
+ *
+ * When there are no resizable structures or on the platforms that do not
+ * support resizable structures, all three totals are the same.
*
* This is called at postmaster startup, before the shared memory segment has
* been created.
*/
-size_t
-ShmemGetRequestedSize(void)
+void
+ShmemGetRequestedSize(size_t *initial, size_t *min, size_t *max)
{
- size_t size;
+ size_t initial_size;
+ size_t min_size;
+ size_t max_size;
/* memory needed for the ShmemIndex */
- size = hash_estimate_size(list_length(pending_shmem_requests) + SHMEM_INDEX_ADDITIONAL_SIZE,
- sizeof(ShmemIndexEnt));
- size = CACHELINEALIGN(size);
+ initial_size = hash_estimate_size(list_length(pending_shmem_requests) + SHMEM_INDEX_ADDITIONAL_SIZE,
+ sizeof(ShmemIndexEnt));
+ initial_size = CACHELINEALIGN(initial_size);
+ min_size = initial_size;
+ max_size = initial_size;
/* memory needed for all the requested areas */
foreach_ptr(ShmemRequest, request, pending_shmem_requests)
{
size_t alignment = request->options->alignment;
+ size_t req_min;
+ size_t req_max;
+ size_t req_initial = request->options->size;
+
+ if (request->options->maximum_size > 0)
+ {
+ req_min = request->options->minimum_size;
+ req_max = request->options->maximum_size;
+ }
+ else
+ {
+ req_min = req_initial;
+ req_max = req_initial;
+ }
/* pad the start address for alignment like ShmemAllocRaw() does */
if (alignment < PG_CACHE_LINE_SIZE)
alignment = PG_CACHE_LINE_SIZE;
- size = TYPEALIGN(alignment, size);
+ initial_size = TYPEALIGN(alignment, initial_size);
+ min_size = TYPEALIGN(alignment, min_size);
+ max_size = TYPEALIGN(alignment, max_size);
- size = add_size(size, request->options->size);
+ initial_size = add_size(initial_size, req_initial);
+ min_size = add_size(min_size, req_min);
+ max_size = add_size(max_size, req_max);
}
- return size;
+ *initial = initial_size;
+ *min = min_size;
+ *max = max_size;
}
/*
@@ -431,6 +526,23 @@ ShmemInitRequested(void)
* Initialize the ShmemIndex entries and perform basic initialization of
* all the requested memory areas. There are no concurrent processes yet,
* so no need for locking.
+ *
+ * TODO: If we have resizable structures, we will mmap with MAP_NORESERVE.
+ * In case there is not enough memory to cover the initial sizes of the
+ * shared structures, the mmap will succeed but the initialization will
+ * fail with SIGBUS. Instead we should allocate the memory worth the
+ * initial size of each structure using madvise(MADV_WRITE_POPOULATE). But
+ * instead of doing that for each structure, we should combine contiguous
+ * structures and do it once for the whole range. The algorithm to use is
+ * as follows 1. start from the first request 2. note the start address of
+ * the first structure in the range 3. walk upto a resizable structure or
+ * the last request, noting the initial end address of the last structure
+ * in the range 4. call madvise(MADV_WRITE_POPULATE) for the range (start
+ * address, end address) 5. Next structure becomes the first structure in
+ * the next range, repeat from step 2
+ *
+ * If there are no resizable structure, no need to do anything, as all the
+ * memory is reserved during mmap itself.
*/
foreach_ptr(ShmemRequest, request, pending_shmem_requests)
{
@@ -514,6 +626,7 @@ InitShmemIndexEntry(ShmemRequest *request)
ShmemIndexEnt *index_entry;
bool found;
size_t allocated_size;
+ size_t requested_size;
void *structPtr;
/* look it up in the shmem index */
@@ -531,10 +644,19 @@ InitShmemIndexEntry(ShmemRequest *request)
}
/*
- * We inserted the entry to the shared memory index. Allocate requested
- * amount of shared memory for it, and initialize the index entry.
+ * We inserted the entry to the shared memory index. Allocate requested
+ * amount of address space in the shared memory segment for it, and do
+ * basic initializion. The memory gets allocated during initialization as
+ * the corresponding memory pages are written to. Allocate enough space
+ * for a resizable structure to grow to its maximum size. It is expected
+ * that the initialization callback will use only as much memory as the
+ * initial size of the resizable structure. (Well, if it doesn't, more
+ * memory will be allocated initially than expected, no further harm is
+ * done.)
*/
- structPtr = ShmemAllocRaw(request->options->size,
+ requested_size = request->options->maximum_size > 0 ?
+ request->options->maximum_size : request->options->size;
+ structPtr = ShmemAllocRaw(requested_size,
request->options->alignment,
&allocated_size);
if (structPtr == NULL)
@@ -543,13 +665,36 @@ InitShmemIndexEntry(ShmemRequest *request)
hash_search(ShmemIndex, name, HASH_REMOVE, NULL);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("not enough shared memory for data structure"
+ errmsg("not enough shared memory space for data structure"
" \"%s\" (%zd bytes requested)",
- name, request->options->size)));
+ name, requested_size)));
}
index_entry->size = request->options->size;
index_entry->allocated_size = allocated_size;
index_entry->location = structPtr;
+ index_entry->reserved_space = allocated_size;
+ if (request->options->maximum_size > 0)
+ {
+ index_entry->minimum_size = request->options->minimum_size;
+ index_entry->maximum_size = request->options->maximum_size;
+
+ /* Adjust allocated size of a resizable structure. */
+ index_entry->allocated_size = EstimateAllocatedSize(index_entry);
+
+ /*
+ * Protect the unused part of the reserved address space for a
+ * resizable structure. If the structure has same minimum and maximum
+ * size, it is effectively a fixed-size structure without any unused
+ * space; no protection is required.
+ */
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ ShmemProtectStructInternal(index_entry);
+ }
+ else
+ {
+ index_entry->minimum_size = request->options->size;
+ index_entry->maximum_size = request->options->size;
+ }
/* Initialize depending on the kind of shmem area it is */
switch (request->kind)
@@ -594,7 +739,7 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
return false;
}
- /* Check that the size in the index matches the request */
+ /* Check that the sizes in the index match the request. */
if (index_entry->size != request->options->size &&
request->options->size != SHMEM_ATTACH_UNKNOWN_SIZE)
{
@@ -604,6 +749,43 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
name, index_entry->size, request->options->size)));
}
+ /*
+ * For resizable structures, also check that minimum_size and maximum_size
+ * match. For fixed-size structures, these are derived (set to size) in
+ * the index entry and not meaningful in the request.
+ */
+ if (request->options->maximum_size != 0)
+ {
+ if (index_entry->minimum_size != request->options->minimum_size &&
+ request->options->minimum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ {
+ ereport(ERROR,
+ errmsg("shared memory struct \"%s\" was created with"
+ " different minimum_size: existing %zu, requested %zu",
+ name, index_entry->minimum_size,
+ request->options->minimum_size));
+ }
+
+ if (index_entry->maximum_size != request->options->maximum_size &&
+ request->options->maximum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ {
+ ereport(ERROR,
+ errmsg("shared memory struct \"%s\" was created with"
+ " different maximum_size: existing %zu, requested %zu",
+ name, index_entry->maximum_size,
+ request->options->maximum_size));
+ }
+ }
+ else
+ {
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ elog(ERROR, "shared memory struct \"%s\" was created as resizable, but requested as fixed-size",
+ name);
+ }
+
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ ShmemProtectStructInternal(index_entry);
+
/*
* Re-establish the caller's pointer variable, or do other actions to
* attach depending on the kind of shmem area it is.
@@ -625,6 +807,291 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
return true;
}
+/*
+ * Estimate the actual memory allocated for a resizable structure.
+ *
+ * ... based on the assumption that the memory is allocated in pages.
+ *
+ * The memory pages covered by the current size of a resizable structure are
+ * considered to be allocated. The memory page where the maximal structure ends
+ * also hosts the next structure, unless the maximal structure ends on a page
+ * boundary. Hence that page is allocated because of the next structure. The
+ * memory pages between the page where the current structure ends and the page
+ * where the next structure starts remain unallocated. Thus the memory allocated
+ * for a resizable structure can be estimated as the total address space
+ * reserved for the structure minus the unallocated memory pages between the
+ * current end and the next structure.
+ */
+static Size
+EstimateAllocatedSize(ShmemIndexEnt *entry)
+{
+ Size page_size = GetOSPageSize();
+ char *align_end = (char *) TYPEALIGN(page_size, (char *) entry->location + entry->size);
+ char *floor_max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) entry->location + entry->maximum_size);
+
+ Assert(entry->maximum_size >= entry->size);
+ Assert(entry->reserved_space >= entry->maximum_size);
+
+ if (align_end < floor_max_end)
+ return entry->reserved_space - (floor_max_end - align_end);
+
+ return entry->reserved_space;
+}
+
+/*
+ * ShmemResizeStruct() --- resize a resizable shared memory structure.
+ *
+ * The new size must be within [minimum_size, maximum_size]. If the structure
+ * is being shrunk, the memory pages that are no longer needed are freed. If
+ * the structure is being expanded, the memory pages that are needed for the
+ * new size are allocated. See EstimateAllocatedSize() for explanation of which
+ * pages are allocated for a resizable structure.
+ *
+ * The caller must ensure that no other backend is accessing the part of the
+ * structure between the old size and the new size while this function is
+ * running. It should also ensure that all backends that may access the
+ * structure have observed the new size before they access the range between
+ * the old size and the new size.
+ *
+ * If we can not allocate memory pages when expanding the structure, this
+ * function will return false. On success it returns true. Instead
+ * of returning false, an error is raised if we can not free the memory pages
+ * when shrinking the structure, which should not happen in practice.
+ */
+bool
+ShmemResizeStruct(const char *name, Size new_size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ pg_unreachable();
+#else
+ ShmemIndexEnt *result;
+ bool found;
+ Size page_size = GetOSPageSize();
+ char *new_end;
+ bool success = true;
+
+ Assert(new_size > 0);
+
+ /*
+ * Resizable shared memory structures are only supported with mmap'ed
+ * memory.
+ */
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
+
+ /* look it up in the shmem index */
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+ result = (ShmemIndexEnt *) hash_search(ShmemIndex, name, HASH_FIND, &found);
+ if (!found)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shmem struct \"%s\" is not initialized", name));
+
+ Assert(result);
+
+ if (result->minimum_size == result->maximum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shared memory struct \"%s\" is not resizable", name));
+
+ if (new_size < result->minimum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("cannot shrink shared memory structure \"%s\" below minimum size"
+ "(requested %zu bytes, minimum %zu bytes)",
+ name, new_size, result->minimum_size));
+
+ if (result->maximum_size < new_size)
+ ereport(ERROR,
+ errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough address space is reserved for resizing structure \"%s\""
+ " (required %zu bytes, reserved %zu bytes)",
+ name, new_size, result->maximum_size));
+
+ /*
+ * A structure requires memory pages from the page containing the start of
+ * the structure and current end of the structure to be allocated. When
+ * expanding, we make sure that pages are allocated up to the new end.
+ * When shrinking, release memory pages beyond the new end, but not the
+ * page containing maximal end of the structure, as it may be used by the
+ * next structure.
+ *
+ * We do not consider the current end of the structure as it simplifies
+ * the calculations. Instead we rely on the underlying APIs not to touch
+ * the memory pages that will not be affected by the change in size.
+ */
+ new_end = (char *) TYPEALIGN(page_size, (char *) result->location + new_size);
+ if (new_size < result->size)
+ {
+ char *max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+
+ if (max_end > new_end)
+ {
+ if (!PGSharedMemoryEnsureFreed(new_end, max_end - new_end))
+ ereport(ERROR,
+ errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not free %zu bytes of shared memory from structure \"%s\"",
+ result->size - new_size, name));
+ }
+ }
+ else if (new_size > result->size)
+ {
+ char *struct_start = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location);
+
+ if (new_end > struct_start)
+ {
+ ShmemIndexEnt entry_copy = *result;
+
+ /*
+ * Allocating memory pages in the expanded range may require the
+ * corresponding address space to have read-write access.
+ */
+ entry_copy.size = new_size;
+ ShmemProtectStructInternal(&entry_copy);
+
+ if (!PGSharedMemoryEnsureAllocated(struct_start, new_end - struct_start))
+ {
+ ShmemProtectStructInternal(result);
+ ereport(WARNING,
+ errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("could not allocate %zu bytes of shared memory to structure \"%s\"",
+ new_size - result->size, name));
+ success = false;
+ }
+ }
+ }
+
+ /* Update shmem index entry. */
+ if (success)
+ {
+ result->size = new_size;
+ result->allocated_size = EstimateAllocatedSize(result);
+ }
+
+ LWLockRelease(ShmemIndexLock);
+
+ return success;
+#endif
+}
+
+/*
+ * ShmemProtectStruct() --- protect the unused portion of the given resizable
+ * structure.
+ *
+ * Makes the region beyond the current size up to maximum_size inaccessible, and
+ * ensures the region up to the current size is readable and writable. Depending
+ * upon the platform, the protection honours the page boundaries. So it may be
+ * more permissible than strictly needed.
+ *
+ * This function only affects the calling backend's address space. After each
+ * ShmemResizeStruct(), every backend that may access the structure should call
+ * this function before its next access. When backends are still in the middle
+ * of that round of ShmemProtectStruct() calls, ShmemResizeStruct() should not
+ * be called.
+ */
+void
+ShmemProtectStruct(const char *name)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ ShmemIndexEnt *result;
+ bool found;
+
+ LWLockAcquire(ShmemIndexLock, LW_SHARED);
+
+ result = (ShmemIndexEnt *) hash_search(ShmemIndex, name, HASH_FIND, &found);
+ if (!found)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shmem struct \"%s\" is not initialized", name));
+
+ if (result->minimum_size == result->maximum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shared memory struct \"%s\" is not resizable", name));
+
+ ShmemProtectStructInternal(result);
+
+ LWLockRelease(ShmemIndexLock);
+#endif
+}
+
+/*
+ * ShmemProtectStructInternal --- same as ShmemProtectStruct()
+ *
+ * ..., but called when the ShmemIndexEnt of the struct is available. The caller
+ * should hold ShmemIndexLock if required.
+ */
+static void
+ShmemProtectStructInternal(ShmemIndexEnt *entry)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ Size page_size = GetOSPageSize();
+ char *rw_start;
+ char *rw_end;
+ char *prot_end;
+
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
+ Assert(entry->minimum_size != entry->maximum_size);
+
+ /* Make at least [location, location+size) readable and writable */
+ rw_start = (char *) TYPEALIGN_DOWN(page_size, entry->location);
+ rw_end = (char *) TYPEALIGN(page_size,
+ (char *) entry->location + entry->size);
+
+ /*
+ * Make remaining portion inaccessible while making sure that the portion
+ * after maximum_size is not affected since it may be used by other
+ * structures.
+ */
+ prot_end = (char *) TYPEALIGN_DOWN(page_size,
+ (char *) entry->location + entry->maximum_size);
+
+ if (!PGSharedMemoryProtect(rw_start, rw_end, prot_end))
+ ereport(ERROR,
+ errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not protect shared memory structure \"%s\"", entry->key));
+#endif
+}
+
+/*
+ * ShmemReprotectResizableStructs() --- re-apply per-backend protection to every
+ * resizable shmem structure.
+ *
+ * On !EXEC_BACKEND platforms, a new backend inherits shmem mappings and their
+ * per-process protections from the postmaster. If any resizable structure has
+ * been resized since postmaster start, the inherited protections no longer
+ * match the current size. This is called early in per-backend startup so the
+ * new backend sees the protections according to the current sizes of the
+ * resizable structures.
+ */
+void
+ShmemReprotectResizableStructs(void)
+{
+#ifdef HAVE_RESIZABLE_SHMEM
+ HASH_SEQ_STATUS status;
+ ShmemIndexEnt *entry;
+
+ LWLockAcquire(ShmemIndexLock, LW_SHARED);
+ hash_seq_init(&status, ShmemIndex);
+ while ((entry = (ShmemIndexEnt *) hash_seq_search(&status)) != NULL)
+ {
+ if (entry->minimum_size != entry->maximum_size)
+ ShmemProtectStructInternal(entry);
+ }
+ LWLockRelease(ShmemIndexLock);
+#endif
+}
+
/*
* InitShmemAllocator() --- set up basic pointers to shared memory.
*
@@ -731,6 +1198,9 @@ InitShmemAllocator(PGShmemHeader *seghdr)
Assert(!found);
result->size = ShmemAllocator->index_size;
result->allocated_size = ShmemAllocator->index_size;
+ result->minimum_size = result->size;
+ result->maximum_size = result->size;
+ result->reserved_space = result->allocated_size;
result->location = ShmemAllocator->index;
}
}
@@ -1047,7 +1517,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Datum
pg_get_shmem_allocations(PG_FUNCTION_ARGS)
{
-#define PG_GET_SHMEM_SIZES_COLS 4
+#define PG_GET_SHMEM_SIZES_COLS 7
ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
HASH_SEQ_STATUS hstat;
ShmemIndexEnt *ent;
@@ -1069,7 +1539,20 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
- named_allocated += ent->allocated_size;
+ values[4] = Int64GetDatum(ent->minimum_size);
+ values[5] = Int64GetDatum(ent->maximum_size);
+ values[6] = Int64GetDatum(ent->reserved_space);
+
+ /*
+ * Anonymous areas are allocated in the area remaining after all named
+ * areas have been allocated. Thus the amount of shared memory
+ * allocated for anonymous areas can be calculated as the total amount
+ * of space allocated minus the amount of space allocated for named
+ * areas. The amount of free shared memory at the end of the segment
+ * can be calculated as the total size of the segment minus the total
+ * amount of space allocated.
+ */
+ named_allocated += ent->reserved_space;
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
values, nulls);
@@ -1080,6 +1563,9 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
nulls[1] = true;
values[2] = Int64GetDatum(ShmemAllocator->free_offset - named_allocated);
values[3] = values[2];
+ values[4] = values[2];
+ values[5] = values[2];
+ values[6] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
@@ -1088,6 +1574,9 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
nulls[1] = false;
values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemAllocator->free_offset);
values[3] = values[2];
+ values[4] = values[2];
+ values[5] = values[2];
+ values[6] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
LWLockRelease(ShmemIndexLock);
@@ -1274,23 +1763,9 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size
pg_get_shmem_pagesize(void)
{
- Size os_page_size;
-#ifdef WIN32
- SYSTEM_INFO sysinfo;
-
- GetSystemInfo(&sysinfo);
- os_page_size = sysinfo.dwPageSize;
-#else
- os_page_size = sysconf(_SC_PAGESIZE);
-#endif
-
Assert(IsUnderPostmaster);
- Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
-
- if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
- return os_page_size;
+ return GetOSPageSize();
}
Datum
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 9d6e69175a5..780cdb5e9ff 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -575,6 +575,14 @@ InitProcess(void)
if (IsUnderPostmaster)
AttachSharedMemoryStructs();
#endif
+
+ /*
+ * Update access to the address space occupied by the resizable shared
+ * structures. We do it here so that the structures can be accessed safely
+ * by this backend. But we will do this again after ProcSignalInit() for
+ * the reasons mentioned there.
+ */
+ ShmemReprotectResizableStructs();
}
/*
@@ -756,6 +764,14 @@ InitAuxiliaryProcess(void)
if (IsUnderPostmaster)
AttachSharedMemoryStructs();
#endif
+
+ /*
+ * Update access to the address space occupied by the resizable shared
+ * structures. We do it here so that the structures can be accessed safely
+ * by this backend. But we will do this again after ProcSignalInit() for
+ * the reasons mentioned there.
+ */
+ ShmemReprotectResizableStructs();
}
/*
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 3d8c9bdebd5..815b865aa38 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -787,6 +787,16 @@ InitPostgres(const char *in_dbname, Oid dboid,
*/
InitLocalDataChecksumState();
+ /*
+ * Refresh per-backend protections for resizable shmem structures. Usually
+ * the subsystems using resizable shared structures will use
+ * ProcSignalBarrier mechanism to coordinate resizing which would involve
+ * adjusting the protections as well. Like InitLocalDataChecksumState()
+ * above, this must run after ProcSignalInit so as not to miss a barrier
+ * for protection change.
+ */
+ ShmemReprotectResizableStructs();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index 5afa3e47552..db64f8b2601 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -1226,6 +1226,13 @@
max => '1000.0',
},
+{ name => 'have_resizable_shmem', type => 'bool', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows whether the running server supports resizable shared memory.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE',
+ variable => 'have_resizable_shmem_enabled',
+ boot_val => 'HAVE_RESIZABLE_SHMEM_ENABLED',
+},
+
{ name => 'hba_file', type => 'string', context => 'PGC_POSTMASTER', group => 'FILE_LOCATIONS',
short_desc => 'Sets the server\'s "hba" configuration file.',
flags => 'GUC_SUPERUSER_ONLY',
@@ -2731,20 +2738,58 @@
max => 'INT_MAX / 2',
},
-{ name => 'shared_memory_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
- short_desc => 'Shows the size of the server\'s main shared memory area (rounded up to the nearest MB).',
+{ name => 'shared_memory_initial_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the amount of memory allocated at server startup in the main shared memory area (rounded up to the nearest MB).',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_initial_size_mb',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_initial_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the number of huge pages needed in the main shared memory area at server startup.',
+ long_desc => '-1 means huge pages are not supported.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_initial_size_in_huge_pages',
+ boot_val => '-1',
+ min => '-1',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_maximum_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the size of the main shared memory area and maximum memory that can be allocated in that area (rounded up to the nearest MB).',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_maximum_size_mb',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_maximum_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the maximum number of huge pages needed in the main shared memory area.',
+ long_desc => '-1 means huge pages are not supported.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_maximum_size_in_huge_pages',
+ boot_val => '-1',
+ min => '-1',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_minimum_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the minimum amount of memory required in the main shared memory area (rounded up to the nearest MB).',
flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
- variable => 'shared_memory_size_mb',
+ variable => 'shared_memory_minimum_size_mb',
boot_val => '0',
min => '0',
max => 'INT_MAX',
},
-{ name => 'shared_memory_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
- short_desc => 'Shows the number of huge pages needed for the main shared memory area.',
+{ name => 'shared_memory_minimum_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the minimum number of huge pages needed in the main shared memory area.',
long_desc => '-1 means huge pages are not supported.',
flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
- variable => 'shared_memory_size_in_huge_pages',
+ variable => 'shared_memory_minimum_size_in_huge_pages',
boot_val => '-1',
min => '-1',
max => 'INT_MAX',
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index c6d9b2a6f89..8b18d1f8eca 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -644,8 +644,12 @@ static int max_index_keys;
static int max_identifier_length;
static int block_size;
static int segment_size;
-static int shared_memory_size_mb;
-static int shared_memory_size_in_huge_pages;
+static int shared_memory_initial_size_mb;
+static int shared_memory_minimum_size_mb;
+static int shared_memory_maximum_size_mb;
+static int shared_memory_initial_size_in_huge_pages;
+static int shared_memory_minimum_size_in_huge_pages;
+static int shared_memory_maximum_size_in_huge_pages;
static int wal_block_size;
static int num_os_semaphores;
static int effective_wal_level = WAL_LEVEL_REPLICA;
@@ -665,6 +669,13 @@ static bool assert_enabled = DEFAULT_ASSERT_ENABLED;
#endif
static bool exec_backend_enabled = EXEC_BACKEND_ENABLED;
+#ifdef HAVE_RESIZABLE_SHMEM
+#define HAVE_RESIZABLE_SHMEM_ENABLED true
+#else
+#define HAVE_RESIZABLE_SHMEM_ENABLED false
+#endif
+static bool have_resizable_shmem_enabled = HAVE_RESIZABLE_SHMEM_ENABLED;
+
static char *recovery_target_timeline_string;
static char *recovery_target_string;
static char *recovery_target_xid_string;
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 66c3c9a04cf..e1f5945a971 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8753,8 +8753,8 @@
{ oid => '5052', descr => 'allocations from the main shared memory segment',
proname => 'pg_get_shmem_allocations', prorows => '50', proretset => 't',
provolatile => 'v', prorettype => 'record', proargtypes => '',
- proallargtypes => '{text,int8,int8,int8}', proargmodes => '{o,o,o,o}',
- proargnames => '{name,off,size,allocated_size}',
+ proallargtypes => '{text,int8,int8,int8,int8,int8,int8}', proargmodes => '{o,o,o,o,o,o,o}',
+ proargnames => '{name,off,size,allocated_size,minimum_size,maximum_size,reserved_space}',
prosrc => 'pg_get_shmem_allocations',
proacl => '{POSTGRES=X,pg_read_all_stats=X}' },
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 661c4a9b168..661c12e9bf3 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,14 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `MADV_POPULATE_WRITE', and to 0
+ if you don't. */
+#undef HAVE_DECL_MADV_POPULATE_WRITE
+
+/* Define to 1 if you have the declaration of `MADV_REMOVE', and to 0 if you
+ don't. */
+#undef HAVE_DECL_MADV_REMOVE
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/include/pg_config_manual.h b/src/include/pg_config_manual.h
index 521b49b8888..ab944babe2b 100644
--- a/src/include/pg_config_manual.h
+++ b/src/include/pg_config_manual.h
@@ -131,6 +131,20 @@
#define EXEC_BACKEND
#endif
+/*
+ * HAVE_RESIZABLE_SHMEM indicates whether resizable shared memory structures are
+ * supported. The implementation requires Linux-specific madvise constants
+ * (MADV_REMOVE and MADV_POPULATE_WRITE) and existence of mprotect() API.
+ *
+ * TODO: We may want to remove EXEC_BACKEND from the condition to test attaching
+ * and resizing resizable shared memory structures in EXEC_BACKEND mode. Windows
+ * will anyway won't have HAVE_RESIZABLE_SHMEM defined since it won't have
+ * MADV_REMOVE and MADV_POPULATE_WRITE.
+ */
+#if HAVE_DECL_MADV_REMOVE && HAVE_DECL_MADV_POPULATE_WRITE && !defined(EXEC_BACKEND)
+#define HAVE_RESIZABLE_SHMEM
+#endif
+
/*
* USE_POSIX_FADVISE controls whether Postgres will attempt to use the
* posix_fadvise() kernel call. Usually the automatic configure tests are
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index b205b00e7a1..46ae87fe863 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -78,7 +78,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
extern void RegisterBuiltinShmemCallbacks(void);
-extern Size CalculateShmemSize(void);
+extern void CalculateShmemSize(size_t *initial, size_t *min, size_t *max);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 10c7b065861..50474c2d579 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -85,10 +85,14 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(Size size,
+extern PGShmemHeader *PGSharedMemoryCreate(Size initial_size, Size max_size,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
+extern bool PGSharedMemoryEnsureFreed(void *addr, Size size);
+extern bool PGSharedMemoryEnsureAllocated(void *addr, Size size);
+extern bool PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern size_t GetOSPageSize(void);
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 43b636868c7..947debcede4 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -57,6 +57,22 @@ typedef struct ShmemStructOpts
*/
size_t alignment;
+ /*
+ * Minimum size this structure can shrink to. Should be set to 0 for
+ * fixed-size structures.
+ */
+ ssize_t minimum_size;
+
+ /*
+ * Maximum size this structure can grow upto in future. The memory is not
+ * allocated right away but the corresponding address space is reserved so
+ * that memory can be mapped to it when the structure grows. Typically
+ * should be used for large resizable structures which need several pages
+ * worth of contiguous memory. Should be set to 0 for fixed-size
+ * structures.
+ */
+ ssize_t maximum_size;
+
/*
* When the shmem area is initialized or attached to, pointer to it is
* stored in *ptr. It usually points to a global variable, used to access
@@ -168,6 +184,9 @@ typedef struct ShmemCallbacks
extern void RegisterShmemCallbacks(const ShmemCallbacks *callbacks);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemResizeStruct(const char *name, Size new_size);
+extern void ShmemProtectStruct(const char *name);
+extern void ShmemReprotectResizableStructs(void);
/*
* These macros provide syntactic sugar for calling the underlying functions
diff --git a/src/include/storage/shmem_internal.h b/src/include/storage/shmem_internal.h
index 8746b614fa3..6c8c81812a8 100644
--- a/src/include/storage/shmem_internal.h
+++ b/src/include/storage/shmem_internal.h
@@ -36,7 +36,7 @@ extern void ResetShmemAllocator(void);
extern void ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind);
-extern size_t ShmemGetRequestedSize(void);
+extern void ShmemGetRequestedSize(size_t *initial, size_t *min, size_t *max);
extern void ShmemInitRequested(void);
#ifdef EXEC_BACKEND
extern void ShmemAttachRequested(void);
diff --git a/src/test/modules/test_shmem/meson.build b/src/test/modules/test_shmem/meson.build
index fb4bf328b8f..1f795aae7eb 100644
--- a/src/test/modules/test_shmem/meson.build
+++ b/src/test/modules/test_shmem/meson.build
@@ -27,7 +27,8 @@ tests += {
'bd': meson.current_build_dir(),
'tap': {
'tests': [
- 't/001_late_shmem_alloc.pl',
+ 't/001_fixed_shmem_struct.pl',
+ 't/002_resizable_shmem_struct.pl',
],
},
}
diff --git a/src/test/modules/test_shmem/t/001_late_shmem_alloc.pl b/src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
similarity index 58%
rename from src/test/modules/test_shmem/t/001_late_shmem_alloc.pl
rename to src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
index 5cf07d071ec..38e687e33e9 100644
--- a/src/test/modules/test_shmem/t/001_late_shmem_alloc.pl
+++ b/src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
@@ -56,5 +56,45 @@ else
);
}
+###
+# Test that a fixed-size shared memory structure cannot be resized.
+# Only relevant on platforms that support resizable shmem.
+###
+my $have_resizable_shmem =
+ $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+if ($have_resizable_shmem)
+{
+ # Try expanding the fixed-size structure
+ my ($ret, $stdout, $stderr) =
+ $node->psql("postgres", "SELECT test_shmem_resize_fixed(1000);");
+ isnt($ret, 0, "expanding a fixed-size structure fails");
+ like(
+ $stderr,
+ qr/is not resizable/,
+ "expand error message mentions not resizable");
+
+ # Try shrinking the fixed-size structure
+ ($ret, $stdout, $stderr) =
+ $node->psql("postgres", "SELECT test_shmem_resize_fixed(1);");
+ isnt($ret, 0, "shrinking a fixed-size structure fails");
+ like(
+ $stderr,
+ qr/is not resizable/,
+ "shrink error message mentions not resizable");
+}
+
+###
+# Test that minimum_size and maximum_size equal size for a fixed-size structure
+# in pg_shmem_allocations.
+###
+is( $node->safe_psql(
+ 'postgres',
+ "SELECT minimum_size = size AND maximum_size = size FROM pg_shmem_allocations WHERE name = 'test_shmem area';"
+ ),
+ 't',
+ "fixed-size structure has minimum_size = maximum_size = size");
+
$node->stop;
+
done_testing();
diff --git a/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl b/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
new file mode 100644
index 00000000000..5a972383517
--- /dev/null
+++ b/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
@@ -0,0 +1,450 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Test resizable shared memory functionality, both when loaded at startup via
+# shared_preload_libraries and when loaded after startup (late allocation).
+
+# Verify that enough shared memory is allocated to cover the resizable_shmem
+# structure at its current size but does not exceed the memory required by the
+# current sizes of all shared memory structures. We expect that the backend
+# where we run the query will have touched the entire resizable_shmem structure,
+# so that all the memory pages covering the resizable structure are mapped to
+# the backend's address space.
+#
+# Since we have configured the server so that resizable shared struture
+# dominates the main shared memory segment, the total memory allocated to other
+# shared memory structures does not result in false positive tests below.
+sub check_shmem_usage
+{
+ my ($session, $label, $node) = @_;
+
+ my $shmem_usage =
+ $session->query_safe('SELECT test_shmem_usage();', verbose => 0);
+ my $total_alloc = $node->safe_psql('postgres',
+ "SELECT sum(allocated_size) FROM pg_shmem_allocations;");
+ my $resizable_alloc = $node->safe_psql('postgres',
+ "SELECT allocated_size FROM pg_shmem_allocations WHERE name = 'resizable_shmem';"
+ );
+
+ diag
+ "$label: shmem_usage=$shmem_usage, resizable_shmem allocated=$resizable_alloc, sum(allocated_size)=$total_alloc";
+ ok( $shmem_usage <= $total_alloc,
+ "$label: allocated shared memory does not exceed total allocated size"
+ );
+ ok($shmem_usage >= $resizable_alloc,
+ "$label: shared memory usage covers the resizable_shmem allocation");
+}
+
+# Test a resize operation: resize, verify old data, write new data, verify
+# new data, and check shmem usage. Returns updated ($num_entries, $value).
+sub test_resize
+{
+ my ($node, $prefix, $old_num_entries, $old_value, $new_num_entries,
+ $new_value, $label)
+ = @_;
+
+ $label = "$prefix: $label";
+
+ my $session1 = $node->background_psql('postgres');
+ my $session2 = $node->background_psql('postgres');
+
+ $session1->query_safe("SELECT resizable_shmem_resize($new_num_entries);",
+ verbose => 0);
+
+ # Old data should still be intact in the (possibly smaller) area
+ my $readable_entries =
+ ($new_num_entries < $old_num_entries)
+ ? $new_num_entries
+ : $old_num_entries;
+ is( $session1->query_safe(
+ "SELECT resizable_shmem_read($readable_entries, $old_value);",
+ verbose => 0),
+ 't',
+ "old data readable after $label");
+
+ $session2->query_safe("SELECT resizable_shmem_write($new_value);",
+ verbose => 0);
+ is( $session1->query_safe(
+ "SELECT resizable_shmem_read($new_num_entries, $new_value);",
+ verbose => 0),
+ 't',
+ "new data readable after $label");
+
+ check_shmem_usage($session1, "$label (session 1)", $node);
+ check_shmem_usage($session2, "$label (session 2)", $node);
+
+ $session1->quit;
+ $session2->quit;
+
+ return ($new_num_entries, $new_value);
+}
+
+# Verify that reads or writes past the current size, but within the reserved
+# maximum, fault when they reach the protected region.
+sub test_fault_beyond_size
+{
+ my ($node, $initial_entries, $prefix) = @_;
+
+ # Enable restart_after_crash to test postmaster driven restart with
+ # resizable shared memory.
+ $node->safe_psql('postgres',
+ 'ALTER SYSTEM SET restart_after_crash = on;');
+ $node->reload;
+
+ for my $mode ('write', 'read')
+ {
+ my ($ret, $stdout, $stderr) = $node->psql('postgres',
+ "SELECT resizable_shmem_access_beyond_size('$mode');");
+ ok($ret != 0, "$prefix: $mode past current size crashes the backend");
+ like(
+ $stderr,
+ qr/server closed the connection unexpectedly|connection to server was lost/,
+ "$prefix: $mode crash reports lost connection");
+
+ $node->poll_query_until('postgres', 'SELECT 1', '1')
+ or die "server did not come back after $mode crash";
+ }
+
+ is( $node->safe_psql(
+ 'postgres', "SELECT resizable_shmem_read($initial_entries, 0);"),
+ 't',
+ "$prefix: read succeeds after crash recovery");
+
+ $node->safe_psql('postgres', 'ALTER SYSTEM RESET restart_after_crash;');
+ $node->reload;
+}
+
+# Run the full suite of resizable shared memory tests on the given node.
+sub run_resizable_tests
+{
+ my ($node, $initial_entries, $max_entries, $prefix) = @_;
+ my $have_resizable_shmem =
+ $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+ my $num_entries = $initial_entries;
+
+ # Basic read/write should work on all platforms
+ my $value = 100;
+ $node->safe_psql('postgres', "SELECT resizable_shmem_write($value);");
+ is( $node->safe_psql(
+ 'postgres', "SELECT resizable_shmem_read($num_entries, $value);"),
+ 't',
+ "$prefix: data read after write successful");
+
+ if ($have_resizable_shmem)
+ {
+ # Initial structure state
+ my $session1 = $node->background_psql('postgres');
+ my $session2 = $node->background_psql('postgres');
+
+ $value = 100;
+ # Write and read the initial set of entries.
+ $session1->query_safe("SELECT resizable_shmem_write($value);",
+ verbose => 0);
+ is( $session2->query_safe(
+ "SELECT resizable_shmem_read($num_entries, $value);",
+ verbose => 0),
+ 't',
+ "$prefix: data read after write successful");
+ check_shmem_usage($session1, "$prefix: initial write (session 1)",
+ $node);
+ check_shmem_usage($session2, "$prefix: initial write (session 2)",
+ $node);
+ $session1->quit;
+ $session2->quit;
+
+ # Verify no other structure is resizable
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT count(*) FROM pg_shmem_allocations WHERE name <> 'resizable_shmem' AND maximum_size <> minimum_size;"
+ ),
+ '0',
+ "$prefix: no other resizable structures");
+
+ # Resize to maximum
+ ($num_entries, $value) =
+ test_resize($node, $prefix, $num_entries, $value,
+ $max_entries, 500, 'resize to maximum');
+
+ # Shrink to 75% of max
+ my $shrink_entries = int($max_entries * 3 / 4);
+ ($num_entries, $value) =
+ test_resize($node, $prefix, $num_entries, $value,
+ $shrink_entries, 999, 'shrinking');
+
+ # Resize to the same size (no-op)
+ ($num_entries, $value) =
+ test_resize($node, $prefix, $num_entries, $value,
+ $num_entries, 1999, 'no-op resize');
+
+ # Shrink to minimum i.e. zero entries and grow back
+ ($num_entries, $value) =
+ test_resize($node, $prefix, $num_entries, $value, 0, 0,
+ 'shrink to minimum');
+ ($num_entries, $value) =
+ test_resize($node, $prefix, $num_entries, $value,
+ $initial_entries, 2999, 'grow back from minimum');
+
+ # Test resize failure (attempt to resize beyond max - should fail)
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres',
+ "SELECT resizable_shmem_resize(" . ($max_entries * 2) . ");");
+ ok( $ret != 0 || $stderr =~ /ERROR/,
+ "$prefix: Resize beyond maximum fails");
+
+ # Resize to a size below minimum_size must fail.
+ ($ret, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT resizable_shmem_resize(-1);');
+ ok($ret != 0, "$prefix: resize below minimum_size fails");
+ like(
+ $stderr,
+ qr/cannot shrink shared memory structure "resizable_shmem" below minimum size/,
+ "$prefix: resize-below-minimum error comes from ShmemResizeStruct"
+ );
+
+ # The fault test relies on a hole being present between the current end
+ # of the structure and its maximal end. Skip when the structure does not
+ # span multiple pages.
+ my $spans_pages = $node->safe_psql(
+ 'postgres', qq{
+ SELECT (maximum_size - minimum_size) >= test_shmem_pagesize()
+ FROM pg_shmem_allocations WHERE name = 'resizable_shmem';
+ });
+ if ($spans_pages ne 't')
+ {
+ diag
+ "$prefix: skipping fault-beyond-size test: resizable_shmem does not span multiple shmem pages";
+ }
+ else
+ {
+ test_fault_beyond_size($node, $initial_entries, $prefix);
+ }
+ }
+ else
+ {
+ # On unsupported platforms, resizing should fail with a clear error
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres',
+ "SELECT resizable_shmem_resize($num_entries);");
+ ok($ret != 0, "$prefix: resize fails on unsupported platform");
+ like(
+ $stderr,
+ qr/not supported/,
+ "$prefix: resize error mentions not supported");
+ }
+}
+
+# Check the runtime-computed shared_memory_{initial,minimum,maximum}_size GUC
+# invariants. min <= initial <= max must always hold. When a resizable
+# structure has been registered on a server that supports resizable shared
+# memory structures, min must additionally be strictly less than max;
+# otherwise all three GUCs must be equal.
+sub check_shmem_size_gucs
+{
+ my ($node, $label) = @_;
+ my $pgdata = $node->data_dir;
+
+ my $get = sub {
+ my ($guc) = @_;
+ my ($stdout, $stderr) =
+ run_command([ 'postgres', '-D' => $pgdata, '-C' => $guc ]);
+
+ return $stdout;
+ };
+
+ my $have_resizable_shmem = $get->('have_resizable_shmem');
+ my $ini = 0 + $get->('shared_memory_initial_size');
+ my $min = 0 + $get->('shared_memory_minimum_size');
+ my $max = 0 + $get->('shared_memory_maximum_size');
+ my $have_resizable_struct = ($get->('resizable_shmem.max_entries') ne '');
+
+ ok($min <= $ini && $ini <= $max,
+ "$label: shared_memory size GUCs in expected order");
+
+ if ($have_resizable_struct && $have_resizable_shmem eq 'on')
+ {
+ ok($min < $max, "$label: min < max with resizable structures");
+ }
+ else
+ {
+ ok( $min == $ini && $ini == $max,
+ "$label: all shared_memory size GUCs equal when no resizable structures"
+ );
+ }
+}
+
+# Log the runtime shared_memory_{initial,minimum,maximum}_size GUCs and huge
+# pages usage information for easier debugging.
+sub diag_shmem_sizes
+{
+ my ($node, $label) = @_;
+
+ my $vals = $node->safe_psql(
+ 'postgres', q{
+ SELECT format('initial=%s minimum=%s maximum=%s huge_pages_status=%s huge_page_size=%s shmem_page_size=%s',
+ current_setting('shared_memory_initial_size'),
+ current_setting('shared_memory_minimum_size'),
+ current_setting('shared_memory_maximum_size'),
+ current_setting('huge_pages_status'),
+ current_setting('huge_page_size'),
+ test_shmem_pagesize());
+ });
+ diag "$label: $vals";
+}
+
+### Set up a test node.
+#
+# Configure minimal shared memory so that the resizable_shmem structure dominates
+# and any unexpected increase is easy to detect.
+#
+# If we turn on huge pages and the machine where the test is running does not
+# have huge pages available, the test will fail midway because it will not be
+# able to allocate memory pages when expanding the resizable_shmem structure.
+# Hence we turn off huge pages for this test. The test outputs the GUCs
+# shared_memory_{initial,minimum,maximum}_size and information about huge pages.
+# By provisioning enough huge pages, and by changing huge_pages = try/on, the
+# test can be run with huge pages enabled.
+###
+my $node = PostgreSQL::Test::Cluster->new('resizable_shmem');
+$node->init;
+
+$node->append_conf('postgresql.conf', 'huge_pages = off');
+$node->append_conf('postgresql.conf', 'shared_buffers = 128kB');
+$node->append_conf('postgresql.conf', 'max_connections = 5');
+$node->append_conf('postgresql.conf', 'max_worker_processes = 0');
+$node->append_conf('postgresql.conf', 'max_wal_senders = 0');
+$node->append_conf('postgresql.conf', 'max_prepared_transactions = 0');
+$node->append_conf('postgresql.conf', 'max_locks_per_transaction = 10');
+$node->append_conf('postgresql.conf', 'max_pred_locks_per_transaction = 10');
+$node->append_conf('postgresql.conf', 'wal_buffers = 32kB');
+
+###
+# Test 1: Startup allocation via shared_preload_libraries
+###
+my $startup_initial = 25 * 1024 * 1024;
+my $startup_max = 100 * 1024 * 1024;
+
+$node->append_conf('postgresql.conf',
+ 'shared_preload_libraries = test_shmem');
+$node->append_conf('postgresql.conf',
+ "resizable_shmem.initial_entries = $startup_initial");
+$node->append_conf('postgresql.conf',
+ "resizable_shmem.max_entries = $startup_max");
+
+check_shmem_size_gucs($node, 'startup preload');
+
+$node->start;
+$node->safe_psql('postgres', 'CREATE EXTENSION test_shmem;');
+diag_shmem_sizes($node, 'startup');
+run_resizable_tests($node, $startup_initial, $startup_max, 'startup');
+
+my $have_resizable_shmem =
+ $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+###
+# Test 2: Late allocation (loaded after startup, not in shared_preload_libraries).
+# Use much smaller sizes since only ~100KB of shared memory is available for
+# structures allocated after startup.
+###
+my $late_initial = 5 * 1024;
+my $late_max = 12 * 1024;
+
+$node->safe_psql(
+ 'postgres', qq{
+ ALTER SYSTEM RESET shared_preload_libraries;
+ ALTER SYSTEM SET resizable_shmem.initial_entries = $late_initial;
+ ALTER SYSTEM SET resizable_shmem.max_entries = $late_max;
+});
+$node->safe_psql('postgres', 'DROP EXTENSION test_shmem;');
+$node->restart;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION test_shmem;');
+diag_shmem_sizes($node, 'late');
+run_resizable_tests($node, $late_initial, $late_max, 'late');
+
+###
+# Test sysv shared memory does not support resizable shmem. Only relevant on
+# platforms that support resizable shmem (HAVE_RESIZABLE_SHMEM), since the
+# module only sets maximum_size in that case.
+###
+if ($have_resizable_shmem)
+{
+ ###
+ # Test 3: Verify that CREATE EXTENSION fails with sysv shared memory
+ # when loaded after startup (not in shared_preload_libraries).
+ ###
+ $node->safe_psql('postgres', 'DROP EXTENSION test_shmem;');
+
+ # Remove settings that would cause the library to auto-load at startup:
+ # shared_preload_libraries and module-prefixed GUCs. ALTER SYSTEM RESET
+ # only affects postgresql.auto.conf, so we must use adjust_conf to remove
+ # from postgresql.conf.
+ $node->adjust_conf('postgresql.conf', 'shared_preload_libraries', undef);
+ $node->adjust_conf('postgresql.conf', 'resizable_shmem.initial_entries',
+ undef);
+ $node->adjust_conf('postgresql.conf', 'resizable_shmem.max_entries',
+ undef);
+ $node->adjust_conf('postgresql.auto.conf', 'shared_preload_libraries',
+ undef);
+ $node->adjust_conf('postgresql.auto.conf',
+ 'resizable_shmem.initial_entries', undef);
+ $node->adjust_conf('postgresql.auto.conf', 'resizable_shmem.max_entries',
+ undef);
+ $node->safe_psql(
+ 'postgres', qq{
+ ALTER SYSTEM SET shared_memory_type = 'sysv';
+ });
+
+ $node->stop;
+
+ check_shmem_size_gucs($node, 'sysv');
+
+ $node->start;
+
+ is($node->safe_psql('postgres', 'SHOW have_resizable_shmem;'),
+ 'off',
+ 'have_resizable_shmem reports off with shared_memory_type = sysv');
+
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', 'CREATE EXTENSION test_shmem;');
+ ok($ret != 0, 'CREATE EXTENSION fails with resizable shmem on sysv');
+ like(
+ $stderr,
+ qr/resizable shared memory requires shared_memory_type = mmap/,
+ 'CREATE EXTENSION error mentions shared_memory_type = mmap requirement'
+ );
+
+ ###
+ # Test 4: Verify that resizable structures are also rejected with sysv
+ # shared memory when loaded at startup via shared_preload_libraries.
+ ###
+ $node->safe_psql(
+ 'postgres', qq{
+ ALTER SYSTEM SET shared_preload_libraries = 'test_shmem';
+ ALTER SYSTEM SET resizable_shmem.initial_entries = $startup_initial;
+ ALTER SYSTEM SET resizable_shmem.max_entries = $startup_max;
+ });
+ $node->stop;
+
+ ok(!$node->start(fail_ok => 1),
+ 'server fails to start with resizable shmem on sysv');
+
+ my $log = slurp_file($node->logfile);
+ like(
+ $log,
+ qr/resizable shared memory requires shared_memory_type = mmap/,
+ 'log mentions shared_memory_type = mmap requirement');
+}
+
+done_testing();
+
+#TODO: Add a test to test the behavior of resizable shared memory when memory
+#allocation fails by simulating a memory allocation failure through injection
+#points. Add an injection point in ShmemResizeStruct in expansion branch or in
+#the underlying platform specific memory allocation function.
diff --git a/src/test/modules/test_shmem/test_shmem--1.0.sql b/src/test/modules/test_shmem/test_shmem--1.0.sql
index 2d01fd9256c..eb695604b1d 100644
--- a/src/test/modules/test_shmem/test_shmem--1.0.sql
+++ b/src/test/modules/test_shmem/test_shmem--1.0.sql
@@ -4,6 +4,61 @@
\echo Use "CREATE EXTENSION test_shmem" to load this file. \quit
+-- ===================================================================
+-- Fixed-size shared memory structure
+-- ===================================================================
+
CREATE FUNCTION get_test_shmem_attach_count()
RETURNS pg_catalog.int4 STRICT
AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION test_shmem_resize_fixed(pg_catalog.int4)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+-- ===================================================================
+-- Resizable shared memory structure
+-- ===================================================================
+
+-- Function to resize the test structure in the shared memory
+CREATE FUNCTION resizable_shmem_resize(new_entries integer)
+RETURNS bool
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to write data to all entries in the test structure in shared memory
+-- Writing all the entries makes sure that the memory is actually allocated and
+-- mapped to the process, so that we can later measure the memory usage.
+CREATE FUNCTION resizable_shmem_write(entry_value integer)
+RETURNS void
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to verify that specified number of initial entries have expected value.
+-- Reading all the entries makes sure that the memory is actually mapped to the
+-- process, so that we can later measure the memory usage.
+CREATE FUNCTION resizable_shmem_read(entry_count integer, entry_value integer)
+RETURNS boolean
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to report memory mapped against the main shared memory segment in
+-- the backend where this function runs.
+CREATE FUNCTION test_shmem_usage()
+RETURNS bigint
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to get the shared memory page size
+CREATE FUNCTION test_shmem_pagesize()
+RETURNS integer
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to crash the backend by walking entries past the current size up to
+-- the reserved maximum, reading or writing each one as decided by mode.
+CREATE FUNCTION resizable_shmem_access_beyond_size(mode text)
+RETURNS integer
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
diff --git a/src/test/modules/test_shmem/test_shmem.c b/src/test/modules/test_shmem/test_shmem.c
index 9bd4012b435..83f8ea9cc3c 100644
--- a/src/test/modules/test_shmem/test_shmem.c
+++ b/src/test/modules/test_shmem/test_shmem.c
@@ -1,11 +1,10 @@
/*-------------------------------------------------------------------------
*
* test_shmem.c
- * Helpers to test shmem allocation routines
+ * Helpers to test shmem management routines
*
- * Test basic memory allocation in an extension module. One notable feature
- * that is not exercised by any other module in the repository is the
- * allocating (non-DSM) shared memory after postmaster startup.
+ * Test fixed-size and resizable shared memory structures created during
+ * postmaster startup and after startup respectively.
*
* Copyright (c) 2020-2026, PostgreSQL Global Development Group
*
@@ -17,13 +16,26 @@
#include "postgres.h"
+#include <limits.h>
+
+#include "commands/extension.h"
#include "fmgr.h"
#include "miscadmin.h"
+#include "storage/fd.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
+#include "utils/builtins.h"
+#include "utils/guc.h"
PG_MODULE_MAGIC;
+
+/* ----------------------------------------------------------------
+ * Fixed-size shared memory structure
+ * ----------------------------------------------------------------
+ */
+
typedef struct TestShmemData
{
int value;
@@ -35,17 +47,6 @@ static TestShmemData *TestShmem;
static bool attached_or_initialized = false;
-static void test_shmem_request(void *arg);
-static void test_shmem_init(void *arg);
-static void test_shmem_attach(void *arg);
-
-static const ShmemCallbacks TestShmemCallbacks = {
- .flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP,
- .request_fn = test_shmem_request,
- .init_fn = test_shmem_init,
- .attach_fn = test_shmem_attach,
-};
-
static void
test_shmem_request(void *arg)
{
@@ -60,6 +61,17 @@ static void
test_shmem_init(void *arg)
{
elog(LOG, "init callback called");
+
+ /*
+ * Reset the per-process flag and the shared "initialized" marker during
+ * postmaster induced restart.
+ */
+ if (!IsUnderPostmaster)
+ {
+ attached_or_initialized = false;
+ TestShmem->initialized = false;
+ }
+
if (TestShmem->initialized)
elog(ERROR, "shmem area already initialized");
TestShmem->initialized = true;
@@ -82,12 +94,12 @@ test_shmem_attach(void *arg)
attached_or_initialized = true;
}
-void
-_PG_init(void)
-{
- elog(LOG, "test_shmem module's _PG_init called");
- RegisterShmemCallbacks(&TestShmemCallbacks);
-}
+static const ShmemCallbacks TestShmemCallbacks = {
+ .flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP,
+ .request_fn = test_shmem_request,
+ .init_fn = test_shmem_init,
+ .attach_fn = test_shmem_attach,
+};
PG_FUNCTION_INFO_V1(get_test_shmem_attach_count);
Datum
@@ -99,3 +111,453 @@ get_test_shmem_attach_count(PG_FUNCTION_ARGS)
elog(ERROR, "shmem area not yet initialized");
PG_RETURN_INT32(TestShmem->attach_count);
}
+
+/*
+ * Attempt to resize the fixed-size shared memory structure. This should
+ * fail because the structure was not allocated with a maximum_size.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_resize_fixed);
+Datum
+test_shmem_resize_fixed(PG_FUNCTION_ARGS)
+{
+ int32 new_size = PG_GETARG_INT32(0);
+
+ ShmemResizeStruct("test_shmem area", new_size);
+ PG_RETURN_VOID();
+}
+
+
+/* ----------------------------------------------------------------
+ * Resizable shared memory structure
+ * ----------------------------------------------------------------
+ */
+
+/*
+ * The test module may be loaded after postmaster startup in which case only
+ * 100K of shared memory is available for the extension. Keep the default
+ * initial and maximum sizes small enough to fit in that space.
+ */
+#define TEST_INITIAL_ENTRIES_DEFAULT 1
+#define TEST_MAX_ENTRIES_DEFAULT 1024
+
+#define TEST_ENTRY_SIZE sizeof(int32) /* Size of each entry */
+
+/*
+ * Resizable test data structure stored in shared memory.
+ *
+ * The test performs resizing, reads or writes, only one at a time and never
+ * concurrently. Hence, there is no need for locks in the test structure.
+ */
+typedef struct TestResizableShmemStruct
+{
+ /* Metadata */
+ int32 num_entries; /* Number of entries that can fit */
+
+ /* Data area - variable size */
+ int32 data[FLEXIBLE_ARRAY_MEMBER];
+} TestResizableShmemStruct;
+
+static TestResizableShmemStruct *resizable_shmem = NULL;
+
+/* GUC variables controlling the size of the test structure */
+static int test_initial_entries;
+static int test_max_entries;
+
+/* Whether to use SHMEM_ATTACH_UNKNOWN_SIZE when attaching to the shared memory */
+/* TODO: We may use opaque_arg to pass this value to the request function.*/
+static bool use_unknown_size = false;
+
+/*
+ * Request shared memory resources.
+ */
+static void
+resizable_shmem_request(void *arg)
+{
+ Size initial_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(test_initial_entries, TEST_ENTRY_SIZE));
+
+/*
+ * Create resizable structure on the platforms which support it. Otherwise create
+ * as a fixed-size structure. Other way would be to conditionally include
+ * .maximum_size in the call to ShmemRequestStruct().
+ */
+#ifdef HAVE_RESIZABLE_SHMEM
+ Size max_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(test_max_entries, TEST_ENTRY_SIZE));
+ Size min_size = offsetof(TestResizableShmemStruct, data);
+#else
+ Size max_size = 0;
+ Size min_size = 0;
+#endif
+
+ ShmemRequestStruct(.name = "resizable_shmem",
+ .size = use_unknown_size ? SHMEM_ATTACH_UNKNOWN_SIZE : initial_size,
+ .minimum_size = min_size,
+ .maximum_size = max_size,
+ .ptr = (void **) &resizable_shmem,
+ );
+}
+
+/*
+ * Initialize shared memory structure.
+ */
+static void
+resizable_shmem_shmem_init(void *arg)
+{
+ Assert(resizable_shmem != NULL);
+
+ resizable_shmem->num_entries = test_initial_entries;
+ memset(resizable_shmem->data, 0, mul_size(test_initial_entries, TEST_ENTRY_SIZE));
+}
+
+/*
+ * Attach to the already-allocated shared memory structure.
+ */
+static void
+resizable_shmem_shmem_attach(void *arg)
+{
+ Assert(resizable_shmem != NULL);
+}
+
+static ShmemCallbacks resizable_shmem_callbacks = {
+ .request_fn = resizable_shmem_request,
+ .init_fn = resizable_shmem_shmem_init,
+ .attach_fn = resizable_shmem_shmem_attach,
+};
+
+/*
+ * Resize the shared memory structure to accommodate the specified number of
+ * entries.
+ *
+ * Negative value for new_entries can be used to test resizing below the
+ * minimum size.
+ *
+ * Returns true if the resize was successful, false if ShmemResizeStruct()
+ * could not allocate the requested memory. On platforms that do not support
+ * resizable shared memory, ShmemResizeStruct() raises an error.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_resize);
+Datum
+resizable_shmem_resize(PG_FUNCTION_ARGS)
+{
+ int32 new_entries = PG_GETARG_INT32(0);
+ Size new_size;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (new_entries < 0)
+ new_size = 1;
+ else
+ new_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(new_entries, TEST_ENTRY_SIZE));
+ if (!ShmemResizeStruct("resizable_shmem", new_size))
+ PG_RETURN_BOOL(false);
+
+ ShmemProtectStruct("resizable_shmem");
+ resizable_shmem->num_entries = new_entries;
+
+ PG_RETURN_BOOL(true);
+}
+
+/*
+ * Write the given integer value to all entries in the data array.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_write);
+Datum
+resizable_shmem_write(PG_FUNCTION_ARGS)
+{
+ int32 entry_value = PG_GETARG_INT32(0);
+ int32 i;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (i = 0; i < resizable_shmem->num_entries; i++)
+ resizable_shmem->data[i] = entry_value;
+
+ PG_RETURN_VOID();
+}
+
+/*
+ * Check whether the first 'entry_count' entries all have the expected 'entry_value'.
+ * Returns true if all match, false otherwise.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_read);
+Datum
+resizable_shmem_read(PG_FUNCTION_ARGS)
+{
+ int32 entry_count = PG_GETARG_INT32(0);
+ int32 entry_value = PG_GETARG_INT32(1);
+ int32 i;
+
+ if (resizable_shmem == NULL)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (entry_count < 0 || entry_count > resizable_shmem->num_entries)
+ ereport(ERROR,
+ errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("entry_count %d is out of range (0..%d)", entry_count, resizable_shmem->num_entries));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (i = 0; i < entry_count; i++)
+ {
+ if (resizable_shmem->data[i] != entry_value)
+ PG_RETURN_BOOL(false);
+ }
+
+ PG_RETURN_BOOL(true);
+}
+
+/*
+ * Return the memory mapped against the main shared memory segment in this
+ * backend.
+ *
+ * The VMA containing our resizable_shmem pointer identifies the start of the
+ * main shared-memory segment.
+ *
+ * mprotect() calls issued when the resizable structure grows and shrinks can
+ * split the original mmap into several adjacent VMAs, so we sum the accounting
+ * fields across the base VMA and any VMAs contiguous with it.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_usage);
+Datum
+test_shmem_usage(PG_FUNCTION_ARGS)
+{
+ FILE *f;
+ char line[256];
+ uintptr_t target = (uintptr_t) resizable_shmem;
+ bool in_target_vma = false;
+ bool use_hugetlb = (huge_pages_status == HUGE_PAGES_ON);
+ unsigned long prev_end = 0;
+ int64 total_rss_kb = 0;
+ int64 total_swap_kb = 0;
+ int64 total_shared_hugetlb_kb = 0;
+ int64 val;
+ size_t result;
+
+ f = AllocateFile("/proc/self/smaps", "r");
+ if (f == NULL)
+ ereport(ERROR,
+ errcode_for_file_access(),
+ errmsg("could not open /proc/self/smaps: %m"));
+
+ while (fgets(line, sizeof(line), f) != NULL)
+ {
+ unsigned long start;
+ unsigned long end;
+
+ if (sscanf(line, "%lx-%lx", &start, &end) == 2)
+ {
+ if (in_target_vma)
+ {
+ /*
+ * Continue accumulating only across VMAs that are contiguous
+ * with the previous one; stop as soon as we hit a gap or a
+ * different mapping.
+ */
+ if (start != prev_end)
+ break;
+ }
+ else
+ in_target_vma = (target >= start && target < end);
+
+ prev_end = end;
+ }
+ else if (in_target_vma)
+ {
+ if (use_hugetlb)
+ {
+ if (sscanf(line, "Shared_Hugetlb: %ld kB", &val) == 1)
+ total_shared_hugetlb_kb += val;
+ }
+ else
+ {
+ if (sscanf(line, "Rss: %ld kB", &val) == 1)
+ total_rss_kb += val;
+ else if (sscanf(line, "Swap: %ld kB", &val) == 1)
+ total_swap_kb += val;
+ }
+ }
+ }
+
+ FreeFile(f);
+
+ if (use_hugetlb)
+ result = mul_size(total_shared_hugetlb_kb, 1024);
+ else
+ {
+ result = mul_size(total_rss_kb, 1024);
+ result = add_size(result, mul_size(total_swap_kb, 1024));
+ }
+
+ PG_RETURN_INT64(result);
+}
+
+/*
+ * Return the shared memory page size.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_pagesize);
+Datum
+test_shmem_pagesize(PG_FUNCTION_ARGS)
+{
+ PG_RETURN_INT32(pg_get_shmem_pagesize());
+}
+
+/*
+ * Walk the entries between the current size and the reserved maximum, accessing
+ * each one. Ideally, this function should (seg)fault the moment we try to access
+ * the entry outside the currently allocated size, but the memory allocation and
+ * protection mechanisms work on page basis. Hence it may only (seg)fault when a
+ * page boundary is crossed. The mode argument selects between "read" and
+ * "write" access.
+ *
+ * When the current end of the structure and end of maximal structure are on the
+ * same page, this function may not (seg)fault at all.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_access_beyond_size);
+Datum
+resizable_shmem_access_beyond_size(PG_FUNCTION_ARGS)
+{
+ text *mode_txt = PG_GETARG_TEXT_PP(0);
+ const char *mode = text_to_cstring(mode_txt);
+ bool do_write;
+ int32 sink = 0;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (strcmp(mode, "read") == 0)
+ do_write = false;
+ else if (strcmp(mode, "write") == 0)
+ do_write = true;
+ else
+ ereport(ERROR,
+ errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("mode must be \"read\" or \"write\""));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (int i = resizable_shmem->num_entries; i < test_max_entries; i++)
+ {
+ if (do_write)
+ resizable_shmem->data[i] = 0xdead;
+ else
+ sink = resizable_shmem->data[i];
+ }
+
+ /*
+ * Return the last read value so that compiler doesn't optimize away the
+ * assignment to sink.
+ */
+ PG_RETURN_INT32(sink);
+}
+
+
+/* ----------------------------------------------------------------
+ * Module initialization
+ * ----------------------------------------------------------------
+ */
+
+void
+_PG_init(void)
+{
+ int guc_context;
+
+ elog(LOG, "test_shmem module's _PG_init called");
+
+ RegisterShmemCallbacks(&TestShmemCallbacks);
+
+ /*
+ * Use PGC_POSTMASTER when loaded at startup so the values are fixed once
+ * the shared memory segment is created. When loaded after startup
+ * PGC_POSTMASTER is not allowed, so we use PGC_SIGHUP instead. Although
+ * we do not intend to change these values at config reload, PGC_SIGHUP is
+ * the least permissive context that allows defining the GUC after startup
+ * and still prevents it from being changed via SET.
+ */
+ if (process_shared_preload_libraries_in_progress)
+ guc_context = PGC_POSTMASTER;
+ else
+ {
+ guc_context = PGC_SIGHUP;
+ resizable_shmem_callbacks.flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP;
+ }
+
+ DefineCustomIntVariable("resizable_shmem.initial_entries",
+ "Initial number of entries in the test structure.",
+ NULL,
+ &test_initial_entries,
+ TEST_INITIAL_ENTRIES_DEFAULT,
+ 1,
+ INT_MAX,
+ guc_context,
+ 0,
+ NULL, NULL, NULL);
+
+ DefineCustomIntVariable("resizable_shmem.max_entries",
+ "Maximum number of entries in the test structure.",
+ NULL,
+ &test_max_entries,
+ TEST_MAX_ENTRIES_DEFAULT,
+ 1,
+ INT_MAX,
+ guc_context,
+ 0,
+ NULL, NULL, NULL);
+
+ /*
+ * When loaded after startup by a backend that is not creating the
+ * extension, the shared memory might have been resized to a size other
+ * than the initial size. Use SHMEM_ATTACH_UNKNOWN_SIZE to attach without
+ * knowing the exact size.
+ */
+ if (!process_shared_preload_libraries_in_progress && !creating_extension)
+ use_unknown_size = true;
+
+ RegisterShmemCallbacks(&resizable_shmem_callbacks);
+}
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 1a29d46213e..2e9da6cbf87 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1770,8 +1770,11 @@ pg_shadow| SELECT pg_authid.rolname AS usename,
pg_shmem_allocations| SELECT name,
off,
size,
- allocated_size
- FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size);
+ allocated_size,
+ minimum_size,
+ maximum_size,
+ reserved_space
+ FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size, minimum_size, maximum_size, reserved_space);
pg_shmem_allocations_numa| SELECT name,
numa_node,
size
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 298a3d586e7..0ff41ec129c 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -3192,6 +3192,7 @@ TestDSMRegistryHashEntry
TestDSMRegistryStruct
TestDecodingData
TestDecodingTxnData
+TestResizableShmemStruct
TestShmemData
TestSpec
TestValueType
--
2.34.1
[text/x-patch] v20260817-0007-Allow-to-resize-shared-buffers-without-res.patch (223.6K, ../../CAExHW5ts93Rnof7pjFFYrY9aTOmCo+xsdQazqeqN1BaKozTBvA@mail.gmail.com/6-v20260817-0007-Allow-to-resize-shared-buffers-without-res.patch)
download | inline diff:
From 06a020eebdea8a0beadc404c9da5a63e47ac28f4 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Sat, 6 Jun 2026 13:21:50 +0530
Subject: [PATCH v20260817 7/7] Allow to resize shared buffers without restart
User interface
==============
shared_buffers is now PGC_SIGHUP instead of PGC_POSTMASTER.
When a server is running, the new value of GUC (set using ALTER SYSTEM
... SET shared_buffers = ...; followed by SELECT pg_reload_conf()) does
not come into effect immediately. Instead a function
pg_resize_shared_buffers() is used to resize the buffer pool. The
function uses the current value of GUC in the backend where it is
executed. The function also coordinates the buffer access in other
backends while resizing the buffer pool.
SHOW shared_buffers now shows the current size of the shared buffer pool
but it also shows pending size of shared buffers, if any.
A new GUC max_shared_buffers is introduced to control the maximum value
of shared_buffers that can be set. By default it is 0. When explicitly
set, it needs to be higher than 'shared_buffers'. When
max_shared_buffers is set to 0, it assumes the same value as GUC
shared_buffers. This GUC determines the size of address space reserved
for future buffer pool sizes and the size of buffer look up table.
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking. If a buffer being evicted is pinned, we abort the
resizing operation. There are other alternative which are not
implemented in the current patches 1. to wait for the pinned buffer to
get unpinned, 2. the backend is killed or it itself cancels the query
or 3. rollback the operation which is implemented currently. Note that
option 1 and 2 would require the pinning related local and shared
records to be accessed. But we need infrastructure to do either of this
right now.
pg_resize_shared_buffers() throws an error when the server does not
support resizable shared memory structures.
Passing current buffer pool state to a new backend
==================================================
So far the buffer pool metdata (NBuffers and the shared memory segment
address space) is saved in process local heap memory since it's static
for the life of a server. It is passed to a new backend through
Postmaster. But with buffer pool being resized while the server running,
we need Postmaster to update its buffer pool metadata as the resizing
progresses and pass it to the new backend. This has few complications:
1. Postmaster does not receive ProcSignalBarrier. So we need to signal
it separately.
2. Postmaster's local state is inherited by the new backend when
fork()ed. But we need more complex implementation to pass it to an
exec()ed backend.
3. If Postmaster can not attend to its core functionality while it is
busy responding to the resizing signal
Instead, we maintain the buffer manager state in the shared memory.
Every backend maintains its own local copy of the state. When a backend
starts, it fetches syncs its local state with the global state before it
accesses any shared buffers. During run time it updates the local state
in response to the ProcSignalBarriers. The resizing does not advance
until the recent changes to the buffer pool state have been absorbed by
all the concurrent backends. This allows the backends to continue with
their regular activity without bothering about the buffer pool state
becoming inconsistent underneath.
Discussion points
=================
Removing the evicted buffers from buffer ring
---------------------------------------------
If the buffer pool has been shrunk, the buffers in the buffer ring may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition and masks any bug that
introduces an invalid buffer into the ring. The alternative fix is more
complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure hence inaccessible when processing
buffer resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work. So defering it
to v2.
Buffer lookup table
-------------------
In order to shrink the buffer lookup table, we need to compact the hash
table directory and the hash table entries so that the free space is
moved to the end of the memory allocated to the buffer lookup table.
This requires exclusively locking the shared hash table for a longer
duration, freezing the server for that duration. Furthermore the
compaction operation itself requires significant code. Hence we setup
the buffer lookup table considering the maximum possible size of the
buffer pool which is MaxAvailableMemory only once at the beginning. It
is not resized even though buffer pool is resized. We will need separate
effort later to implement a hash table which can be resized without
locking it for a longer duration.
BgWriter reset
--------------
The background writer makes sure to free buffers ahead of the clock
hand. For this it keeps track of the clock hand and jump forward if it
fells behind the clock hand. The position of clock hand is maintained as
a couple (size of buffer pool, position of next victim buffer). When
buffer pool size changes, the position of clock hand is adjusted
according to the new size. The background's knowledge of clock hand goes
out of sync when resize happens. Hence when resize happens the
background writer resets its knowledge of clock hand and jumps to the
clock hand's position. In case the background writer was ahead of clock
hand when resize happens, it will loose that advantage causing a
momentary glitch which may not be noticeable. It may be possible to
compare background writer's past knowledge of clock hand with the
current position by adjusting the previous according to the new size of
the buffer pool and avoid reset. But it requires more investigation
into background worker and clock hand synchronization. Hence deferred it
to a future version.
Fault tolerance of pg_resize_shared_buffers()
---------------------------------------------
If the backend executing pg_resize_shared_buffers() is interrupted
because of an ERROR, query cancellation, terminate signal, timeout etc.
the server is restarted to avoid leaving the buffer pool in an
inconsistent state. The server restarts with the buffer pool setup with
the new size. If the chances of such interruption are very low, this
solution might work for the first version. However, the restart can be
avoided by in following ways:
1. roll back the resize operation, if ERRORs are recoverable
2. In case of timeout and query cancellation, leave the resizing
operation with the buffer pool in degenerate but functioning state.
Re-attempting the resize would complete the operation or roll it
back.
3. In case of non-recoverable errors or terminate signal, it's better to
restart the operation.
The implementation depends upon the resizing protocol and use of
ProcSignalBarrier to maintain the process local buffer manager state
(shadow variables and address map protections). Hence require an
agreement on those things.
GUC type of shared_buffers
--------------------------
With the current code, on the platforms where have_resizable_shmem is
OFF, the users will be able to load new value of shared_buffers but not
resize the buffer pool. Probably we should change that set the type of
GUC to be POSTMASTER on those platforms. However, that still leaves the
server running one of the platforms which usually supports resizable
shared memory structure, but can not do so at run time because the
shared memory type does not support it, we can not do the same. Will
tackle this as the patches get finalized.
Calling BufferManagerInitProc()
-------------------------------
This function gets called twice in a backend startup sequence. Once so
that BufferManagerInitalizeAccess can access the buffer pool and second
time after ProcSignalInit() for the reasons mentioned there. I think we
need a better place so that we avoid calling it twice and set it only
once properly. Also it's not clear whether the function has been called
at all the right places.
I am actually not sure why don't we call ProcSignalInit() right after
InitProcess() in a backend startup sequence. If that happens, we don't
need to worry about it.
Updating local activeNBuffers when allocating new buffer
--------------------------------------------------------
Local activeNBuffers is updated in response to ProcSignalBarrier. It is
also updated by ClockSweepTick() when choosing a victim buffer for the
reasons mentioned there. The GetStrategyBuffer() function which
allocates a new buffer from the buffer ring doesn't update the local
activeNBuffers. I think it should be harmless to just rely on the
ProcSignalBarrier to update the local activeNBuffers and everyone uses
the same value. But it might interfere with the way Clock hand is
wrapped.
Barrier optimization
--------------------
Each resize operation uses three barriers to synchronize the buffer pool
state across all the backends. As noted at a few places in the code, we
may be able to use lesser number of barriers to reduce waiting time.
However, we need to be sure that we are not introducing any hazard by
doing so. Hence we will keep the current implementation for now and
optimize it after we have enough tests to cover all the scenarios.
Testing
-------
The stress tests give us a higher confidence that the buffer resizing
does not introduce hazards in the subsystems which use buffer pool
concurrently with the resizing. Once we have an agrement on protocol and
barrier mechanism, we will add more white box tests using injection
point to test specific hazardous scenarios.
Please note that the stress tests added by this or the earlier commits
are not necessarily meant to be committed to the core. But they are in
the patch so that reviewers have some readily available stress tests.
Author: Ashutosh Bapat
Initial patches by: Dmitry Dolgov <9erthalion6@gmail.com>,
Inspired by patches from: Haoyu Huang <haoyu.huang.68@gmail.com>
Author of some tests: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
Reviewed-by: Tomas Vondra
---
contrib/pg_buffercache/pg_buffercache_pages.c | 31 +-
contrib/pg_prewarm/autoprewarm.c | 27 +-
doc/src/sgml/config.sgml | 62 +-
doc/src/sgml/func/func-admin.sgml | 69 ++
doc/src/sgml/regress.sgml | 11 +
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/catalog/storage.c | 4 +
src/backend/postmaster/auxprocess.c | 7 +
src/backend/postmaster/postmaster.c | 5 +
src/backend/storage/buffer/Makefile | 3 +-
src/backend/storage/buffer/README | 62 +-
src/backend/storage/buffer/buf_init.c | 299 +++++-
src/backend/storage/buffer/buf_resize.c | 539 +++++++++++
src/backend/storage/buffer/buf_table.c | 27 +-
src/backend/storage/buffer/bufmgr.c | 221 ++++-
src/backend/storage/buffer/freelist.c | 150 ++-
src/backend/storage/buffer/meson.build | 1 +
src/backend/storage/ipc/procsignal.c | 10 +
src/backend/storage/lmgr/proc.c | 31 +
src/backend/tcop/postgres.c | 3 +
src/backend/utils/init/globals.c | 2 +
src/backend/utils/init/postinit.c | 77 +-
src/backend/utils/misc/guc.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 15 +-
src/include/catalog/pg_proc.dat | 15 +
src/include/miscadmin.h | 3 +
src/include/storage/buf_internals.h | 48 +-
src/include/storage/bufmgr.h | 15 +
src/include/storage/procsignal.h | 5 +
src/include/utils/guc.h | 2 +
src/include/utils/guc_hooks.h | 2 +
src/test/Makefile | 3 +-
src/test/README | 3 +
src/test/buffermgr/Makefile | 40 +
src/test/buffermgr/README | 43 +
src/test/buffermgr/buffermgr_test--1.0.sql | 52 +
src/test/buffermgr/buffermgr_test.conf | 11 +
src/test/buffermgr/buffermgr_test.control | 3 +
src/test/buffermgr/expected/buffer_resize.out | 290 ++++++
src/test/buffermgr/meson.build | 37 +
src/test/buffermgr/sql/buffer_resize.sql | 106 ++
.../buffermgr/t/001_resize_fault_tolerance.pl | 916 ++++++++++++++++++
.../t/002_client_join_buffer_resize.pl | 286 ++++++
src/test/buffermgr/t/003_resize_failures.pl | 202 ++++
.../buffermgr/t/004_resize_with_syslogger.pl | 68 ++
.../buffermgr/t/005_resize_unsupported.pl | 50 +
.../buffermgr/t/010_stress_resize_buffer.pl | 34 +
.../t/011_stress_drop_relation_buffers.pl | 99 ++
.../t/012_stress_drop_database_buffers.pl | 82 ++
src/test/buffermgr/t/013_stress_checkpoint.pl | 47 +
.../t/014_stress_flush_relation_buffers.pl | 68 ++
.../buffermgr/t/015_stress_pg_buffercache.pl | 60 ++
src/test/buffermgr/t/016_stress_pg_prewarm.pl | 70 ++
src/test/buffermgr/t/StressUtil.pm | 628 ++++++++++++
src/test/meson.build | 1 +
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 77 ++
src/tools/pgindent/typedefs.list | 1 +
57 files changed, 4877 insertions(+), 150 deletions(-)
create mode 100644 src/backend/storage/buffer/buf_resize.c
create mode 100644 src/test/buffermgr/Makefile
create mode 100644 src/test/buffermgr/README
create mode 100644 src/test/buffermgr/buffermgr_test--1.0.sql
create mode 100644 src/test/buffermgr/buffermgr_test.conf
create mode 100644 src/test/buffermgr/buffermgr_test.control
create mode 100644 src/test/buffermgr/expected/buffer_resize.out
create mode 100644 src/test/buffermgr/meson.build
create mode 100644 src/test/buffermgr/sql/buffer_resize.sql
create mode 100644 src/test/buffermgr/t/001_resize_fault_tolerance.pl
create mode 100644 src/test/buffermgr/t/002_client_join_buffer_resize.pl
create mode 100644 src/test/buffermgr/t/003_resize_failures.pl
create mode 100644 src/test/buffermgr/t/004_resize_with_syslogger.pl
create mode 100644 src/test/buffermgr/t/005_resize_unsupported.pl
create mode 100644 src/test/buffermgr/t/010_stress_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/011_stress_drop_relation_buffers.pl
create mode 100644 src/test/buffermgr/t/012_stress_drop_database_buffers.pl
create mode 100644 src/test/buffermgr/t/013_stress_checkpoint.pl
create mode 100644 src/test/buffermgr/t/014_stress_flush_relation_buffers.pl
create mode 100644 src/test/buffermgr/t/015_stress_pg_buffercache.pl
create mode 100644 src/test/buffermgr/t/016_stress_pg_prewarm.pl
create mode 100644 src/test/buffermgr/t/StressUtil.pm
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 9512f1efa2f..a1e72d79108 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -150,8 +150,6 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
Datum values[NUM_BUFFERCACHE_PAGES_ELEM];
bool nulls[NUM_BUFFERCACHE_PAGES_ELEM];
- CHECK_FOR_INTERRUPTS();
-
bufHdr = GetBufferDescriptor(i);
/* Lock each buffer header before inspecting. */
buf_state = LockBufHdr(bufHdr);
@@ -220,6 +218,12 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
}
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+
+ /*
+ * Check for interrupts here, at the end of the loop, so that the
+ * buffer index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
return (Datum) 0;
@@ -289,6 +293,13 @@ pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
HeapTuple tuple;
Datum result;
+ /*
+ * TODO: This allocates memory using NBuffers which may change while this
+ * function is executed. We need to change this function so that it
+ * doesn't rely on NBuffers being static throughout the execution of this
+ * function.
+ */
+
if (SRF_IS_FIRSTCALL())
{
int i,
@@ -450,8 +461,6 @@ pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
char *startptr_buff,
*endptr_buff;
- CHECK_FOR_INTERRUPTS();
-
bufHdr = GetBufferDescriptor(i);
/* Lock each buffer header before inspecting. */
@@ -481,6 +490,12 @@ pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
++idx;
++page_num;
}
+
+ /*
+ * Check for interrupts here, at the end of the loop, so that the
+ * buffer index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
Assert(idx <= max_entries);
@@ -592,8 +607,6 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
BufferDesc *bufHdr;
uint64 buf_state;
- CHECK_FOR_INTERRUPTS();
-
/*
* This function summarizes the state of all headers. Locking the
* buffer headers wouldn't provide an improved result as the state of
@@ -616,6 +629,12 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
buffers_pinned++;
+
+ /*
+ * Check for interrupts here, at the end of the loop, so that the
+ * buffer index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
memset(nulls, 0, sizeof(nulls));
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index 87959af3c54..7cdfd3b49a4 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -675,6 +675,7 @@ static int
apw_dump_now(bool is_bgworker, bool dump_unlogged)
{
int num_blocks;
+ int max_blocks;
int i;
int ret;
BlockInfoRecord *block_info_array;
@@ -704,19 +705,33 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
/*
* With sufficiently large shared_buffers, allocation will exceed 1GB, so
- * allow for a huge allocation to prevent outright failure.
+ * allow for a huge allocation to prevent outright failure. Use the
+ * current size of the buffer pool as the estimate of the number of blocks
+ * to dump, and grow the array if necessary.
*
* (In the future, it might be a good idea to redesign this to use a more
* memory-efficient data structure.)
*/
+ max_blocks = NBuffers;
block_info_array = (BlockInfoRecord *)
- palloc_extended((sizeof(BlockInfoRecord) * NBuffers), MCXT_ALLOC_HUGE);
+ palloc_extended((sizeof(BlockInfoRecord) * max_blocks), MCXT_ALLOC_HUGE);
for (num_blocks = 0, i = 0; i < NBuffers; i++)
{
uint64 buf_state;
- CHECK_FOR_INTERRUPTS();
+ /*
+ * Expand the array if necessary using the latest size of the buffer
+ * pool as the estimate of the number of blocks to dump.
+ */
+ if (num_blocks >= max_blocks)
+ {
+ max_blocks = NBuffers;
+ block_info_array = (BlockInfoRecord *)
+ repalloc_extended(block_info_array,
+ sizeof(BlockInfoRecord) * max_blocks,
+ MCXT_ALLOC_HUGE);
+ }
bufHdr = GetBufferDescriptor(i);
@@ -742,6 +757,12 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
}
UnlockBufHdr(bufHdr);
+
+ /*
+ * Check for interrupts here, at the end of the loop, so that the
+ * buffer index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
snprintf(transient_dump_file_path, MAXPGPATH, "%s.tmp", AUTOPREWARM_FILE);
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index f06e0f2e1a5..1bf06b4143f 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -1815,7 +1815,6 @@ include_dir 'conf.d'
that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
(Non-default values of <symbol>BLCKSZ</symbol> change the minimum
value.)
- This parameter can only be set at server start.
</para>
<para>
@@ -1838,6 +1837,67 @@ include_dir 'conf.d'
appropriate, so as to leave adequate space for the operating system.
</para>
+ <para>
+ The shared memory consumed by the buffer pool is allocated and
+ initialized according to the value of the GUC at the time of starting
+ the server. A desired new value of GUC can be loaded while the server is
+ running using <systemitem>SIGHUP</systemitem>. But the buffer pool will
+ not be resized immediately. Use
+ <function>pg_resize_shared_buffers()</function> to dynamically resize
+ the shared buffer pool (see <xref linkend="functions-admin"/> for
+ details). If the running server has
+ <varname>have_resizable_shmem</varname> set to OFF,
+ <function>pg_resize_shared_buffers()</function> throws an error since
+ resizing a shared memory structure is not supported on that server. In
+ such a case the new value of <varname>shared_buffers</varname> gets
+ loaded using <function>pg_reload_conf</function> but the new size of the
+ pool takes effect only after restarting the server. <command>SHOW
+ shared_buffers</command> shows the currently effective value and any
+ pending value of the GUC. Please note that when the GUC is changed, the
+ other GUCS which use this GUCs value to set their defaults will not be
+ changed. They may still require a server restart to consider new value.
+ </para>
+
+ <para>
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-max-shared-buffers" xreflabel="max_shared_buffers">
+ <term><varname>max_shared_buffers</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>max_shared_buffers</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Sets the upper limit for the <varname>shared_buffers</varname> value.
+ The default value is <literal>0</literal>,
+ which means no explicit limit is set and <varname>max_shared_buffers</varname>
+ will be automatically set to the value of <varname>shared_buffers</varname>
+ at server startup.
+ If this value is specified without units, it is taken as blocks,
+ that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
+ This parameter can only be set at server start.
+ </para>
+
+ <para>
+ This parameter determines the amount of memory address space to reserve
+ in each backend for expanding the buffer pool in future. While the
+ memory for buffer pool is allocated on demand as it is resized, the
+ memory required for the buffer lookup table and the array used to sort
+ buffers during a checkpoint is allocated at the server start
+ considering the largest buffer pool size allowed by this parameter.
+ <!-- TODO: Provide a numeric example of how much extra memory say max_shared_buffers = 1GB consume. -->
+ </para>
+
+ <para>
+ When <varname>have_resizable_shmem</varname> is OFF, this parameter does
+ not have any effect except limiting the maximum value of
+ <varname>shared_buffers</varname>. It does not decide the address space
+ to be reserved or the size of the buffer lookup table and the array
+ used to sort buffers during a checkpoint.
+ </para>
</listitem>
</varlistentry>
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 0eae1c1f616..2a621fff060 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -99,6 +99,75 @@
<returnvalue>off</returnvalue>
</para></entry>
</row>
+
+ <row>
+ <entry role="func_table_entry"><para role="func_signature">
+ <indexterm>
+ <primary>pg_resize_shared_buffers</primary>
+ </indexterm>
+ <function>pg_resize_shared_buffers</function> ()
+ <returnvalue>boolean</returnvalue>
+ </para>
+ <para>
+ Dynamically resizes the shared buffer pool to match the current value of
+ the <varname>shared_buffers</varname> parameter in the client backend
+ where it is run. This function implements a coordinated resize process
+ that ensures all backend processes continue to operate without causing
+ any hazard. The resize happens in multiple phases to maintain data
+ consistency and system stability. Returns <literal>true</literal> if the
+ resize was successful, otherwise <literal>false</literal>. In the latter
+ case, the buffer pool is left at its previous size; this can happen if a
+ buffer being evicted during shrink could not be released, or if the
+ operating system could not supply enough shared memory to expand the
+ pool. Consult the server log for the underlying reason. This function
+ can only be called by superusers.
+ </para>
+ <para>
+ To resize shared buffers, first update the <varname>shared_buffers</varname>
+ setting and reload the configuration, then verify the new value is loaded
+ before calling this function. For example:
+<programlisting>
+postgres=# ALTER SYSTEM SET shared_buffers = '256MB'; -- Step 1
+ALTER SYSTEM
+postgres=# SELECT pg_reload_conf(); -- Step 2
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers; -- Step 3
+ shared_buffers
+-------------------------
+ 128MB (pending: 256MB)
+(1 row)
+
+postgres=# SELECT pg_resize_shared_buffers(); -- Step 4
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers; -- Step 5 (verification)
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+</programlisting>
+ The <command>SHOW shared_buffers</command> at Step 3 is important to
+ verify that the configuration reload was successful and the new value is
+ available to the current session before attempting the resize. The
+ output shows both the current and pending values when the GUC change is
+ pending to be applied.
+ </para>
+ <para>
+ <!-- TODO: Document behaviour when the function is called on platforms that do not support resizable shared memory -->
+ linkend="functions-admin-signal-table"/> send control signals to
+ other server processes. Use of these functions is restricted to
+ superusers by default but access may be granted to others using
+ <command>GRANT</command>, with noted exceptions.
+ </para>
+ </entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/doc/src/sgml/regress.sgml b/doc/src/sgml/regress.sgml
index c74941bfbf2..30092fd820d 100644
--- a/doc/src/sgml/regress.sgml
+++ b/doc/src/sgml/regress.sgml
@@ -275,6 +275,17 @@ make check-world PG_TEST_EXTRA='kerberos ldap ssl load_balance libpq_encryption'
</programlisting>
The following values are currently supported:
<variablelist>
+ <varlistentry>
+ <term><literal>bufmgr_stress</literal></term>
+ <listitem>
+ <para>
+ Runs the shared_buffers resize stress tests under
+ <filename>src/test/buffermgr</filename>. Not enabled by default because
+ they are long-running and resource-intensive.
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry>
<term><literal>checksum</literal>, <literal>checksum_extended</literal></term>
<listitem>
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index a678f345230..fa9a01a7ec3 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -376,6 +376,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ InitializeMaxNBuffers();
+
ShmemCallRequestCallbacks();
CreateSharedMemoryAndSemaphores();
diff --git a/src/backend/catalog/storage.c b/src/backend/catalog/storage.c
index e443a4993c5..5c36820d58f 100644
--- a/src/backend/catalog/storage.c
+++ b/src/backend/catalog/storage.c
@@ -33,6 +33,7 @@
#include "storage/proc.h"
#include "storage/smgr.h"
#include "utils/hsearch.h"
+#include "utils/injection_point.h"
#include "utils/memutils.h"
#include "utils/rel.h"
@@ -383,6 +384,9 @@ RelationTruncate(Relation rel, BlockNumber nblocks)
*
* (See also visibilitymap.c if changing this code.)
*/
+
+ /* Load the injection point before entering the critical section */
+ INJECTION_POINT_LOAD("drop-relation-buffers-scan");
START_CRIT_SECTION();
if (RelationNeedsWAL(rel))
diff --git a/src/backend/postmaster/auxprocess.c b/src/backend/postmaster/auxprocess.c
index 07a3b5c5923..a09d376d333 100644
--- a/src/backend/postmaster/auxprocess.c
+++ b/src/backend/postmaster/auxprocess.c
@@ -19,6 +19,7 @@
#include "miscadmin.h"
#include "pgstat.h"
#include "postmaster/auxprocess.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/proc.h"
@@ -106,6 +107,12 @@ AuxiliaryProcessMainCommon(void)
*/
ShmemReprotectResizableStructs();
+ /*
+ * Update buffer manager's local state, which might have been changed by
+ * an online resize after startup.
+ */
+ BufferManagerInitProc();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 1711743f1ac..087556fe6b5 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -959,6 +959,11 @@ PostmasterMain(int argc, char *argv[])
*/
InitializeFastPathLocks();
+ /*
+ * Calculate MaxNBuffers after NBuffersGUC has been set.
+ */
+ InitializeMaxNBuffers();
+
/*
* Also call any legacy shmem request hooks that might've been installed
* by preloaded libraries.
diff --git a/src/backend/storage/buffer/Makefile b/src/backend/storage/buffer/Makefile
index fd7c40dcb08..3bc9aee85de 100644
--- a/src/backend/storage/buffer/Makefile
+++ b/src/backend/storage/buffer/Makefile
@@ -17,6 +17,7 @@ OBJS = \
buf_table.o \
bufmgr.o \
freelist.o \
- localbuf.o
+ localbuf.o \
+ buf_resize.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index b332e002ba1..9ee05d562eb 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -181,8 +181,11 @@ buffer header spinlock, which would have to be taken anyway to increment the
buffer reference count, so it's nearly free.)
The "clock hand" is a buffer index, nextVictimBuffer, that moves circularly
-through all the available buffers. nextVictimBuffer is protected by the
-buffer_strategy_lock.
+through all the available buffers. Usually a victim can be chosen from the whole
+buffer pool, except when resizing the buffer pool. During resizing the victim
+can be chosen from a range of buffer pool which will not be affected by the
+resizing. See "Resizing shared buffers section" below. nextVictimBuffer is
+protected by the buffer_strategy_lock.
The algorithm for a process that needs to obtain a victim buffer is:
@@ -275,3 +278,58 @@ As of 8.4, background writer starts during recovery mode when there is
some form of potentially extended recovery to perform. It performs an
identical service to normal processing, except that checkpoints it
writes are technically restartpoints.
+
+Resizing shared buffers at runtime
+----------------------------------
+
+Before PostgreSQL 20, the size of the shared buffer pool (i.e. the number of
+shared buffers) was given by the global variable NBuffers and was fixed at
+server start time using GUC 'shared_buffers'. In order to change the size of the
+shared buffer pool, one needed to change the GUC and restart the server.
+Starting PostgreSQL 20, PostgreSQL supports resizing the buffer pool without a
+server restart. The value of the GUC 'shared_buffers' is stored in a new GUC
+variable called NBuffersGUC. The old variable NBuffers now strictly reflects the
+number of buffers in the buffer pool. The new GUC variable max_shared_buffers
+defines the maximum size of the shared buffer pool. See configure.sgml for more
+details about these GUCs.
+
+Resizing buffer pool involves resizing the data structures used by the buffer
+manager and coordinating the resize with all the backends.
+
+The buffer manager maintains following data structures in shared memory.
+1. Buffer Descriptors: An array of BufferDesc structures, one per buffer.
+2. Buffer Blocks: An array of buffer blocks, forming the buffer pool.
+3. Buffer Lookup Table: A hash table mapping a page to the buffer containing
+ that page.
+4. IO conditional variables: An array of conditional variables, one per buffer.
+5. Checkpoint buffer ids: An array of buffer ids used during checkpointing.
+6. Buffer Control: A structure containing information about the size of the
+ buffer pool and whether it is being resized.
+
+Except for the Buffer Lookup Table, Checkpoint buffer ids and Buffer Control
+all other structures are registered as resizable structures. We use
+ShmemResizeStruct() described in doc/src/sgml/xfunc.sgml to resize them. The
+hash table contains multiple shared structures, each of which needs to be
+resized separately. Since these substructures are not registered as separate
+structures, they can not be turned into resizable structures and hence the hash
+table is not registered as a resizable structure. Instead we set it up for
+maximal buffer pool at the server startup. The Checkpoint buffer ids array is
+filled in by the checkpointer with buffers to be written out during a
+checkpoint; if it were resized alongside the buffer pool, a shrink concurrent
+with an ongoing checkpoint could discard entries the checkpointer still needs
+to process. To avoid that, it is also sized for the maximal buffer pool at
+server startup.
+
+For ease of resizing, we differentiate between the number of buffers in the
+whole buffer pool (BufferControl::currentNBuffers) and the number of buffers at
+the start of the buffer pool from which a victim can be chosen for replacement
+(BufferControl::activeNBuffers). Each backend maintains a local copy of these
+two numbers. These copies provide a local view of the buffer pool that the
+backend can rely upon without worrying about the concurrent changes happening to
+the buffer pool. These copies are updated through a barrier mechanism when a
+resize is performed.
+
+To resize shared buffers at runtime a user performs the steps mentioned in the
+description of shared_buffers GUC variable in configure.sgml. Actual resizing
+protocol is documented in the prologue of function pg_resize_shared_buffers() in
+src/backend/storage/buffer/buf_resize.c
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 9ddf6551fcd..91671bce888 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,10 +17,15 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proclist.h"
#include "storage/shmem.h"
#include "storage/subsystems.h"
+#include "utils/guc.h"
+#include "utils/guc_hooks.h"
+#include "utils/injection_point.h"
+BufferControlBlock *BufferControl;
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
ConditionVariableMinimallyPadded *BufferIOCVArray;
@@ -37,6 +42,31 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
.attach_fn = BufferManagerShmemAttach,
};
+/*
+ * Resizable shared memory structures backing the buffer pool.
+ */
+static const struct
+{
+ const char *name;
+ size_t element_size;
+ size_t alignment;
+ void **ptr;
+} BufferManagerResizableStructs[] = {
+
+ {"Buffer Descriptors",
+ sizeof(BufferDescPadded),
+ PG_CACHE_LINE_SIZE,
+ (void **) &BufferDescriptors},
+ {"Buffer Blocks",
+ BLCKSZ,
+ PG_IO_ALIGN_SIZE,
+ (void **) &BufferBlocks},
+ {"Buffer IO Condition Variables",
+ sizeof(ConditionVariableMinimallyPadded),
+ PG_CACHE_LINE_SIZE,
+ (void **) &BufferIOCVArray},
+};
+
/*
* Data Structures:
* buffers live in a freelist and a lookup data structure.
@@ -69,6 +99,27 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
* multiple times. Check the PrivateRefCount infrastructure in bufmgr.c.
*/
+/*
+ * Initialize a single buffer.
+ */
+static void
+InitializeBuffer(int buf_id)
+{
+ /*
+ * Do not use GetBufferDescriptor here since it relies on the buffer
+ * descriptor being initialized.
+ */
+ BufferDesc *buf = &(BufferDescriptors[buf_id]).bufferdesc;
+
+ ClearBufferTag(&buf->tag);
+ pg_atomic_init_u64(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ buf->buf_id = buf_id;
+ pgaio_wref_clear(&buf->io_wref);
+ proclist_init(&buf->lock_waiters);
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+}
+
/*
* Register shared memory area for the buffer pool.
@@ -76,25 +127,21 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
static void
BufferManagerShmemRequest(void *arg)
{
- ShmemRequestStruct(.name = "Buffer Descriptors",
- .size = NBuffersGUC * sizeof(BufferDescPadded),
- /* Align descriptors to a cacheline boundary. */
- .alignment = PG_CACHE_LINE_SIZE,
- .ptr = (void **) &BufferDescriptors,
- );
-
- ShmemRequestStruct(.name = "Buffer Blocks",
- .size = NBuffersGUC * (Size) BLCKSZ,
- /* Align buffer pool on IO page size boundary. */
- .alignment = PG_IO_ALIGN_SIZE,
- .ptr = (void **) &BufferBlocks,
- );
+ /*
+ * Fall back to fixed sized shared buffer pool if resizable shared memory
+ * is not supported on this platform.
+ */
+#ifdef HAVE_RESIZABLE_SHMEM
+ bool resizable = (shared_memory_type == SHMEM_TYPE_MMAP);
+#else
+ bool resizable = false;
+#endif
+ int min_nbuffers = resizable ? MIN_NUM_BUFFERS : 0;
+ int max_nbuffers = resizable ? MaxNBuffers : 0;
- ShmemRequestStruct(.name = "Buffer IO Condition Variables",
- .size = NBuffersGUC * sizeof(ConditionVariableMinimallyPadded),
- /* Align descriptors to a cacheline boundary. */
- .alignment = PG_CACHE_LINE_SIZE,
- .ptr = (void **) &BufferIOCVArray,
+ ShmemRequestStruct(.name = "Buffer Control",
+ .size = sizeof(BufferControl),
+ .ptr = (void **) &BufferControl,
);
/*
@@ -103,11 +150,29 @@ BufferManagerShmemRequest(void *arg)
* memory at runtime. As that'd be in the middle of a checkpoint, or when
* the checkpointer is restarted, memory allocation failures would be
* painful.
+ *
+ * When the buffer pool is resizable, it is sized for MaxNBuffers up front
+ * so that the entries filled in by the checkpointer are not freed even
+ * when the buffer pool is shrunk.
*/
ShmemRequestStruct(.name = "Checkpoint BufferIds",
- .size = NBuffersGUC * sizeof(CkptSortItem),
+ .size = (size_t) (resizable ? MaxNBuffers : NBuffersGUC) * sizeof(CkptSortItem),
+ .alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &CkptBufferIds,
);
+
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
+ {
+ size_t elem_size = BufferManagerResizableStructs[i].element_size;
+
+ ShmemRequestStruct(.name = BufferManagerResizableStructs[i].name,
+ .minimum_size = min_nbuffers * elem_size,
+ .size = NBuffersGUC * elem_size,
+ .maximum_size = max_nbuffers * elem_size,
+ .alignment = BufferManagerResizableStructs[i].alignment,
+ .ptr = BufferManagerResizableStructs[i].ptr,
+ );
+ }
}
/*
@@ -129,34 +194,194 @@ BufferManagerShmemInit(void *arg)
* Initialize all the buffer headers.
*/
for (int i = 0; i < NBuffers; i++)
+ InitializeBuffer(i);
+
+ /* Initialize BufferControl */
+ pg_atomic_init_u32(&BufferControl->currentNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->activeNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->targetNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->resizer_pid, 0);
+
+ /* Need to perform per backend steps in this backend too. */
+ BufferManagerShmemAttach(arg);
+}
+
+static void
+BufferManagerShmemAttach(void *arg)
+{
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+
+ BufferManagerInitProc();
+}
+
+/*
+ * Fetch latest buffer pool sizes (NBuffers and activeNBuffers) shared state
+ * into process local globals.
+ */
+void
+BufferManagerInitProc(void)
+{
+ NBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ elog(DEBUG1, "setting process local buffer pool sizes: currentNBuffers = %d, activeNBuffers = %d", NBuffers, activeNBuffers);
+}
+
+/*
+ * Protect unused shared memory reserved address space.
+ *
+ * Protect the parts of the shared memory address space reserved by the buffer
+ * manager which are not used by current structures from being accessed by
+ * backends.
+ *
+ * Unused address spaces of all resizable shared structures, including the
+ * buffer manager ones, are protected at the server startup together using
+ * ShmemProtectResizableStructs(). We need this function only during resizing of
+ * the buffer pool when we specifically adjust protections of buffer manager
+ * structures.
+ */
+void
+BufferManagerShmemProtect(void)
+{
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
{
- BufferDesc *buf = GetBufferDescriptor(i);
+ if (i == 2)
+ INJECTION_POINT("buffer-mgr-protect-struct", NULL);
+ ShmemProtectStruct(BufferManagerResizableStructs[i].name);
+ }
+}
+
+/*
+ * Resize and reinitialize shared buffer manager structures when resizing the
+ * buffer pool.
+ *
+ * Returns true if all the structures were resized successfully, false
+ * otherwise. We expect shrink to always succeed, but expansion may fail if the
+ * system is out of memory.
+ *
+ * The caller will always see all the structures resized consistently. If
+ * expanding a structure fails, all the expanded structures are shrunk back to
+ * their original sizes and no barrier is sent to the other backends.
+ */
+bool
+BufferManagerShmemResize(int currentNBuffers, int targetNBuffers)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizing shared buffer pool is not supported on this platform"));
+ pg_unreachable();
+#else
+ int resized = 0;
- ClearBufferTag(&buf->tag);
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
- pg_atomic_init_u64(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
+ {
+ const char *name = BufferManagerResizableStructs[i].name;
+ size_t elem_size = BufferManagerResizableStructs[i].element_size;
+ bool resize_ok = true;
+
+#ifdef USE_INJECTION_POINTS
+ if (i == 2)
+ {
+ /* Injection point to simulate an interruption in this function. */
+ INJECTION_POINT("buffer-mgr-resize-struct", NULL);
- buf->buf_id = i;
+ /*
+ * Injection point to simulate a failure in resizing a structure
+ * like memory allocation failure without actually running out of
+ * memory.
+ */
+ if (IS_INJECTION_POINT_ATTACHED("buffer-mgr-resize-struct-fail"))
+ resize_ok = false;
+ }
+#endif
- pgaio_wref_clear(&buf->io_wref);
+ if (resize_ok)
+ resize_ok = ShmemResizeStruct(name, (size_t) targetNBuffers * elem_size);
- proclist_init(&buf->lock_waiters);
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ if (!resize_ok)
+ {
+ Assert(targetNBuffers > currentNBuffers);
+ for (int j = 0; j < resized; j++)
+ ShmemResizeStruct(BufferManagerResizableStructs[j].name,
+ (size_t) currentNBuffers * BufferManagerResizableStructs[j].element_size);
+ return false;
+ }
+ resized++;
}
- /* Initialize per-backend file flush context */
- WritebackContextInit(&BackendWritebackContext,
- &backend_flush_after);
+ /* Initialize the headers for new buffers. */
+ for (int i = currentNBuffers; i < targetNBuffers; i++)
+ InitializeBuffer(i);
+
+ return true;
+#endif
}
-static void
-BufferManagerShmemAttach(void *arg)
+/*
+ * check_shared_buffers
+ * GUC check_hook for shared_buffers
+ *
+ * When reloading the configuration, shared_buffers should not be set to a value
+ * higher than max_shared_buffers fixed at the boot time.
+ */
+bool
+check_shared_buffers(int *newval, void **extra, GucSource source)
{
- /* Update the size of the buffer pool. */
- NBuffers = NBuffersGUC;
+ if (finalMaxNBuffers && *newval > MaxNBuffers)
+ {
+ GUC_check_errdetail("\"shared_buffers\" must be less than \"max_shared_buffers\".");
+ return false;
+ }
+ return true;
+}
- /* Initialize per-backend file flush context */
- WritebackContextInit(&BackendWritebackContext,
- &backend_flush_after);
+/*
+ * show_shared_buffers
+ * GUC show_hook for shared_buffers
+ *
+ * Shows both current and pending buffer counts with proper unit formatting.
+ */
+const char *
+show_shared_buffers(bool use_units)
+{
+ static char buffer[128];
+ int64 current_value;
+ const char *current_unit;
+ int currentNBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+
+ if (use_units)
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ else
+ {
+ current_unit = "";
+ current_value = currentNBuffers;
+ }
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s", current_value, current_unit);
+
+ if (currentNBuffers != NBuffersGUC)
+ {
+ int64 pending_value;
+ const char *pending_unit;
+
+ /*
+ * Shared buffer pool is pending to be resized, show both current and
+ * pending sizes.
+ */
+ if (use_units)
+ convert_int_from_base_unit(NBuffersGUC, GUC_UNIT_BLOCKS, &pending_value, &pending_unit);
+ else
+ {
+ pending_value = NBuffersGUC;
+ pending_unit = "";
+ }
+ snprintf(buffer + strlen(buffer), sizeof(buffer) - strlen(buffer), " (pending: " INT64_FORMAT "%s)",
+ pending_value, pending_unit);
+ }
+
+ return buffer;
}
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
new file mode 100644
index 00000000000..90f5fb1c71d
--- /dev/null
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -0,0 +1,539 @@
+/*-------------------------------------------------------------------------
+ *
+ * buf_resize.c
+ * shared buffer pool resizing functionality
+ *
+ * This module contains the implementation of shared buffer pool resizing,
+ * including the main resize coordination function and barrier processing
+ * functions that synchronize all backends during resize operations.
+ *
+ * Portions Copyright (c) 1996-2026, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/storage/buffer/buf_resize.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "access/htup_details.h"
+#include "fmgr.h"
+#include "funcapi.h"
+#include "miscadmin.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
+#include "utils/builtins.h"
+#include "utils/injection_point.h"
+
+static volatile sig_atomic_t safe_exit = true;
+
+#ifdef HAVE_RESIZABLE_SHMEM
+static bool resize_shared_buffers_internal(void);
+static void buf_resize_shmem_exit(int code, Datum arg);
+
+/*
+ * Set the new buffer allocation pool size, broadcast it to all the backends
+ * and wait for them to acknowledge it.
+ */
+static void
+buf_resize_set_new_alloc_size(int alloc_size)
+{
+ uint64 generation;
+
+ pg_atomic_write_u32(&BufferControl->activeNBuffers, alloc_size);
+ StrategyAdjustNewBufAllocSize();
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC);
+ INJECTION_POINT("pgrsb-new-buffer-alloc-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC barrier");
+}
+
+/*
+ * Update the buffer pool size, broadcast it to all the backends and wait for
+ * them to acknowledge the change.
+ */
+static void
+buf_resize_set_buffer_pool_size(int new_size)
+{
+ uint64 generation;
+
+ pg_atomic_write_u32(&BufferControl->currentNBuffers, new_size);
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE);
+ INJECTION_POINT("pgrsb-buffer-pool-size-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE barrier");
+}
+
+/*
+ * Resize the shared buffer manager structures, broadcast the change to all
+ * the backends and wait for them to acknowledge it.
+ *
+ * If memory is not available when expanding the buffer pool, this function
+ * returns false without sending the barrier. When shrinking the buffer pool, we
+ * don't expect any failure, so this function always returns true.
+ */
+static bool
+buf_resize_shmem_resize(int currentNBuffers, int targetNBuffers)
+{
+ uint64 generation;
+
+ if (!BufferManagerShmemResize(currentNBuffers, targetNBuffers))
+ {
+ Assert(targetNBuffers > currentNBuffers);
+ return false;
+ }
+
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE);
+ INJECTION_POINT("pgrsb-buffer-pool-resize-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE barrier");
+ return true;
+}
+#endif
+
+/*
+ * C implementation of SQL interface to update the shared buffers according to
+ * the current values of shared_buffers GUC.
+ *
+ * Atomic BufferControl::resizer_pid holds the PID of the backend currently
+ * performing a resize, or 0 when no resize is in progress. Using
+ * compare-and-exchange to set and reset this field, we make sure that only one
+ * resize is in progress at a time.
+ *
+ * Shrinking the buffer pool involves the following steps:
+ * - s1: Set BufferControl::activeNBuffers to the new size of the buffer pool
+ * and send SHBUF_NEW_BUFFER_ALLOC barrier to all backends. Every backend is
+ * expected to update their local buffer allocation pool size and acknowledge
+ * the barrier.
+ * - s2: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, new buffer allocations will be restricted
+ * to the new size of the buffer pool.
+ * - s3: Evict the buffers beyond the new size. A backend which still requires
+ * a previously allocated buffer which is being evicted, must have pinned it.
+ * If a pinned buffer is encountered, the resize operation is rolled back and
+ * the function returns false.
+ * - s4: If eviction succeeds, no backend should be using the buffers beyond
+ * the new size of the buffer pool and no new buffers can be allocated in
+ * that range. Update BufferControl::currentNBuffers to the new size of the
+ * buffer pool and send SHBUF_BUFFER_POOL_SIZE barrier to all backends. In
+ * response, all the backends should update their local buffer pool size and
+ * acknowledge the barrier.
+ * - s5: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, no backend will be accessing the shared
+ * buffer manager structures beyond the new size of the buffer pool.
+ * - s6: Resize the shared buffer manager structures to the new size (using
+ * ShmemResizeStruct()) and send SHBUF_BUFFER_POOL_RESIZE barrier to all
+ * backends. In response, all the backends should call ShmemProtectStruct()
+ * to update the memory address space protection of the shared buffer
+ * manager structures.
+ * - s7: Wait for all backends to acknowledge the barrier, before returning
+ * true to indicate successful resizing.
+ *
+ * Expanding the buffer pool involves the following steps:
+ * - e1: Resize the shared buffer manager structures to the new size (using
+ * ShmemResizeStruct()) and send SHBUF_BUFFER_POOL_RESIZE barrier to all the
+ * backends. In response, all the backends should call ShmemProtectStruct()
+ * to update the memory address space protection of the shared buffer
+ * manager structures. If expanding the shared buffer manager structures
+ * fails because of lack of memory, the function returns false without
+ * sending the barrier.
+ * - e2: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, every backend will be able to access the
+ * shared buffer manager structures beyond the old size of the buffer pool.
+ * - e3: Update BufferControl::currentNBuffers to the new size of the buffer
+ * pool and send SHBUF_BUFFER_POOL_SIZE barrier to all backends. In response,
+ * all the backends should update their local buffer pool size and
+ * acknowledge the barrier.
+ * - e4: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, every backend is setup to use the new size
+ * of the buffer pool.
+ * - e5: Update BufferControl::activeNBuffers to the new size of the buffer
+ * pool, so that backends can start allocating from the new area of the
+ * buffer pool. Send SHBUF_NEW_BUFFER_ALLOC barrier to all backends. In
+ * response, all the backends should update their local buffer allocation
+ * pool size and acknowledge the barrier.
+ * - e6: Wait for all backends to acknowledge the barrier, before returning
+ * true to indicate successful resizing.
+ *
+ * Reason we introduce e4:
+ * Once we expand the new buffer allocation area, all the backends will start
+ * allocating buffers from the new area. Since this happens asynchronously,
+ * there is a chance that some backends may see buffers from outside their
+ * known buffer pool size. To avoid that, first set the new buffer pool size,
+ * broadcast it to all the backends and wait for them to update their
+ * knowledge of buffer pool size. We may be able to avoid sending the barrier
+ * after the first step or avoid them altogether but it's not clear that that
+ * is completely hazard free. It feels safer this way, even though it takes
+ * longer.
+ *
+ * If a timeout happens or request to cancel query arrives while the function is
+ * being executed, we need to abort the operation immediately. The state of the
+ * buffer pool and its state as viewed by the backends may not be consistent at
+ * that point. Hence we escalate it to PANIC to restart the server and avoid
+ * inconsistent state. We may improve this situation by leaving the buffer pool
+ * in a consistent but degenerate state and allowing a subsequent resize
+ * operation to rollback or continue the operation.
+ *
+ * If an ERROR is raised while the function is being executed, we may have
+ * already entered an inconsistent state. Hence we escalate it to PANIC to
+ * restart the server and avoid inconsistent state.
+ */
+Datum
+pg_resize_shared_buffers(PG_FUNCTION_ARGS)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizing shared buffer pool is not supported on this platform"));
+ pg_unreachable();
+#else
+ bool success = false;
+
+ if (shared_memory_type != SHMEM_TYPE_MMAP)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizing shared buffer pool is not supported on this platform"));
+
+ /*
+ * Register the exit hook before claiming resizer_pid, so that if we exit
+ * after claiming resizer_pid, the hook is in place to reset it.
+ */
+ before_shmem_exit(buf_resize_shmem_exit, 0);
+
+ PG_TRY();
+ {
+ uint32 expected_pid = 0;
+
+ if (!pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, MyProcPid))
+ {
+ /*
+ * Another backend holds resizer_pid; expected_pid was updated by
+ * the CAS to reflect its PID.
+ */
+ elog(LOG, "shared buffer resize already in progress in backend %u",
+ expected_pid);
+ /* No shared memory was touched, so it should be safe to exit. */
+ Assert(safe_exit);
+ }
+ else
+ {
+ INJECTION_POINT("pg-resize-shared-buffers-flag-set", NULL);
+
+ /*
+ * We are about to make changes to the shared memory which can not
+ * be rolled back easily since we need all the backends to
+ * acknowledge these changes. Indicate that a sudden exit in this
+ * state can leave the server in an inconsistent state.
+ */
+ safe_exit = false;
+ success = resize_shared_buffers_internal();
+
+ /*
+ * The changes to shared memory are in a consistent state across
+ * all the backends, so it should be safe to exit.
+ */
+ safe_exit = true;
+ }
+ }
+ PG_FINALLY();
+ {
+ uint32 expected_pid = MyProcPid;
+
+ /*
+ * We are in the middle of resizing and caught an error. Without
+ * knowing the reason and exact state of resizing it's not safe to
+ * continue or to exit. Restarting the server is the safest option
+ * here. Emit the error to the server log and raise PANIC to restart
+ * the server.
+ */
+ if (!safe_exit)
+ {
+ HOLD_INTERRUPTS();
+ errcontext("during shared buffer resize");
+ EmitErrorReport();
+ ereport(PANIC,
+ errmsg("shared buffer resize caught an error when shared memory was in an inconsistent state"));
+ pg_unreachable();
+ }
+
+ /*
+ * Reset the PID, if we set it before removing the shmem_exit hook so
+ * as not to leave it set after the backend has exited.
+ */
+ (void) pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, 0);
+ cancel_before_shmem_exit(buf_resize_shmem_exit, 0);
+ }
+ PG_END_TRY();
+
+ if (success)
+ elog(LOG, "shared buffer resizing to %d buffers completed successfully", NBuffersGUC);
+ else
+ elog(WARNING, "shared buffer resizing to %d buffers failed", NBuffersGUC);
+
+ PG_RETURN_BOOL(success);
+#endif
+}
+
+#ifdef HAVE_RESIZABLE_SHMEM
+/*
+ * Workhorse function for the C implementation.
+ */
+static bool
+resize_shared_buffers_internal(void)
+{
+ int currentNBuffers;
+ int targetNBuffers;
+ bool resize_success;
+
+ currentNBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+ targetNBuffers = NBuffersGUC;
+ if (currentNBuffers == targetNBuffers)
+ {
+ elog(LOG, "shared buffers are already at %d, no need to resize", currentNBuffers);
+ return true;
+ }
+
+ /*
+ * TODO: What if the NBuffersGUC value seen here is not the desired one
+ * because somebody did a pg_reload_conf() between the last
+ * pg_reload_conf() and execution of this function?
+ */
+
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, targetNBuffers);
+ elog(LOG, "resizing shared buffers from %d to %d", currentNBuffers, targetNBuffers);
+
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * step s1, s2: Restrict new buffer allocations to the new buffer pool
+ * size.
+ *
+ * TODO: Alternate design idea by Andres (as I understand it): Set
+ * BufferControl::activeNBuffers and send the barrier. Instead of
+ * waiting for barrier, start evicting buffers but don't unpin the
+ * evicted buffers so that they will not be considered for new
+ * allocations. Once all the buffers are evicted wait for the barrier
+ * to be acknowledged. This will reduce the time taken to shrink the
+ * buffer pool.
+ */
+ elog(LOG, "shrinking buffer pool, restricting allocations to %d buffers", targetNBuffers);
+ buf_resize_set_new_alloc_size(targetNBuffers);
+
+ /* Step s3: Evict buffers in the area being shrunk */
+ elog(LOG, "evicting buffers %u..%u", targetNBuffers + 1, currentNBuffers);
+ if (!EvictExtraBuffers(targetNBuffers, currentNBuffers))
+ {
+ elog(WARNING, "failed to evict extra buffers during shrinking");
+
+ /* Eviction failed, rollback the buffer resize operation. */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, currentNBuffers);
+ buf_resize_set_new_alloc_size(currentNBuffers);
+ return false;
+ }
+
+ /* Step s4, s5: Update the buffer pool size. */
+ buf_resize_set_buffer_pool_size(targetNBuffers);
+ }
+
+ /* Step s6, s7 or e1, e2: Resize the buffer manager structures. */
+ resize_success = buf_resize_shmem_resize(currentNBuffers, targetNBuffers);
+
+ if (targetNBuffers > currentNBuffers)
+ {
+ if (!resize_success)
+ {
+ elog(WARNING, "failed to expand buffer pool structures");
+
+ /* Revert any changes to the shared memory in this function. */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, currentNBuffers);
+ return false;
+ }
+
+ /* Step e3, e4: Declare new buffer pool size. */
+ buf_resize_set_buffer_pool_size(targetNBuffers);
+
+ /* Step e5, e6: Let expanded buffer pool be used by all backends. */
+ buf_resize_set_new_alloc_size(targetNBuffers);
+ }
+
+ return true;
+}
+
+/*
+ * Function to handle process exit when buffer resizing is in progress.
+ */
+static void
+buf_resize_shmem_exit(int code, Datum arg)
+{
+ uint32 expected_pid;
+
+ /*
+ * If resizer_pid does not match our PID, either we never claimed it or we
+ * have already released it. Nothing to do.
+ */
+ if (pg_atomic_read_u32(&BufferControl->resizer_pid) != MyProcPid)
+ return;
+
+ /*
+ * Resize is in progress and the process crashed. We do not know exactly
+ * at which step of the resizing we are. Just restart the server to be
+ * safe.
+ *
+ * TODO: If we can perform heavy operations in this callback like waiting
+ * for barriers, we could set the current status of resizing in the
+ * process local memory and use this callback to rollback every operation
+ * that was performed, except buffer eviction.
+ */
+ if (!safe_exit)
+ ereport(PANIC,
+ errmsg("buffer resize operation interrupted, restarting to avoid inconsistent state"));
+
+ /*
+ * safe_exit should be set to true when new allocations are not using the
+ * whole buffer pool, so the following condition should never happen. But
+ * be on the safer side.
+ */
+ if (pg_atomic_read_u32(&BufferControl->currentNBuffers) != pg_atomic_read_u32(&BufferControl->activeNBuffers))
+ ereport(PANIC,
+ (errmsg("buffer resize operation interrupted at an unexpected stage, restarting to avoid inconsistent state")));
+
+ /*
+ * Reset targetNBuffers before releasing resizer_pid, so that a backend
+ * claiming ownership immediately afterwards does not have its own
+ * targetNBuffers overwritten by us.
+ */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, pg_atomic_read_u32(&BufferControl->currentNBuffers));
+
+ expected_pid = MyProcPid;
+ (void) pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, 0);
+}
+#endif
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC.
+ */
+bool
+ProcessBarrierNewBufferAlloc(void)
+{
+ elog(DEBUG2, "processing barrier to restrict new buffer allocations to %d buffers (target = %d)",
+ pg_atomic_read_u32(&BufferControl->activeNBuffers), pg_atomic_read_u32(&BufferControl->targetNBuffers));
+
+ INJECTION_POINT("pgrsb-handle-new-buffer-alloc-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(NBuffers == pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ return true;
+}
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE.
+ */
+bool
+ProcessBarrierBufferPoolResize(void)
+{
+ elog(DEBUG2, "processing barrier to propagate resized shared buffer pool structures");
+
+ INJECTION_POINT("pgrsb-handle-buffer-pool-resize-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(NBuffers == pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ Assert(activeNBuffers == pg_atomic_read_u32(&BufferControl->activeNBuffers));
+
+ /*
+ * Access permissions to address range covered by a resizable structure is
+ * maintained consistently across all the backends right from the time a
+ * backend is started. We maintain that consistency as the buffer pool is
+ * resized.So ideally modifying the access permissions in this backend
+ * should not fail. But if it does, the address space accessible to this
+ * backend may be inconsistent with the new buffer pool size and also with
+ * the other backends. This may cause data corruption and other memory
+ * access issues, if we let this backend continue to run and access the
+ * buffer pool. Better to quit from the faulty backend.
+ */
+ PG_TRY();
+ {
+ BufferManagerShmemProtect();
+ }
+ PG_CATCH();
+ {
+ /*
+ * We don't know what caused the error and so avoid using further
+ * resources. Emit the original error to the server log so that it's
+ * not lost and raise a FATAL to terminate this backend.
+ */
+ HOLD_INTERRUPTS();
+ errcontext("during shared buffer pool resize barrier");
+ EmitErrorReport();
+ ereport(FATAL,
+ (errmsg("shared buffer pool resize barrier caught an error while updating buffer pool protection")));
+ pg_unreachable();
+ }
+ PG_END_TRY();
+
+ return true;
+}
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE.
+ */
+bool
+ProcessBarrierBufferPoolSize(void)
+{
+ elog(DEBUG2, "processing barrier to establish new size of the buffer pool to %d", pg_atomic_read_u32(&BufferControl->currentNBuffers));
+
+ INJECTION_POINT("pgrsb-handle-buffer-pool-size-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(activeNBuffers == pg_atomic_read_u32(&BufferControl->activeNBuffers));
+ NBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+
+ return true;
+}
+
+/*
+ * SQL-callable function reporting the current shared buffer pool resize
+ * status.
+ */
+Datum
+pg_get_buffer_resize_status(PG_FUNCTION_ARGS)
+{
+#define PG_GET_BUFFER_RESIZE_STATUS_COLS 4
+ TupleDesc tupdesc;
+ Datum values[PG_GET_BUFFER_RESIZE_STATUS_COLS];
+ bool nulls[PG_GET_BUFFER_RESIZE_STATUS_COLS] = {0};
+ HeapTuple tuple;
+
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+ tupdesc = BlessTupleDesc(tupdesc);
+
+ values[0] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->activeNBuffers));
+ values[1] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ values[2] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->targetNBuffers));
+ values[3] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->resizer_pid));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ PG_RETURN_DATUM(HeapTupleGetDatum(tuple));
+#undef PG_GET_BUFFER_RESIZE_STATUS_COLS
+}
+
+/*
+ * TODO: add progress report facility if required.
+ */
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index c82b71deaa6..85664990b3c 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -28,6 +28,7 @@
#include "utils/builtins.h"
#include "storage/lwlock.h"
#include "storage/subsystems.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -58,15 +59,25 @@ BufTableShmemRequest(void *arg)
* Request the shared buffer lookup hashtable.
*
* Since we can't tolerate running out of lookup table entries, we must be
- * sure to specify an adequate table size here. The maximum steady-state
- * usage is of course as many entries as the number of buffers in the
- * pool, but BufferAlloc() tries to insert a new entry before deleting the
- * old. In principle this could be happening in each partition
- * concurrently, so we could need as many as (number of buffers in the
- * pool) + NUM_BUFFER_PARTITIONS entries. Since we are still requesting
- * shared memory, use the GUC value instead of the actual size.
+ * sure to specify an adequate table size here. The maximum number of
+ * entries we could need is the number of buffers, but BufferAlloc() tries
+ * to insert a new entry before deleting the old. In principle this could
+ * be happening in each partition concurrently, so we could need as many
+ * as (number of buffers) + NUM_BUFFER_PARTITIONS entries. The number of
+ * buffers can be increased upto MaxNBuffers at run time. So we need to
+ * make sure that the table can accommodate MaxNBuffers +
+ * NUM_BUFFER_PARTITIONS entries.
+ *
+ * See storage/buffer/README for reasons why we don't register the hash
+ * table as a resizable structure.
+ *
*/
- size = NBuffersGUC + NUM_BUFFER_PARTITIONS;
+#ifdef HAVE_RESIZABLE_SHMEM
+ size = (shared_memory_type == SHMEM_TYPE_MMAP) ? MaxNBuffers : NBuffersGUC;
+#else
+ size = NBuffersGUC;
+#endif
+ size = size + NUM_BUFFER_PARTITIONS;
ShmemRequestHash(.name = "Shared Buffer Lookup Table",
.nelems = size,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index bb6902777cd..2aac77ccc35 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -64,6 +64,7 @@
#include "storage/read_stream.h"
#include "storage/smgr.h"
#include "storage/standby.h"
+#include "utils/injection_point.h"
#include "utils/memdebug.h"
#include "utils/ps_status.h"
#include "utils/rel.h"
@@ -223,7 +224,10 @@ int io_max_combine_limit = DEFAULT_IO_COMBINE_LIMIT;
int checkpoint_flush_after = DEFAULT_CHECKPOINT_FLUSH_AFTER;
int bgwriter_flush_after = DEFAULT_BGWRITER_FLUSH_AFTER;
int backend_flush_after = DEFAULT_BACKEND_FLUSH_AFTER;
+
int NBuffers = 0; /* number of buffers in the buffer pool */
+int activeNBuffers = 0; /* number of buffers at the start of the pool
+ * from which new buffer allocations happnen. */
/* local state for LockBufferForCleanup */
static BufferDesc *PinCountWaitBuf = NULL;
@@ -814,6 +818,15 @@ PrefetchBuffer(Relation reln, ForkNumber forkNum, BlockNumber blockNum)
* Compared to ReadBuffer(), this avoids a buffer mapping lookup when it's
* successful. Return true if the buffer is valid and still has the expected
* tag. In that case, the buffer is pinned and the usage count is bumped.
+ *
+ * The callers of this function should make sure that the buffer is valid even
+ * if the shared buffer pool has undergone a resize. Buffer pool resizing waits
+ * for all backends to acknowledge the barrier before changing the buffer pool
+ * size. Hence the caller should call this function after validation without an
+ * intervening ProcSignalBarrier processing.
+ *
+ * TODO: This function could perform the validation in this function itself
+ * instead of relying on the two callers who do it currently.
*/
bool
ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockNum,
@@ -3655,6 +3668,11 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_START(NBuffers, num_to_scan);
+ /*
+ * TODO: Test the case when buffer pool is shrunk after CkptBufferIds is
+ * filled and num_to_scan is higher than the new NBuffers?
+ */
+
/*
* Sort buffers that need to be written to reduce the likelihood of random
* IO. The sorting is also important for the implementation of balancing
@@ -3768,29 +3786,43 @@ BufferSync(int flags)
buf_id = CkptBufferIds[ts_stat->index].buf_id;
Assert(buf_id != -1);
- bufHdr = GetBufferDescriptor(buf_id);
-
- num_processed++;
+ /*
+ * TODO: We need to test the scenario when the buffer pool is shrunk
+ * after checkpointer has collected the buffer ids and one or more of
+ * the buffer ids is out of range.
+ */
/*
- * We don't need to acquire the lock here, because we're only looking
- * at a single bit. It's possible that someone else writes the buffer
- * and clears the flag right after we check, but that doesn't matter
- * since SyncOneBuffer will then do nothing. However, there is a
- * further race condition: it's conceivable that between the time we
- * examine the bit here and the time SyncOneBuffer acquires the lock,
- * someone else not only wrote the buffer but replaced it with another
- * page and dirtied it. In that improbable case, SyncOneBuffer will
- * write the buffer though we didn't need to. It doesn't seem worth
- * guarding against this, though.
+ * The buffer pool might have been shrunk between the time the
+ * checkpoint collected the buffer ids and now. Skip any buffers that
+ * are out of range now; they were written when they were evicted
+ * during resizing.
*/
- if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+ if (buf_id < NBuffers)
{
- if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
+ bufHdr = GetBufferDescriptor(buf_id);
+
+ /*
+ * We don't need to acquire the lock here, because we're only
+ * looking at a single bit. It's possible that someone else writes
+ * the buffer and clears the flag right after we check, but that
+ * doesn't matter since SyncOneBuffer will then do nothing.
+ * However, there is a further race condition: it's conceivable
+ * that between the time we examine the bit here and the time
+ * SyncOneBuffer acquires the lock, someone else not only wrote
+ * the buffer but replaced it with another page and dirtied it. In
+ * that improbable case, SyncOneBuffer will write the buffer
+ * though we didn't need to. It doesn't seem worth guarding
+ * against this, though.
+ */
+ if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
{
- TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id);
- PendingCheckpointerStats.buffers_written++;
- num_written++;
+ if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
+ {
+ TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id);
+ PendingCheckpointerStats.buffers_written++;
+ num_written++;
+ }
}
}
@@ -3801,6 +3833,7 @@ BufferSync(int flags)
ts_stat->progress += ts_stat->progress_slice;
ts_stat->num_scanned++;
ts_stat->index++;
+ num_processed++;
/* Have all the buffers from the tablespace been processed? */
if (ts_stat->num_scanned == ts_stat->num_to_scan)
@@ -3843,6 +3876,13 @@ BufferSync(int flags)
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
+ * Usually new buffers are allocated from the whole buffer pool. However, when
+ * resizing the buffer pool, the new buffer allocations are restricted to some
+ * initial portion of the buffer pool. The rest of pool is either being evicted
+ * when shrinking or not utilized yet when growing, hence not interesting to the
+ * background writer. Hence the background writer always restricts its activity
+ * to the same portion of the buffer pool as new buffer allocations.
+ *
* This is called periodically by the background writer process.
*
* Returns true if it's appropriate for the bgwriter process to go into
@@ -3894,6 +3934,7 @@ BgBufferSync(WritebackContext *wb_context)
/* Variables for final smoothed_density update */
long new_strategy_delta;
uint32 new_recent_alloc;
+ static int prev_activeNBuffers;
Assert(AmBackgroundWriterProcess());
@@ -3903,6 +3944,24 @@ BgBufferSync(WritebackContext *wb_context)
*/
strategy_buf_id = StrategySyncStart(&strategy_passes, &recent_alloc);
+ if (prev_activeNBuffers != activeNBuffers)
+ {
+ /*
+ * Previous clock sweep position does not make sense if the size of
+ * the active buffer pool has changed.
+ *
+ * TODO: actually we have added a fix in
+ * StrategyAdjustNewBufAllocSize() to adjust the complete_passes so
+ * that we don't have to reset the saved info here, but still the
+ * Assert(StrategyDelta >= 0) below fails without resetting the saved
+ * info. Resetting the saved info means we loose the current position
+ * of the bgwriter and thus may have to unnecessarily scan already
+ * scanned buffers and the new allocations may have to find victim
+ * themselves. Needs further investigation.
+ */
+ saved_info_valid = false;
+ }
+
/* Report buffer alloc counts to pgstat */
PendingBgWriterStats.buf_alloc += recent_alloc;
@@ -3930,7 +3989,7 @@ BgBufferSync(WritebackContext *wb_context)
int32 passes_delta = strategy_passes - prev_strategy_passes;
strategy_delta = strategy_buf_id - prev_strategy_buf_id;
- strategy_delta += (long) passes_delta * NBuffers;
+ strategy_delta += (long) passes_delta * activeNBuffers;
if ((int32) (next_passes - strategy_passes) > 0)
{
@@ -3947,7 +4006,7 @@ BgBufferSync(WritebackContext *wb_context)
next_to_clean >= strategy_buf_id)
{
/* on same pass, but ahead or at least not behind */
- bufs_to_lap = NBuffers - (next_to_clean - strategy_buf_id);
+ bufs_to_lap = activeNBuffers - (next_to_clean - strategy_buf_id);
#ifdef BGW_DEBUG
elog(DEBUG2, "bgwriter ahead: bgw %u-%u strategy %u-%u delta=%ld lap=%d",
next_passes, next_to_clean,
@@ -3969,7 +4028,7 @@ BgBufferSync(WritebackContext *wb_context)
#endif
next_to_clean = strategy_buf_id;
next_passes = strategy_passes;
- bufs_to_lap = NBuffers;
+ bufs_to_lap = activeNBuffers;
}
/*
@@ -3996,12 +4055,13 @@ BgBufferSync(WritebackContext *wb_context)
strategy_delta = 0;
next_to_clean = strategy_buf_id;
next_passes = strategy_passes;
- bufs_to_lap = NBuffers;
+ bufs_to_lap = activeNBuffers;
}
/* Update saved info for next time */
prev_strategy_buf_id = strategy_buf_id;
prev_strategy_passes = strategy_passes;
+ prev_activeNBuffers = activeNBuffers;
saved_info_valid = true;
/*
@@ -4022,7 +4082,7 @@ BgBufferSync(WritebackContext *wb_context)
* strategy point and where we've scanned ahead to, based on the smoothed
* density estimate.
*/
- bufs_ahead = NBuffers - bufs_to_lap;
+ bufs_ahead = activeNBuffers - bufs_to_lap;
reusable_buffers_est = (float) bufs_ahead / smoothed_density;
/*
@@ -4060,7 +4120,7 @@ BgBufferSync(WritebackContext *wb_context)
* the BGW will be called during the scan_whole_pool time; slice the
* buffer pool into that many sections.
*/
- min_scan_buffers = (int) (NBuffers / (scan_whole_pool_milliseconds / BgWriterDelay));
+ min_scan_buffers = (int) (activeNBuffers / (scan_whole_pool_milliseconds / BgWriterDelay));
if (upcoming_alloc_est < (min_scan_buffers + reusable_buffers_est))
{
@@ -4082,13 +4142,24 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
+ /*
+ * Execute the LRU scan on the part of the buffer pool from where the new
+ * allocations will happen.
+ *
+ * Note that activeNBuffers may change during this loop if buffer pool
+ * gets resized concurrently. This may invalidate the num_to_scan count.
+ * If the buffer pool was shrunk this would make background writing more
+ * aggressive, which might be desired since the buffer pool is smaller. If
+ * the buffer pool was grown this would make the background writer not
+ * scan the newly added buffers, which should not have any dirty buffers
+ * yet.
+ */
while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
- if (++next_to_clean >= NBuffers)
+ if (++next_to_clean >= activeNBuffers)
{
next_to_clean = 0;
next_passes++;
@@ -4906,6 +4977,8 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
if (j >= nforks)
UnlockBufHdr(bufHdr);
}
+
+ INJECTION_POINT_CACHED("drop-relation-buffers-scan", NULL);
}
/* ---------------------------------------------------------------------
@@ -5073,6 +5146,8 @@ DropRelationsAllBuffers(SMgrRelation *smgr_reln, int nlocators)
UnlockBufHdr(bufHdr);
}
+ INJECTION_POINT("drop-relations-all-buffers-scan", NULL);
+
pfree(locators);
pfree(rels);
}
@@ -5173,6 +5248,8 @@ DropDatabaseBuffers(Oid dbid)
else
UnlockBufHdr(bufHdr);
}
+
+ INJECTION_POINT("drop-database-buffers-scan", NULL);
}
/* ---------------------------------------------------------------------
@@ -5270,6 +5347,7 @@ FlushRelationBuffers(Relation rel)
else
UnlockBufHdr(bufHdr);
}
+ INJECTION_POINT("flush-relation-buffers-scan", NULL);
}
/* ---------------------------------------------------------------------
@@ -5365,6 +5443,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
else
UnlockBufHdr(bufHdr);
}
+ INJECTION_POINT("flush-relations-all-buffers-scan", NULL);
pfree(srels);
}
@@ -8053,8 +8132,6 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
uint64 buf_state;
bool buffer_flushed;
- CHECK_FOR_INTERRUPTS();
-
buf_state = pg_atomic_read_u64(&desc->state);
if (!(buf_state & BM_VALID))
continue;
@@ -8071,6 +8148,13 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
if (buffer_flushed)
(*buffers_flushed)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8105,8 +8189,6 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
uint64 buf_state = pg_atomic_read_u64(&(desc->state));
bool buffer_flushed;
- CHECK_FOR_INTERRUPTS();
-
/* An unlocked precheck should be safe and saves some cycles. */
if ((buf_state & BM_VALID) == 0 ||
!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
@@ -8133,6 +8215,13 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
if (buffer_flushed)
(*buffers_flushed)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8250,8 +8339,6 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
uint64 buf_state = pg_atomic_read_u64(&(desc->state));
bool buffer_already_dirty;
- CHECK_FOR_INTERRUPTS();
-
/* An unlocked precheck should be safe and saves some cycles. */
if ((buf_state & BM_VALID) == 0 ||
!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
@@ -8277,6 +8364,13 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
(*buffers_already_dirty)++;
else
(*buffers_skipped)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8304,8 +8398,6 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
uint64 buf_state;
bool buffer_already_dirty;
- CHECK_FOR_INTERRUPTS();
-
buf_state = pg_atomic_read_u64(&desc->state);
if (!(buf_state & BM_VALID))
continue;
@@ -8321,6 +8413,13 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
(*buffers_already_dirty)++;
else
(*buffers_skipped)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -9017,3 +9116,57 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ *
+ * If this function encounters a pinned buffer towards the end of the buffer
+ * pool, we would have evicted most of the buffers and yet rollback the resize
+ * operation. If we could find pinned buffers before evicting any, we could save
+ * wasted work and avoid performance impact because of evicted buffers. But
+ * there is no guarantee that a buffer won't be pinned after we check it. So we
+ * have to check for pinned buffers while evicting and rollback if we encounter
+ * any.
+ */
+bool
+EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
+{
+ bool result = true;
+
+ for (Buffer buf = targetNBuffers + 1; buf <= currentNBuffers; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint64 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u64(&desc->state);
+
+ /*
+ * Nobody is expected to allocate new buffers while resizing is going
+ * on hence unlocked precheck should be safe and saves some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 4d5ee52ddc0..e397c5ce47a 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -37,8 +37,9 @@ typedef struct
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo size of the buffer
- * pool.
+ * get an actual buffer, it needs to be used modulo the size of the buffer
+ *
+ * allocation area.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -101,11 +102,52 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
static void AddBufferToRing(BufferAccessStrategy strategy,
BufferDesc *buf);
+/*
+ * StrategyWrapAround - Wrap around the clock-sweep hand.
+ *
+ * `new_pos` is the new_pos position of the clock hand after wrap around.
+ * `num_passes` is the number of completed passes when wrapping around.
+ */
+static void
+StrategyWrapAround(uint32 new_pos, uint32 num_passes)
+{
+ bool success = false;
+
+ while (!success)
+ {
+ uint32 wrapped;
+
+ /*
+ * Acquire the spinlock while increasing completePasses. That allows
+ * other readers to read nextVictimBuffer and completePasses in a
+ * consistent manner which is required for StrategySyncStart(). In
+ * theory delaying the increment could lead to an overflow of
+ * nextVictimBuffers, but that's highly unlikely and wouldn't be
+ * particularly harmful.
+ */
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ wrapped = new_pos % activeNBuffers;
+
+ success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
+ &new_pos, wrapped);
+ if (success)
+ StrategyControl->completePasses += num_passes;
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+ }
+}
+
/*
* ClockSweepTick - Helper routine for StrategyGetBuffer()
*
- * Move the clock hand one buffer ahead of its current position and return the
- * id of the buffer now under the hand.
+ * Move the clock hand one buffer ahead of its current position and return the id
+ * of the buffer now under the hand.
+ *
+ * We use the same value of activeNBuffers through out the function. Hence, even
+ * if the multiple backends end up wrapping around nextVictimBuffer using
+ * different activeNBuffers, they end up increasing completePasses incrementally
+ * and consistent to the respective victims. Use the latest activeNBuffers so as
+ * to be as consistent with the other allocators as possible.
*/
static inline uint32
ClockSweepTick(void)
@@ -119,13 +161,14 @@ ClockSweepTick(void)
*/
victim =
pg_atomic_fetch_add_u32(&StrategyControl->nextVictimBuffer, 1);
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
- if (victim >= NBuffers)
+ if (victim >= activeNBuffers)
{
uint32 originalVictim = victim;
/* always wrap what we look up in BufferDescriptors */
- victim = victim % NBuffers;
+ victim = victim % activeNBuffers;
/*
* If we're the one that just caused a wraparound, force
@@ -134,35 +177,9 @@ ClockSweepTick(void)
* value consisting of nextVictimBuffer and completePasses.
*/
if (victim == 0)
- {
- uint32 expected;
- uint32 wrapped;
- bool success = false;
-
- expected = originalVictim + 1;
-
- while (!success)
- {
- /*
- * Acquire the spinlock while increasing completePasses. That
- * allows other readers to read nextVictimBuffer and
- * completePasses in a consistent manner which is required for
- * StrategySyncStart(). In theory delaying the increment
- * could lead to an overflow of nextVictimBuffers, but that's
- * highly unlikely and wouldn't be particularly harmful.
- */
- SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
-
- wrapped = expected % NBuffers;
-
- success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
- &expected, wrapped);
- if (success)
- StrategyControl->completePasses++;
- SpinLockRelease(&StrategyControl->buffer_strategy_lock);
- }
- }
+ StrategyWrapAround(originalVictim + 1, 1);
}
+
return victim;
}
@@ -180,6 +197,11 @@ ClockSweepTick(void)
*
* The buffer is pinned and marked as owned, using TrackNewBufferPin(),
* before returning.
+ *
+ * We do not process a ProcSignalBarrier between choosing a victim and pinning
+ * it, so the buffer will remain valid even if the buffer pool is shrunk. For
+ * better safety we may want to disable interrupt handling explicitly during
+ * this time.
*/
BufferDesc *
StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
@@ -238,7 +260,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
pg_atomic_fetch_add_u32(&StrategyControl->numBufferAllocs, 1);
/* Use the "clock sweep" algorithm to find a free buffer */
- trycounter = NBuffers;
+ trycounter = activeNBuffers;
for (;;)
{
uint64 old_buf_state;
@@ -291,7 +313,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
local_buf_state))
{
- trycounter = NBuffers;
+ trycounter = activeNBuffers;
break;
}
}
@@ -334,9 +356,15 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
uint32 nextVictimBuffer;
int result;
+ /*
+ * Update backend local activeNBuffers for the same reason as in
+ * StrategyGetBuffer.
+ */
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
- result = nextVictimBuffer % NBuffers;
+ result = nextVictimBuffer % activeNBuffers;
if (complete_passes)
{
@@ -346,7 +374,7 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
* Additionally add the number of wraparounds that happened before
* completePasses could be incremented. C.f. ClockSweepTick().
*/
- *complete_passes += nextVictimBuffer / NBuffers;
+ *complete_passes += nextVictimBuffer / activeNBuffers;
}
if (num_buf_alloc)
@@ -392,6 +420,29 @@ StrategyCtlShmemRequest(void *arg)
);
}
+/*
+ * StrategyAdjustNewBufAllocSize
+ *
+ * Adjust clock hand when resizing the buffer pool.
+ */
+void
+StrategyAdjustNewBufAllocSize(void)
+{
+ int num_passes;
+ uint32 nextVictimBuffer;
+
+ /* Should be called only when resizing is in progress. */
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ /* Consistently wrap around the clock sweep hand, if necessary. */
+ nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
+ num_passes = nextVictimBuffer / activeNBuffers;
+ if (num_passes > 0)
+ StrategyWrapAround(nextVictimBuffer, num_passes);
+}
+
/*
* StrategyCtlShmemInit -- initialize the buffer cache replacement strategy.
*/
@@ -634,12 +685,27 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot is outside
+ * the buffer allocation area tell the caller to allocate a new buffer
+ * with the normal allocation strategy. He will then fill this slot by
+ * calling AddBufferToRing with the new buffer. Usually the buffers in the
+ * ring will be within the buffer allocation area, but if the buffer pool
+ * has been shrunk since the last time the ring was filled, some of the
+ * buffers in the ring may be outside the new buffer allocation area.
+ *
+ * TODO: buffer ids in the ring will never be greater than the size of
+ * buffer pool, except maybe the first time ring is accessed after
+ * shrinking the buffer pool. Checking the upper bound on buffer id always
+ * may mask a bug bugs that introduces buffer ids higher than the size of
+ * buffer pool in the ring. But performing that check only once after
+ * shrinking seems impossible. The BufferAccessStrategy objects are not
+ * accessible outside the ScanState. Hence we can not purge the buffers
+ * while evicting the buffers. After the resizing is finished, it's not
+ * possible to notice when we touch the first of those objects and the
+ * last of objects. See if this can fixed.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer || bufnum > activeNBuffers)
return NULL;
buf = GetBufferDescriptor(bufnum - 1);
diff --git a/src/backend/storage/buffer/meson.build b/src/backend/storage/buffer/meson.build
index ed84bf08971..f219e29d5ef 100644
--- a/src/backend/storage/buffer/meson.build
+++ b/src/backend/storage/buffer/meson.build
@@ -6,4 +6,5 @@ backend_sources += files(
'bufmgr.c',
'freelist.c',
'localbuf.c',
+ 'buf_resize.c',
)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 21a77f98c1d..f97215dae9c 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -28,6 +28,7 @@
#include "replication/logicalworker.h"
#include "replication/slotsync.h"
#include "replication/walsender.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
@@ -598,6 +599,15 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_CHECKSUM_OFF:
processed = AbsorbDataChecksumsBarrier(type);
break;
+ case PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC:
+ processed = ProcessBarrierNewBufferAlloc();
+ break;
+ case PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE:
+ processed = ProcessBarrierBufferPoolResize();
+ break;
+ case PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE:
+ processed = ProcessBarrierBufferPoolSize();
+ break;
}
/*
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 780cdb5e9ff..8332bc0b252 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -43,6 +43,7 @@
#include "postmaster/autovacuum.h"
#include "replication/slotsync.h"
#include "replication/syncrep.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
@@ -583,6 +584,21 @@ InitProcess(void)
* the reasons mentioned there.
*/
ShmemReprotectResizableStructs();
+
+#ifndef EXEC_BACKEND
+
+ /*
+ * Pick up the current buffer pool size from shared memory. Fork'ed
+ * backends would otherwise inherit the postmaster's potentially stale
+ * values. EXEC_BACKEND children do this via the buffer manager attach
+ * callback above.
+ *
+ * We also call this function after ProcSignalInit() for the reasons
+ * specified there, but we need it here so that InitBufferManagerAccess()
+ * can use the current buffer pool size.
+ */
+ BufferManagerInitProc();
+#endif
}
/*
@@ -772,6 +788,21 @@ InitAuxiliaryProcess(void)
* the reasons mentioned there.
*/
ShmemReprotectResizableStructs();
+
+#ifndef EXEC_BACKEND
+
+ /*
+ * Pick up the current buffer pool size from shared memory. Fork'ed
+ * backends would otherwise inherit the postmaster's potentially stale
+ * values. EXEC_BACKEND children do this via the buffer manager attach
+ * callback above.
+ *
+ * We also call this function after ProcSignalInit() for the reasons
+ * specifid there, but we need it here so that InitBufferManagerAccess()
+ * can use the current buffer pool size.
+ */
+ BufferManagerInitProc();
+#endif
}
/*
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index b6bdfe213fe..2586b7087ac 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4288,6 +4288,9 @@ PostgresSingleUserMain(int argc, char *argv[],
/* Initialize size of fast-path lock cache. */
InitializeFastPathLocks();
+ /* Initialize MaxNBuffers for buffer pool resizing. */
+ InitializeMaxNBuffers();
+
/*
* Also call any legacy shmem request hooks that might'be been installed
* by preloaded libraries.
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index ccf845e87b9..4c5cb3c3908 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -142,6 +142,8 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffersGUC = 16384;
+bool finalMaxNBuffers = false;
+int MaxNBuffers = 0;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 815b865aa38..c129d3a8a70 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -610,6 +610,55 @@ InitializeFastPathLocks(void)
pg_nextpower2_32(FastPathLockGroupsPerBackend));
}
+/*
+ * Initialize MaxNBuffers variable with validation.
+ *
+ * This must be called after GUCs have been loaded but before shared memory size
+ * is determined.
+ *
+ * Since MaxNBuffers limits the size of the buffer pool, it must be at least as
+ * much as NBuffersGUC. If MaxNBuffers is 0 (default), set it to
+ * NBuffersGUC. Otherwise, validate that MaxNBuffers is not less than
+ * NBuffersGUC.
+ */
+void
+InitializeMaxNBuffers(void)
+{
+ if (MaxNBuffers == 0) /* default/boot value */
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", NBuffersGUC);
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+
+ /*
+ * We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
+ * However, if the DBA explicitly set max_shared_buffers = 0 in the
+ * config file, then PGC_S_DYNAMIC_DEFAULT will fail to override that
+ * and we must force the matter with PGC_S_OVERRIDE.
+ */
+ if (MaxNBuffers == 0) /* failed to apply it? */
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+ else
+ {
+ if (MaxNBuffers < NBuffersGUC)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("max_shared_buffers (%d) cannot be less than current shared_buffers (%d)",
+ MaxNBuffers, NBuffersGUC),
+ errhint("Increase max_shared_buffers or decrease shared_buffers.")));
+ }
+ }
+
+ Assert(MaxNBuffers > 0);
+ Assert(!finalMaxNBuffers);
+ finalMaxNBuffers = true;
+}
+
/*
* Early initialization of a backend (either standalone or under postmaster).
* This happens even before InitPostgres.
@@ -760,32 +809,38 @@ InitPostgres(const char *in_dbname, Oid dboid,
SharedInvalBackendInit(false);
/*
- * Prevent consuming interrupts between setting ProcSignalInit and setting
- * the initial local data checksum value. If a barrier is emitted, and
- * absorbed, before local cached state is initialized the state transition
- * can be invalid.
+ * Prevent consuming interrupts between ProcSignalInit() and the
+ * initialization of state that is kept in sync with shared memory via
+ * procsignal-based barriers (currently the data_checksum_version cache
+ * and the local NBuffers/activeNBuffers cache). If a barrier is emitted,
+ * and absorbed, before that local cached state is initialized the state
+ * transition can be invalid.
*/
HOLD_INTERRUPTS();
ProcSignalInit(MyCancelKey, MyCancelKeyLength);
/*
- * Initialize a local cache of the data_checksum_version, to be updated by
- * the procsignal-based barriers.
+ * Initialize the per-backend caches that are kept in sync via
+ * procsignal-based barriers: currently the local data_checksum_version
+ * and the local copy of the buffer pool size (NBuffers/activeNBuffers).
*
- * This intentionally happens after initializing the procsignal, otherwise
- * we might miss a state change. This means we can get a barrier for the
- * state we've just initialized.
+ * These initializations intentionally happen after ProcSignalInit(),
+ * otherwise we might miss a state change. This means we may also receive
+ * a barrier for the state we've just initialized.
*
* The postmaster (which is what gets forked into the new child process)
* does not handle barriers, therefore it may not have the current value
- * of LocalDataChecksumState value (it'll have the value read from the
- * control file, which may be arbitrarily old).
+ * of LocalDataChecksumState (it'll have the value read from the control
+ * file, which may be arbitrarily old) or NBuffers/activeNBuffers (which
+ * may have been changed by an online resize after the postmaster
+ * started).
*
* NB: Even if the postmaster handled barriers, the value might still be
* stale, as it might have changed after this process forked.
*/
InitLocalDataChecksumState();
+ BufferManagerInitProc();
/*
* Refresh per-backend protections for resizable shmem structures. Usually
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 1a5a168bc1a..5f1dd338f56 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -2631,7 +2631,7 @@ convert_to_base_unit(double value, const char *unit,
* the value without loss. For example, if the base unit is GUC_UNIT_KB, 1024
* is converted to 1 MB, but 1025 is represented as 1025 kB.
*/
-static void
+void
convert_int_from_base_unit(int64 base_value, int base_unit,
int64 *value, const char **unit)
{
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index db64f8b2601..ccd111b542c 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2146,6 +2146,15 @@
max => 'MAX_BACKENDS /* XXX? */',
},
+{ name => "max_shared_buffers", type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+ short_desc => 'Sets the upper limit for the shared_buffers value.',
+ flags => 'GUC_UNIT_BLOCKS',
+ variable => 'MaxNBuffers',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX / 2',
+},
+
{ name => 'max_slot_wal_keep_size', type => 'int', context => 'PGC_SIGHUP', group => 'REPLICATION_SENDING',
short_desc => 'Sets the maximum WAL size that can be reserved by replication slots.',
long_desc => 'Replication slots will be marked as failed, and segments released for deletion or recycling, if this much space is occupied by WAL on disk. -1 means no maximum.',
@@ -2729,13 +2738,15 @@
# We sometimes multiply the number of shared buffers by two without
# checking for overflow, so we mustn't allow more than INT_MAX / 2.
-{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
variable => 'NBuffersGUC',
boot_val => '16384',
- min => '16',
+ min => 'MIN_NUM_BUFFERS',
max => 'INT_MAX / 2',
+ check_hook => 'check_shared_buffers',
+ show_hook => 'show_shared_buffers',
},
{ name => 'shared_memory_initial_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index e1f5945a971..ac7d9033a57 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12776,4 +12776,19 @@
proname => 'hashoid8extended', prorettype => 'int8',
proargtypes => 'oid8 int8', prosrc => 'hashoid8extended' },
+{ oid => '9999', descr => 'resize shared buffers according to the value of GUC `shared_buffers`',
+ proname => 'pg_resize_shared_buffers',
+ provolatile => 'v',
+ prorettype => 'bool',
+ proargtypes => '',
+ prosrc => 'pg_resize_shared_buffers'},
+{ oid => '9998', descr => 'report shared buffer pool resize status',
+ proname => 'pg_get_buffer_resize_status',
+ provolatile => 'v',
+ prorettype => 'record',
+ proargtypes => '',
+ proallargtypes => '{int4,int4,int4,int4}',
+ proargmodes => '{o,o,o,o}',
+ proargnames => '{active_nbuffers,current_nbuffers,target_nbuffers,resizer_pid}',
+ prosrc => 'pg_get_buffer_resize_status'},
]
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index e2df51d275a..772f92057e8 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -176,6 +176,8 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffersGUC;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
@@ -515,6 +517,7 @@ extern PGDLLIMPORT ProcessingMode Mode;
extern void pg_split_opts(char **argv, int *argcp, const char *optstr);
extern void InitializeMaxBackends(void);
extern void InitializeFastPathLocks(void);
+extern void InitializeMaxNBuffers(void);
extern void InitPostgres(const char *in_dbname, Oid dboid,
const char *username, Oid useroid,
uint32 flags,
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 678065b3de2..a4971cebde4 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -264,6 +264,33 @@ BufMappingPartitionLockByIndex(uint32 index)
return &MainLWLockArray[BUFFER_MAPPING_LWLOCK_OFFSET + index].lock;
}
+/*
+ * BufferControl -- shared area controlling buffer pool
+ *
+ * This structure stores information about the size of the buffer pool and
+ * whether it is being resized.
+ */
+typedef struct BufferControlBlock
+{
+ /*
+ * size of the part of the buffer pool from where buffers are being
+ * allocated to new requests.
+ */
+ pg_atomic_uint32 activeNBuffers;
+
+ /* current size of the buffer pool, in number of buffers */
+ pg_atomic_uint32 currentNBuffers;
+
+ /* target size of the buffer pool, in number of buffers */
+ pg_atomic_uint32 targetNBuffers;
+
+ /*
+ * PID of the backend currently performing a resize, or 0 when no resize
+ * is in progress. Also acts as a lock prohibiting concurrent resizes.
+ */
+ pg_atomic_uint32 resizer_pid;
+} BufferControlBlock;
+
/*
* BufferDesc -- shared descriptor/state data for a single shared buffer.
*
@@ -411,6 +438,7 @@ typedef struct WritebackContext
} WritebackContext;
/* in buf_init.c */
+extern PGDLLIMPORT BufferControlBlock *BufferControl;
extern PGDLLIMPORT BufferDescPadded *BufferDescriptors;
extern PGDLLIMPORT ConditionVariableMinimallyPadded *BufferIOCVArray;
extern PGDLLIMPORT WritebackContext BackendWritebackContext;
@@ -422,9 +450,26 @@ extern PGDLLIMPORT BufferDesc *LocalBufferDescriptors;
static inline BufferDesc *
GetBufferDescriptor(int id)
{
+ BufferDesc *bdesc;
+
Assert(id >= 0 && id < NBuffers);
- return &(BufferDescriptors[id]).bufferdesc;
+ bdesc = &(BufferDescriptors[id]).bufferdesc;
+
+ /*
+ * TODO: This assertion was proposed in
+ * https://www.postgresql.org/message-id/CAExHW5uzRMYVZsXXS3HXXT0fG_sNrpUhUqwP4NorhaCqH9JDhA@mail.gmail.com,
+ * but was ultimately removed since there was no adequate reason to keep
+ * it in the code without shared buffer resizing. With resizing we may
+ * write and rewrite parts of the buffer descriptor array. So it's better
+ * to make sure that the buffer descriptor is initialized correctly. For
+ * now just make sure that the id in the buffer descriptor is the same as
+ * the id used to access it. Later we may want to expand the assertion to
+ * check the BufferDesc invariants or remove this assertion.
+ */
+ Assert(bdesc->buf_id == id);
+
+ return bdesc;
}
static inline BufferDesc *
@@ -594,6 +639,7 @@ extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
extern int StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc);
extern void StrategyNotifyBgWriter(int bgwprocno);
+extern void StrategyAdjustNewBufAllocSize(void);
/* buf_table.c */
extern uint32 BufTableHashCode(BufferTag *tagPtr);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index f1f6e601f51..187e9b29a9e 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -14,6 +14,7 @@
#ifndef BUFMGR_H
#define BUFMGR_H
+#include "fmgr.h"
#include "port/pg_iovec.h"
#include "storage/aio_types.h"
#include "storage/block.h"
@@ -159,6 +160,7 @@ typedef struct ReadBuffersOperation ReadBuffersOperation;
typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
+#define MIN_NUM_BUFFERS 16
extern PGDLLIMPORT int NBuffersGUC;
/* in bufmgr.c */
@@ -167,6 +169,7 @@ extern PGDLLIMPORT int bgwriter_lru_maxpages;
extern PGDLLIMPORT double bgwriter_lru_multiplier;
extern PGDLLIMPORT bool track_io_timing;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int activeNBuffers;
#define DEFAULT_EFFECTIVE_IO_CONCURRENCY 16
#define DEFAULT_MAINTENANCE_IO_CONCURRENCY 16
@@ -371,6 +374,12 @@ extern void MarkDirtyRelUnpinnedBuffers(Relation rel,
extern void MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
int32 *buffers_already_dirty,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int targetNBuffers, int currentNBuffers);
+
+/* in buf_init.c */
+extern bool BufferManagerShmemResize(int currentNBuffers, int targetNBuffers);
+extern void BufferManagerShmemProtect(void);
+extern void BufferManagerInitProc(void);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
@@ -473,4 +482,10 @@ BufferGetPage(Buffer buffer)
#endif /* FRONTEND */
+/* buf_resize.c */
+extern Datum pg_resize_shared_buffers(PG_FUNCTION_ARGS);
+extern bool ProcessBarrierNewBufferAlloc(void);
+extern bool ProcessBarrierBufferPoolResize(void);
+extern bool ProcessBarrierBufferPoolSize(void);
+
#endif /* BUFMGR_H */
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index aaa158bfd66..78dda9c8d21 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,11 @@ typedef enum
PROCSIGNAL_BARRIER_CHECKSUM_INPROGRESS_ON,
PROCSIGNAL_BARRIER_CHECKSUM_INPROGRESS_OFF,
PROCSIGNAL_BARRIER_CHECKSUM_ON,
+ PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC, /* New buffer allocation pool size
+ * changed */
+ PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE, /* Buffer pool shared structures
+ * resized */
+ PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE, /* Buffer pool size updated */
} ProcSignalBarrierType;
/*
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 2a6e2ed18b3..2e3dc919018 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -462,6 +462,8 @@ extern config_handle *get_config_handle(const char *name);
extern void AlterSystemSetConfigFile(AlterSystemStmt *altersysstmt);
extern char *GetConfigOptionByName(const char *name, const char **varname,
bool missing_ok);
+extern void convert_int_from_base_unit(int64 base_value, int base_unit,
+ int64 *value, const char **unit);
extern void TransformGUCArray(ArrayType *array, List **names,
List **values);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index df048517a0f..7e549a36530 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -181,4 +181,6 @@ extern void assign_synchronized_standby_slots(const char *newval, void *extra);
extern bool check_log_min_messages(char **newval, void **extra, GucSource source);
extern void assign_log_min_messages(const char *newval, void *extra);
+extern const char *show_shared_buffers(bool use_units);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
#endif /* GUC_HOOKS_H */
diff --git a/src/test/Makefile b/src/test/Makefile
index 3eb0a06abb4..7a0d74086c1 100644
--- a/src/test/Makefile
+++ b/src/test/Makefile
@@ -20,7 +20,8 @@ SUBDIRS = \
postmaster \
recovery \
regress \
- subscription
+ subscription \
+ buffermgr
ifeq ($(with_icu),yes)
SUBDIRS += icu
diff --git a/src/test/README b/src/test/README
index afdc7676519..77f11607ff7 100644
--- a/src/test/README
+++ b/src/test/README
@@ -15,6 +15,9 @@ examples/
Demonstration programs for libpq that double as regression tests via
"make check"
+buffermgr/
+ Tests for resizing buffer pool without restarting the server
+
isolation/
Tests for concurrent behavior at the SQL level
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
new file mode 100644
index 00000000000..92a430e736b
--- /dev/null
+++ b/src/test/buffermgr/Makefile
@@ -0,0 +1,40 @@
+#-------------------------------------------------------------------------
+#
+# Makefile for src/test/buffermgr
+#
+# Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+# Portions Copyright (c) 1994, Regents of the University of California
+#
+# src/test/buffermgr/Makefile
+#
+#-------------------------------------------------------------------------
+
+EXTRA_INSTALL = contrib/amcheck \
+ contrib/pg_buffercache \
+ contrib/pg_prewarm \
+ src/test/modules/injection_points \
+ src/test/modules/test_shmem
+
+EXTENSION = buffermgr_test
+DATA = buffermgr_test--1.0.sql
+
+REGRESS = buffer_resize
+
+# Custom configuration for buffer manager tests
+TEMP_CONFIG = $(srcdir)/buffermgr_test.conf
+
+export enable_injection_points
+
+subdir = src/test/buffermgr
+top_builddir = ../../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+
+check:
+ $(prove_check)
+
+installcheck:
+ $(prove_installcheck)
+
+clean distclean:
+ rm -rf tmp_check
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
new file mode 100644
index 00000000000..1010f8e076c
--- /dev/null
+++ b/src/test/buffermgr/README
@@ -0,0 +1,43 @@
+src/test/buffermgr/README
+
+Regression tests for buffer manager
+===================================
+
+This directory contains a test suite for resizing buffer manager without
+restarting the server.
+
+Some of the TAP tests rely on a helper extension buffermgr_test. It bundles SQL
+helpers (such as pg_resize_shared_buffers_sql) used by the tests.
+
+Stress tests
+------------------
+These tests exercise the synchronization between the buffer pool resizing and
+the code that uses buffer manager. They run the SQL commands that exercise the
+specific subsystem functionality in parallel with the resizing of buffer
+manager, both in a tight loop so as to increase the chances of hitting race
+conditions. Optionally they run pgbench to keep the buffer pool busy.
+
+Injection point tests
+------------------
+Any issues discovered during the stress tests can be converted into injection
+point tests. Injection points allow to simulate a scenario precisely and
+deterministically.
+
+Running the tests
+=================
+
+NOTE: You must have given the --enable-tap-tests argument to configure.
+
+Run
+ make check
+or
+ make installcheck
+You can use "make installcheck" if you previously did "make install".
+In that case, the code in the installation tree is tested. With
+"make check", a temporary installation tree is built from the current
+sources and then tested.
+
+Either way, this test initializes, starts, and stops a test Postgres
+cluster.
+
+See src/test/perl/README for more info about running these tests.
diff --git a/src/test/buffermgr/buffermgr_test--1.0.sql b/src/test/buffermgr/buffermgr_test--1.0.sql
new file mode 100644
index 00000000000..51901508523
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test--1.0.sql
@@ -0,0 +1,52 @@
+-- Helper functions used by TAP tests in src/test/buffermgr/t.
+--
+
+\echo Use "CREATE EXTENSION buffermgr_test" to load this file. \quit
+
+-- Retries pg_resize_shared_buffers() until it succeeds, then confirms the
+-- new value is in effect. Returns the number of retries taken along with
+-- the wall-clock times immediately before and after the retry loop.
+--
+-- The new size is expected to be set in shared_buffers GUC before calling this
+-- function.
+create function pg_resize_shared_buffers_sql(
+ new_size int,
+ out num_tries int,
+ out started_at timestamptz,
+ out ended_at timestamptz)
+returns record as $$
+declare
+ success boolean := false;
+ tries int := 0;
+ cur_setting text;
+ target text := new_size::text;
+ pending_pattern text := '%(pending: ' || target || ')%';
+begin
+ select setting into cur_setting
+ from pg_settings where name = 'shared_buffers';
+ if cur_setting <> target and cur_setting not like pending_pattern then
+ raise exception 'shared_buffers change not visible to this backend: setting is %, expected % or matching %',
+ cur_setting, target, pending_pattern;
+ end if;
+
+ started_at := clock_timestamp();
+ while not success loop
+ tries := tries + 1;
+ select pg_resize_shared_buffers() into success;
+ if not success then
+ perform pg_sleep(0.1);
+ end if;
+ end loop;
+ ended_at := clock_timestamp();
+
+ select setting into cur_setting
+ from pg_settings where name = 'shared_buffers';
+ if cur_setting <> target then
+ raise exception 'shared_buffers resize did not take effect: expected %, got %',
+ target, cur_setting;
+ end if;
+
+ num_tries := tries;
+ return;
+end;
+$$ language plpgsql;
diff --git a/src/test/buffermgr/buffermgr_test.conf b/src/test/buffermgr/buffermgr_test.conf
new file mode 100644
index 00000000000..a15f3e442a5
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.conf
@@ -0,0 +1,11 @@
+# Configuration for buffer manager regression tests
+
+# Even if max_shared_buffers is set multiple times only the last one is used to
+# as the limit on shared_buffers.
+max_shared_buffers = 128kB
+# Set initial shared_buffers as expected by test
+shared_buffers = 128MB
+# Set a larger value for max_shared_buffers to allow testing resize operations
+max_shared_buffers = 300MB
+# Turn huge pages off, since that affects the size of memory segments
+huge_pages = off
diff --git a/src/test/buffermgr/buffermgr_test.control b/src/test/buffermgr/buffermgr_test.control
new file mode 100644
index 00000000000..2d13d889ec7
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.control
@@ -0,0 +1,3 @@
+comment = 'Helpers for src/test/buffermgr TAP tests'
+default_version = '1.0'
+relocatable = true
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
new file mode 100644
index 00000000000..1dc3efd830a
--- /dev/null
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -0,0 +1,290 @@
+-- Test buffer pool resizing and shared memory allocation tracking This test
+-- resizes the buffer pool multiple times and monitors shared memory allocations
+-- related to buffer management
+-- TODOs
+--
+-- 1. The test sets shared_buffers values in MBs. Instead it could use values in
+-- kBs so that the test runs on very small machines.
+--
+-- 2. The size, minimum_size and maximum_size columns in pg_shmem_allocations
+-- for "Buffer Blocks" should be same as the value of GUC shared_buffers. We
+-- should test that.
+--
+-- 3. We should make sure that when the shared_buffers value is increased, the
+-- size and allocated_size for all buffer related shared memory allocations
+-- increases and when the shared_buffers value is decreased, the size and
+-- allocated_size for all buffer related shared memory allocations decreases
+-- proportionately.
+--
+-- 4. allocated_size for allocations should be greater than or equal to size for
+-- all buffer related shared memory allocations. Similarly reserved_space should
+-- be greater than or equal to maximum_size for all buffer related shared memory
+-- allocations. We should test these conditions as well.
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+-- Load test_shmem for test_shmem_pagesize().
+CREATE EXTENSION IF NOT EXISTS test_shmem;
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, size,
+ allocated_size >= size AS alloc_size_cmp,
+ allocated_size - size < 2 * test_shmem_pagesize() AS alloc_size_diff,
+ minimum_size, maximum_size, reserved_space
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 134217728 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1048576 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 262144 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 134217728 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1048576 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 262144 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 128MB (pending: 64MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 64MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 67108864 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 524288 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 131072 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 8192
+(1 row)
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 64MB (pending: 256MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 268435456 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 2097152 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 524288 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 32768
+(1 row)
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 256MB (pending: 100MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 100MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 104857600 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 819200 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 204800 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 12800
+(1 row)
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 100MB (pending: 128kB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+--------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 131072 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1024 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 256 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16
+(1 row)
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+ERROR: invalid value for parameter "shared_buffers": 51200
+DETAIL: "shared_buffers" must be less than "max_shared_buffers".
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+-- TODO: Test that a non-superuser can not invoke pg_resize_shared_buffers()
+-- function.
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
new file mode 100644
index 00000000000..3c64d2740ee
--- /dev/null
+++ b/src/test/buffermgr/meson.build
@@ -0,0 +1,37 @@
+# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+
+test_install_data += files(
+ 'buffermgr_test.control',
+ 'buffermgr_test--1.0.sql',
+)
+
+tests += {
+ 'name': 'buffermgr',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'buffer_resize',
+ ],
+ 'regress_args': ['--temp-config', files('buffermgr_test.conf')],
+ },
+ 'tap': {
+ 'env': {
+ 'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
+ },
+ 'tests': [
+ 't/001_resize_fault_tolerance.pl',
+ 't/002_client_join_buffer_resize.pl',
+ 't/003_resize_failures.pl',
+ 't/004_resize_with_syslogger.pl',
+ 't/005_resize_unsupported.pl',
+ 't/010_stress_resize_buffer.pl',
+ 't/011_stress_drop_relation_buffers.pl',
+ 't/012_stress_drop_database_buffers.pl',
+ 't/013_stress_checkpoint.pl',
+ 't/014_stress_flush_relation_buffers.pl',
+ 't/015_stress_pg_buffercache.pl',
+ 't/016_stress_pg_prewarm.pl',
+ ],
+ },
+}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
new file mode 100644
index 00000000000..4219e33a805
--- /dev/null
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -0,0 +1,106 @@
+-- Test buffer pool resizing and shared memory allocation tracking This test
+-- resizes the buffer pool multiple times and monitors shared memory allocations
+-- related to buffer management
+
+-- TODOs
+--
+-- 1. The test sets shared_buffers values in MBs. Instead it could use values in
+-- kBs so that the test runs on very small machines.
+--
+-- 2. The size, minimum_size and maximum_size columns in pg_shmem_allocations
+-- for "Buffer Blocks" should be same as the value of GUC shared_buffers. We
+-- should test that.
+--
+-- 3. We should make sure that when the shared_buffers value is increased, the
+-- size and allocated_size for all buffer related shared memory allocations
+-- increases and when the shared_buffers value is decreased, the size and
+-- allocated_size for all buffer related shared memory allocations decreases
+-- proportionately.
+--
+-- 4. allocated_size for allocations should be greater than or equal to size for
+-- all buffer related shared memory allocations. Similarly reserved_space should
+-- be greater than or equal to maximum_size for all buffer related shared memory
+-- allocations. We should test these conditions as well.
+
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+
+-- Load test_shmem for test_shmem_pagesize().
+CREATE EXTENSION IF NOT EXISTS test_shmem;
+
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, size,
+ allocated_size >= size AS alloc_size_cmp,
+ allocated_size - size < 2 * test_shmem_pagesize() AS alloc_size_diff,
+ minimum_size, maximum_size, reserved_space
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+
+-- TODO: Test that a non-superuser can not invoke pg_resize_shared_buffers()
+-- function.
diff --git a/src/test/buffermgr/t/001_resize_fault_tolerance.pl b/src/test/buffermgr/t/001_resize_fault_tolerance.pl
new file mode 100644
index 00000000000..1f109aece51
--- /dev/null
+++ b/src/test/buffermgr/t/001_resize_fault_tolerance.pl
@@ -0,0 +1,916 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test that only one pg_resize_shared_buffers() call succeeds when multiple
+# sessions attempt to resize buffers concurrently
+
+use strict;
+use warnings;
+use Config;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# =============================================================================
+# Initialization
+# =============================================================================
+my $initial_nbuffers = 16;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf',
+ 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
+$node->append_conf('postgresql.conf', 'max_shared_buffers = 32');
+$node->append_conf('postgresql.conf', 'restart_after_crash = on');
+$node->start;
+
+# Bail out if this build does not support resizable shared memory, which
+# also means that resizing buffer pool is not supported.
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
+# Load injection points extension for test coordination
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# =============================================================================
+# Helper functions
+# =============================================================================
+
+# Setup resize operation to be interrupted.
+#
+# Prepare to resize the buffer pool to a target size. Start a resize session
+# through a background psql session. Adjust GUCs for the mode of interruption.
+# If injection point is provided, setup it up with the injection point and wait
+# for the resize session to reach the injection point. The resize session is
+# returned to the caller.
+sub start_resize_session
+{
+ my ($target_nbuffers, $mode, $injection_point) = @_;
+
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$target_nbuffers'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ my $session = $node->background_psql('postgres', on_error_stop => 0);
+
+ my $injection_action;
+ if (defined $injection_point)
+ {
+ $injection_action = ($mode eq 'error') ? 'error' : 'wait';
+ $session->query_safe('SELECT injection_points_set_local()',
+ verbose => 0);
+ $session->query_safe(
+ "SELECT injection_points_attach('$injection_point', '$injection_action')",
+ verbose => 0);
+ }
+
+ apply_session_gucs_for_mode($session, $mode);
+
+ $session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ ));
+
+ # Wait for the pg_resize_shared_buffers to start waiting at the injection
+ # point.
+ if (defined $injection_point && $injection_action eq 'wait')
+ {
+ my $resize_pid = $session->{backend_pid};
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = '$injection_point' FROM pg_stat_activity WHERE pid = $resize_pid"
+ )
+ or die
+ "timed out waiting for resize backend $resize_pid at $injection_point";
+ }
+
+ return $session;
+}
+
+# Start a backend which can be used to test the barrier handler fault tolerance.
+# We use a long pg_sleep() to simulate a load that checks for interrupts
+# regularly. The given injection point is attached to the peer backend locally
+# to induce a fault in barrier handler.
+sub start_peer_session_with_injection_point
+{
+ my ($injection_point, $action) = @_;
+
+ my $session = $node->background_psql('postgres', on_error_stop => 0);
+
+ $session->query_safe('SELECT injection_points_set_local()', verbose => 0);
+ $session->query_safe(
+ "SELECT injection_points_attach('$injection_point', '$action')",
+ verbose => 0);
+ $session->query_until(
+ qr/starting_sleep/,
+ q(
+ \echo starting_sleep
+ SELECT pg_sleep(60);
+ ));
+
+ my $peer_pid = $session->{backend_pid};
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = 'PgSleep' FROM pg_stat_activity WHERE pid = $peer_pid"
+ ) or die "timed out waiting for peer $peer_pid to enter pg_sleep";
+
+ return $session;
+}
+
+# Apply per-mode session GUCs locally in the given session if required.
+sub apply_session_gucs_for_mode
+{
+ my ($session, $mode) = @_;
+
+ # In timeout mode, set a statement timeout long enough for the resizing
+ # session to reach the injection point and stay there but short enough that
+ # the test doesn't take too long to fail if something goes wrong.
+ if ($mode eq 'timeout')
+ {
+ $session->query_safe("SET statement_timeout = '500ms'", verbose => 0);
+ }
+
+ # Let the resize session detect a client disconnection when testing client
+ # disconnections.
+ if ($mode eq 'disconnect')
+ {
+ $session->query_safe("SET client_connection_check_interval = '100ms'",
+ verbose => 0);
+ }
+}
+
+# Administer the interrupt corresponding to $mode against a resize session
+# that is waiting to be interrupted while resizing the buffer pool.
+sub interrupt_resize_session
+{
+ my ($mode, $session) = @_;
+
+ if ($mode eq 'terminate')
+ {
+ $node->safe_psql('postgres',
+ "SELECT pg_terminate_backend(" . $session->{backend_pid} . ")");
+ }
+ elsif ($mode eq 'cancel')
+ {
+ $node->safe_psql('postgres',
+ "SELECT pg_cancel_backend(" . $session->{backend_pid} . ")");
+ }
+ elsif ($mode eq 'disconnect')
+ {
+ $session->{run}->kill_kill;
+ }
+ elsif ($mode eq 'timeout')
+ {
+ # Nothing to do; statement_timeout will fire from within the resize
+ # session itself.
+ }
+ elsif ($mode eq 'error')
+ {
+ # Nothing to do; the injection point raised ERROR from within the
+ # resize backend itself.
+ }
+ else
+ {
+ die "interrupt_resize_session: unknown mode '$mode'";
+ }
+}
+
+# Function to perform checks after the resize operation has been interrupted. As
+# a result of the interruption, the resize function may finish rolling back the
+# resize or the backend executing that function may exit rolling back the resize
+# or the postmaster may restart all the backends. Perform appropriate checks by
+# detecting the post-interrupt state.
+#
+# - sentinel_session: a background psql session that is used to detect whether the
+# postmaster restarted all backends or not.
+# - resize_session: the background psql session that was executing the resize
+# operation and was interrupted.
+# - log_offset: the offset in the server log file before the resize operation was
+# initiated.
+# - injection_point and mode: the injection point and mode of interruption that
+# was used to interrupt the resize operation.
+# - orig_nbuffers and target_nbuffers: the original and target buffer sizes for
+# the resize operation.
+# - test_label: a label to create unique test names for different tests
+sub check_interrupted_resize
+{
+ my ($sentinel_session, $resize_session, $log_offset, $mode,
+ $injection_point, $orig_nbuffers, $target_nbuffers, $test_label)
+ = @_;
+
+ my $resize_pid = $resize_session->{backend_pid};
+ my $sentinel_pid = $sentinel_session->{backend_pid};
+
+ # Wait until the resize backend is no longer running the resize query.
+ $node->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity "
+ . "WHERE pid = $resize_pid AND state = 'active' "
+ . "AND query LIKE '%pg_resize_shared_buffers%'")
+ or die "timed out waiting for resize backend $resize_pid to finish";
+
+ # Wait for the postmaster to be ready in case it restarted the backends.
+ $node->poll_query_until('postgres', 'SELECT true')
+ or die "timed out waiting for postmaster liveliness check";
+
+ # Confirm the resize backend was interrupted by the intended signal.
+ # Match against the log line emitted by the resize PID so we don't
+ # accidentally pick up an unrelated message.
+ my %expected_msg = (
+ terminate => 'terminating connection due to administrator command',
+ cancel => 'canceling statement due to user request',
+ timeout => 'canceling statement due to statement timeout',
+ disconnect => 'connection to client lost',
+ error => "error triggered for injection point $injection_point",);
+ my $log_pattern = qr/\[$resize_pid\][^\n]*\Q$expected_msg{$mode}\E/;
+
+ $node->wait_for_log($log_pattern, $log_offset);
+ ok( $node->log_contains($log_pattern, $log_offset),
+ "$test_label: server log shows expected $mode message from pid $resize_pid"
+ );
+
+ my $server_restarted = $node->safe_psql('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE pid = $sentinel_pid")
+ eq 't';
+
+ if ($server_restarted)
+ {
+ # Postmaster restarted all backends; sentinel and resize sessions
+ # are dead, just reap their IPC::Run handles.
+ $sentinel_session->finish;
+ $resize_session->finish;
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after crash recovery"
+ );
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"
+ ),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after crash recovery"
+ );
+ }
+ else
+ {
+ $sentinel_session->quit;
+ # The resize session may have exited (e.g. on FATAL or disconnect).
+ if ($node->safe_psql('postgres',
+ "SELECT count(*) = 1 FROM pg_stat_activity WHERE pid = $resize_pid"
+ ) eq 't')
+ {
+ $resize_session->quit;
+ }
+ else
+ {
+ $resize_session->finish;
+ }
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$orig_nbuffers|$orig_nbuffers|$orig_nbuffers|0",
+ "$test_label: buffer resize rolled back after $mode");
+
+ # TODO: Also check that the pg_shmem_allocations values are not changed
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"
+ ),
+ "$orig_nbuffers (pending: $target_nbuffers)",
+ "$test_label: pg_settings reports pending new value after $mode");
+
+ is( $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't',
+ "$test_label: resize succeeds after interrupted resize is cleaned up"
+ );
+ }
+}
+
+# =============================================================================
+# Concurrent resize test functions
+#
+# Verify that only one pg_resize_shared_buffers() call can succeed at a time
+# using injection points.
+# =============================================================================
+
+# Workhorse function:
+#
+# Make the resize session wait at the given injection point and start another
+# concurrent resize session. The concurrent resize should fail.
+sub test_concurrent_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $session = start_resize_session($target_nbuffers, 'concurrent_resize',
+ $injection_point);
+ my $resize_pid = $session->{backend_pid};
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$resize_pid",
+ "$test_label: resizer_pid reports resize backend");
+
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 'f', "$test_label: concurrent resize fails");
+
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('$injection_point')");
+
+ $session->quit;
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool resized to target after wakeup");
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after wakeup");
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points
+sub test_concurrent_resize
+{
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'pgrsb-buffer-pool-resize-barrier-sent',);
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = ([ 'expand', 24 ], [ 'shrink', $initial_nbuffers ]);
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_concurrent_resize_at_injection_point($point, $target,
+ "$name: $point");
+ }
+ }
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test resize operation interruption
+#
+# Verify that an interruption in resize operation does not leave the buffer pool
+# in an inconsistent state.
+# =============================================================================
+
+# Workhorse function:
+#
+# Interrupt pg_resize_shared_buffers() when it is waiting on an injection point.
+# Check that the buffer pool is left in a consistent state as an aftermath.
+#
+# $mode selects how the resize session is interrupted.
+# - 'terminate' - SIGTERM via pg_terminate_backend() from another session.
+# - 'cancel' - SIGINT via pg_cancel_backend() from another session.
+# - 'timeout' - statement_timeout fires inside the resize session itself.
+# - 'disconnect' - the resize session's client connection is closed abruptly.
+# - 'error' - the injection point itself raises ERROR from within the
+# resize backend.
+sub test_interrupt_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $mode, $test_label) = @_;
+
+ my $orig_nbuffers = $node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+ my $log_offset = -s $node->logfile;
+
+ # Start a sentinel session that will be used to detect whether the
+ # postmaster restarted all backends or not after the resize session is
+ # interrupted.
+ my $sentinel_session =
+ $node->background_psql('postgres', on_error_stop => 0);
+
+ my $resize_session =
+ start_resize_session($target_nbuffers, $mode, $injection_point);
+
+ interrupt_resize_session($mode, $resize_session);
+
+ check_interrupted_resize($sentinel_session, $resize_session, $log_offset,
+ $mode, $injection_point, $orig_nbuffers, $target_nbuffers,
+ $test_label);
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points passing it the
+# given mode of interruption.
+sub test_interrupt_resize_session
+{
+ my ($mode) = @_;
+
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'buffer-mgr-resize-struct',
+ 'pgrsb-buffer-pool-resize-barrier-sent',);
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = ([ 'expand', 24 ], [ 'shrink', $initial_nbuffers ]);
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "$mode: buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_interrupt_resize_at_injection_point($point, $target, $mode,
+ "$mode $name: $point");
+ }
+ }
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "$mode: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test error handling in barrier handler
+#
+# Verify that an error in barrier handler does not cause a resize session to
+# fail. The barrier handler may run in a peer backend or the backend which is
+# performing the resize itself.
+# =============================================================================
+
+# Workhorse function:
+#
+# Make a peer session wait at the given injection point in the barrier handler
+# and simulate an error in the handler.
+sub test_error_in_barrier_handler_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $peer_session =
+ start_peer_session_with_injection_point($injection_point, 'error');
+ my $peer_pid = $peer_session->{backend_pid};
+
+ my $log_offset = -s $node->logfile;
+
+ # Resize the buffer pool which will send a barrier to the peer backend
+ # simulating an error in the barrier handler.
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$target_nbuffers'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't');
+
+ # Confirm the peer raised the expected error from inside the handler.
+ my $log_pattern =
+ qr/\[$peer_pid\][^\n]*\Qerror triggered for injection point $injection_point\E/;
+ $node->wait_for_log($log_pattern, $log_offset);
+ ok( $node->log_contains($log_pattern, $log_offset),
+ "$test_label: server log shows error from peer pid $peer_pid at $injection_point"
+ );
+
+ # Check that the resize was completed as expected
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size");
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size");
+
+ # pg_resize_shared_buffers() returns only after every peer has
+ # acknowledged the barrier, so by this point the erroring peer has
+ # already left procArray. Assert that and then reap its IPC::Run handle.
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT count(*) FROM pg_stat_activity WHERE pid = $peer_pid"),
+ '0',
+ "$test_label: peer pid $peer_pid exited after handler error");
+
+ $peer_session->finish;
+}
+
+# Driver function:
+#
+# Simulate a failure to change the protection on the shared memory. This should
+# cause the barrier handler to raise an error. The barrier handler may run in a
+# peer backend or the backend which is performing the resize itself. The resize
+# session should still complete successfully.
+sub test_error_in_barrier_handler
+{
+ my $injection_point = 'buffer-mgr-protect-struct';
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "error-in-handler: buffer pool size is $initial_nbuffers at start");
+
+ # Test error in barrier handler in a peer backend. Expand then shrink so
+ # the pool returns to its starting size.
+ for my $dir ([ 'expand', 24 ], [ 'shrink', $initial_nbuffers ])
+ {
+ my ($name, $target) = @$dir;
+ test_error_in_barrier_handler_at_injection_point($injection_point,
+ $target, "error-in-handler peer $name");
+ }
+
+ # Test the same error in the barrier handler in the resize backend itself.
+ # Expand then shrink so the pool returns to its starting size.
+ for my $dir ([ 'expand', 24 ], [ 'shrink', $initial_nbuffers ])
+ {
+ my ($name, $target) = @$dir;
+ test_interrupt_resize_at_injection_point($injection_point,
+ $target, 'error', "error-in-handler resize-backend $name");
+ }
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "error-in-handler: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test fault tolerance of resize operation waiting for barrier
+#
+# Test that, when interrupted, a resizing operation waiting for a barrier to be
+# acknowledged doesn't leave the buffer pool in an inconsistent state.
+# =============================================================================
+
+# Workhorse function:
+#
+# We start a peer session with the given injection point in the barrier handler
+# code attached locally. Once the resize operation starts, the peer session will
+# hit the injection point and wait there. Interrupt the resize session and check
+# that the buffer pool is left in a consistent state as an aftermath.
+#
+# - injection_point: the injection point to attach to the peer session.
+# - target_nbuffers: the target buffer size for the resize operation.
+# - mode: the mode of interruption to apply to the resize session.
+# - test_label: a label to create unique test names for different tests
+sub test_fault_resize_waiting_barrier
+{
+ my ($injection_point, $target_nbuffers, $mode, $test_label) = @_;
+
+ my $orig_nbuffers = $node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+ my $log_offset = -s $node->logfile;
+
+ my $peer_session =
+ start_peer_session_with_injection_point($injection_point, 'wait');
+ my $peer_pid = $peer_session->{backend_pid};
+
+ # Sentinel session to detect a postmaster restart.
+ my $sentinel_session =
+ $node->background_psql('postgres', on_error_stop => 0);
+
+ my $resize_session = start_resize_session($target_nbuffers, $mode);
+ my $resize_pid = $resize_session->{backend_pid};
+
+ # Wait for the peer to reach the injection point. At this point the resize
+ # backend should be blocked in WaitForProcSignalBarrier.
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = '$injection_point' FROM pg_stat_activity WHERE pid = $peer_pid"
+ )
+ or die
+ "$test_label: timed out waiting for peer $peer_pid at $injection_point";
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT wait_event FROM pg_stat_activity WHERE pid = $resize_pid"
+ ),
+ 'ProcSignalBarrier',
+ "$test_label: resize $resize_pid is waiting at ProcSignalBarrier");
+
+ interrupt_resize_session($mode, $resize_session);
+
+ check_interrupted_resize($sentinel_session, $resize_session, $log_offset,
+ $mode, $injection_point, $orig_nbuffers, $target_nbuffers,
+ $test_label);
+
+ # Cleanup peer session. If the postmaster restarted all backends, the peer
+ # backend is already gone.
+ if ($node->safe_psql('postgres',
+ "SELECT count(*) = 1 FROM pg_stat_activity WHERE pid = $peer_pid")
+ eq 't')
+ {
+ $peer_session->quit;
+ }
+ else
+ {
+ $peer_session->finish;
+ }
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points passing it the
+# given mode of interruption.
+sub test_fault_resize_waiting_barrier_for_mode
+{
+ my ($mode) = @_;
+
+ my @injection_points = (
+ 'pgrsb-handle-new-buffer-alloc-barrier',
+ 'pgrsb-handle-buffer-pool-size-barrier',
+ 'pgrsb-handle-buffer-pool-resize-barrier',);
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = ([ 'expand', 24 ], [ 'shrink', $initial_nbuffers ]);
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "fault-resize-on-peer $mode: buffer pool size is $initial_nbuffers at start"
+ );
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_fault_resize_waiting_barrier($point, $target, $mode,
+ "fault-resize-on-peer $mode $name: $point");
+ }
+ }
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "fault-resize-on-peer $mode: buffer pool size is $initial_nbuffers at end"
+ );
+}
+
+# =============================================================================
+# Functions to test server restart during a resize
+#
+# Verify that a server can be stopped and started while a resize operation is in
+# progress and the server is started with buffer pool in a consistent state that
+# reflects the target size.
+# =============================================================================
+
+# Workhorse function for fast/immediate shutdown:
+#
+# Make the resize session wait at the given injection point and restart the
+# server in the given mode.
+sub test_server_restart_during_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $stop_mode, $test_label) = @_;
+
+ my $resize_session = start_resize_session($target_nbuffers,
+ 'server_restart', $injection_point);
+
+ $node->stop($stop_mode);
+
+ # Cleanup resize session, the backend must have gone now.
+ $resize_session->finish;
+
+ $node->start;
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after $stop_mode restart"
+ );
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after $stop_mode restart"
+ );
+}
+
+# Workhorse function for smart shutdown:
+#
+# Let the resize operation wait at the given injection point, send a smart
+# shutdown asynchronously. Once the postmaster enters smart shutdown, wakeup the
+# resize backend and let it complete. Verify that the pool reflects the target
+# size immediately and also after the restart.
+#
+# - injection_point: the injection point to park the resize backend at.
+# - target_nbuffers: the target buffer size for the resize operation.
+# - test_label: a label to create unique test names for different tests
+sub test_server_restart_smart_during_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $log_offset = -s $node->logfile;
+
+ my $resize_session =
+ start_resize_session($target_nbuffers, 'server_restart',
+ $injection_point);
+ my $resize_pid = $resize_session->{backend_pid};
+
+ # Open another session which can be used to wakeup the resize backend.
+ my $control_session =
+ $node->background_psql('postgres', on_error_stop => 0);
+
+ # Start the process to stop the server in smart mode.
+ local %ENV = $node->_get_env();
+ my @stop_cmd = (
+ 'pg_ctl',
+ '--pgdata' => $node->data_dir,
+ '--mode' => 'smart',
+ 'stop');
+ my ($stop_in, $stop_out, $stop_err) = ('', '', '');
+ my $stop_session =
+ IPC::Run::start(\@stop_cmd, \$stop_in, \$stop_out, \$stop_err);
+
+ # Confirm the postmaster entered smart shutdown.
+ $node->wait_for_log(qr/received smart shutdown request/, $log_offset);
+
+ # Make sure that the resize backend is still alive
+ is( $control_session->query(
+ "SELECT wait_event FROM pg_stat_activity WHERE pid = $resize_pid"
+ ),
+ $injection_point,
+ "$test_label: resize backend alive during smart shutdown");
+
+ # Wake up the resize backnd and let it finish.
+ $control_session->query_safe(
+ "SELECT injection_points_wakeup('$injection_point')",
+ verbose => 0);
+
+ # Check that the resize finished successfully by querying from the same
+ # session. The queries won't return if the resize didn't finish. Accomodate
+ # the output 't' from pg_resize_shared_buffers() in the expected output of
+ # the first query.
+ is( $resize_session->query(
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()",
+ verbose => 0),
+ "t\n$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: pg_resize_shared_buffers() succeeded and pool at target during smart shutdown"
+ );
+ is( $resize_session->query(
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'",
+ verbose => 0),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size during smart shutdown");
+
+ $resize_session->quit;
+ $control_session->quit;
+
+ # Wait for server to stop
+ IPC::Run::finish($stop_session)
+ or die "$test_label: pg_ctl smart stop failed: $stop_err";
+
+ # Sync Cluster.pm internal state and start the cluster back.
+ $node->{_pid} = undef;
+ $node->start;
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after smart restart");
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after smart restart");
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for every injection point in resize operation in
+# both directions for the given stop mode.
+sub test_server_restart_during_resize
+{
+ my ($stop_mode) = @_;
+
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'pgrsb-buffer-pool-resize-barrier-sent',);
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = ([ 'expand', 24 ], [ 'shrink', $initial_nbuffers ]);
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "server-restart $stop_mode: buffer pool size is $initial_nbuffers at start"
+ );
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+ my $label = "server-restart $stop_mode $name: $point";
+
+ if ($stop_mode eq 'smart')
+ {
+ test_server_restart_smart_during_resize_at_injection_point(
+ $point, $target, $label);
+ }
+ else
+ {
+ test_server_restart_during_resize_at_injection_point($point,
+ $target, $stop_mode, $label);
+ }
+ }
+ }
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "server-restart $stop_mode: buffer pool size is $initial_nbuffers at end"
+ );
+}
+
+# =============================================================================
+# Run tests
+# =============================================================================
+test_concurrent_resize();
+test_error_in_barrier_handler();
+
+test_interrupt_resize_session('terminate');
+test_interrupt_resize_session('cancel');
+test_interrupt_resize_session('timeout');
+test_interrupt_resize_session('error');
+
+# A resize session waiting for a barrier to be acknowledged can not be
+# interrupted by an error. Hence don't test that mode.
+test_fault_resize_waiting_barrier_for_mode('terminate');
+test_fault_resize_waiting_barrier_for_mode('cancel');
+test_fault_resize_waiting_barrier_for_mode('timeout');
+
+test_server_restart_during_resize('immediate');
+test_server_restart_during_resize('fast');
+test_server_restart_during_resize('smart');
+
+# client_connection_check_interval is only effective on systems that expose
+# POLLRDHUP/EPOLLRDHUP (Linux, and a few other Unix variants). On other
+# platforms the GUC is silently a no-op, so the disconnect test would hang.
+if ($Config::Config{osname} eq 'linux')
+{
+ test_interrupt_resize_session('disconnect');
+ test_fault_resize_waiting_barrier_for_mode('disconnect');
+}
+else
+{
+ diag( "skipping disconnect interrupt test on $Config::Config{osname} "
+ . "(requires POLLRDHUP support)");
+}
+
+done_testing();
+
+# Few more tests to add but may be somewhere else
+# TODO: test that a non-superuser cannot run pg_resize_shared_buffers()
diff --git a/src/test/buffermgr/t/002_client_join_buffer_resize.pl b/src/test/buffermgr/t/002_client_join_buffer_resize.pl
new file mode 100644
index 00000000000..219c53d59dd
--- /dev/null
+++ b/src/test/buffermgr/t/002_client_join_buffer_resize.pl
@@ -0,0 +1,286 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test shared_buffer resizing coordination with client connections joining using injection points
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+use Time::HiRes qw(sleep);
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Function to calculate the size of test table required to fill up maximum
+# buffer pool when populating it.
+sub calculate_test_sizes
+{
+ my ($node, $block_size) = @_;
+
+ # Get the maximum buffer pool size from configuration
+ my $max_shared_buffers =
+ $node->safe_psql('postgres', "SHOW max_shared_buffers");
+ my ($max_val, $max_unit) = ($max_shared_buffers =~ /(\d+)(\w+)/);
+ my $max_size_bytes;
+ if (lc($max_unit) eq 'kb')
+ {
+ $max_size_bytes = $max_val * 1024;
+ }
+ elsif (lc($max_unit) eq 'mb')
+ {
+ $max_size_bytes = $max_val * 1024 * 1024;
+ }
+ elsif (lc($max_unit) eq 'gb')
+ {
+ $max_size_bytes = $max_val * 1024 * 1024 * 1024;
+ }
+ else
+ {
+ # Default to kB if unit is not recognized
+ $max_size_bytes = $max_val * 1024;
+ }
+
+ # Fill more pages than minimally required to increase the chances of pages
+ # from the test table filling the buffer cache.
+ $max_size_bytes = $max_size_bytes;
+ my $pages_needed = int($max_size_bytes / $block_size) +
+ 10; # Add some extra to ensure buffers are filled
+ my $rows_to_insert = $pages_needed *
+ 100; # Assuming roughly 100 rows per page for our table structure
+ return ($max_size_bytes, $pages_needed, $rows_to_insert);
+}
+
+# Function to calculate expected buffer count from size string
+sub calculate_buffer_count
+{
+ my ($size_string, $block_size) = @_;
+ # Parse size and convert to bytes
+ my ($size_val, $unit) = ($size_string =~ /(\d+)(\w+)/);
+ my $size_bytes;
+ if (lc($unit) eq 'kb')
+ {
+ $size_bytes = $size_val * 1024;
+ }
+ elsif (lc($unit) eq 'mb')
+ {
+ $size_bytes = $size_val * 1024 * 1024;
+ }
+ elsif (lc($unit) eq 'gb')
+ {
+ $size_bytes = $size_val * 1024 * 1024 * 1024;
+ }
+ else
+ {
+ # Default to kB if unit is not recognized
+ $size_bytes = $size_val * 1024;
+ }
+ return int($size_bytes / $block_size);
+}
+
+# Initialize cluster with very small buffer sizes for testing
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Configure for buffer resizing with very small buffer pool sizes for faster tests.
+# TODO: for some reason parallel workers try to load default number of shared_buffers which doesn't work with lower max_shared_buffers. We need to fix that - somewhere it's picking default value of shared buffers. For now disable parallelism
+$node->append_conf('postgresql.conf',
+ 'shared_preload_libraries = injection_points');
+$node->append_conf(
+ 'postgresql.conf', qq{
+max_shared_buffers = 512kB
+shared_buffers = 320kB
+max_parallel_workers_per_gather = 0
+});
+$node->start;
+
+# Bail out if this build does not support resizable shared memory, which
+# also means that resizing buffer pool is not supported.
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
+# Enable injection points
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+
+# Get the block size (this is fixed for the binary)
+my $block_size = $node->safe_psql('postgres', "SHOW block_size");
+
+# Try to create pg_buffercache extension for buffer analysis
+
+# Create a small test table, and fetch its properties for later reference if required.
+$node->safe_psql(
+ 'postgres', qq{
+ CREATE TABLE client_test (c1 int, data char(50));
+});
+my $table_oid = $node->safe_psql('postgres',
+ "SELECT oid FROM pg_class WHERE relname = 'client_test'");
+my $table_relfilenode = $node->safe_psql('postgres',
+ "SELECT relfilenode FROM pg_class WHERE relname = 'client_test'");
+note(
+ "Test table client_test: OID = $table_oid, relfilenode = $table_relfilenode"
+);
+my ($max_size_bytes, $pages_needed, $rows_to_insert) =
+ calculate_test_sizes($node, $block_size);
+
+# Create dedicated sessions for injection point handling and test queries,
+# so that we don't create new backends for test operations after starting
+# resize operation. Only one backend, which tests new backend synchronization
+# with resizing operation, should start after resizing has commenced.
+my $injection_session = $node->background_psql('postgres');
+my $query_session = $node->background_psql('postgres');
+my $resize_session = $node->background_psql('postgres');
+
+# Function to run a single injection point test
+sub run_injection_point_test
+{
+ my ($test_name, $injection_point, $target_size, $operation_type) = @_;
+
+ # Silence the logging of the statements we run to avoid
+ # unnecessarily bloating the test logs. This runs before the
+ # upgrade we're testing, so the details should not be very
+ # interesting for debugging. But if needed, you can make it more
+ # verbose by setting this.
+ my $verbose = 0;
+
+ note("Test with $test_name ($operation_type)");
+
+ # Calculate test parameters before starting resize
+ my ($max_size_bytes, $pages_needed, $rows_to_insert) =
+ calculate_test_sizes($node, $target_size, $block_size);
+
+ # Update buffer pool size and wait for it to reflect pending state
+ $resize_session->query_safe(
+ "ALTER SYSTEM SET shared_buffers = '$target_size'",
+ verbose => $verbose);
+ $resize_session->query_safe("SELECT pg_reload_conf()",
+ verbose => $verbose);
+ my $pending_size_str = "pending: $target_size";
+ $resize_session->poll_query_until(
+ "SELECT substring(current_setting('shared_buffers'), '$pending_size_str')",
+ $pending_size_str,
+ verbose => $verbose);
+
+ # Set up injection point in injection session
+ $injection_session->query_safe(
+ "SELECT injection_points_attach('$injection_point', 'wait')",
+ verbose => $verbose);
+
+ # Trigger resize
+ $resize_session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+ );
+
+ # Wait until resize actually reaches the injection point using the query session
+ $query_session->wait_for_event(
+ 'client backend',
+ $injection_point,
+ verbose => $verbose);
+
+ # Start a client while resize is paused
+ my $client = $node->background_psql('postgres');
+ note("Background client backend PID: "
+ . $client->query_safe("SELECT pg_backend_pid()",
+ verbose => $verbose));
+
+ # Wake up the injection point from injection session
+ $injection_session->query_safe(
+ "SELECT injection_points_wakeup('$injection_point')",
+ verbose => $verbose);
+
+ # Test buffer functionality immediately after waking up injection point
+ # Insert data to test buffer pool functionality during/after resize
+ $client->query_safe(
+ "INSERT INTO client_test SELECT i, 'test_data_' || i FROM generate_series(1, $rows_to_insert) i",
+ verbose => $verbose);
+ # Verify the data was inserted correctly and can be read back
+ is( $client->query_safe(
+ "SELECT COUNT(*) FROM client_test",
+ verbose => $verbose),
+ $rows_to_insert,
+ "inserted $rows_to_insert during $test_name ($operation_type) successful"
+ );
+
+ # Verify table size is reasonable (should be substantial for testing)
+ ok( $query_session->query_safe(
+ "SELECT pg_total_relation_size('client_test')",
+ verbose => $verbose) >= $max_size_bytes,
+ "table size is large enough to overflow buffer pool in test $test_name ($operation_type)"
+ );
+
+ # Wait for the resize operation to complete. There is no direct way to do so
+ # in background_psql. Hence fire a psql command and wait for it to finish
+ $resize_session->query(q(\echo 'done'), verbose => $verbose);
+
+ # Detach injection point from injection session
+ $injection_session->query_safe(
+ "SELECT injection_points_detach('$injection_point')",
+ verbose => $verbose);
+
+ # Verify resize completed successfully
+ is( $query_session->query_safe(
+ "SELECT current_setting('shared_buffers')",
+ verbose => $verbose),
+ $target_size,
+ "resize completed successfully to $target_size");
+
+ # Check buffer pool size using pg_buffercache after resize completion
+ is( $query_session->query_safe(
+ "SELECT COUNT(*) FROM pg_buffercache",
+ verbose => $verbose),
+ calculate_buffer_count($target_size, $block_size),
+ "all buffers in the buffer pool used in $test_name ($operation_type)"
+ );
+
+ # Wait for client to complete
+ ok($client->quit, "client succeeded during $test_name ($operation_type)");
+
+ # Clean up for next test
+ $query_session->query_safe("DELETE FROM client_test",
+ verbose => $verbose);
+}
+
+# Test new client joining during various phases of buffer resizing operation using injection points
+my @injection_tests = (
+ {
+ name => 'flag setting phase',
+ injection_point => 'pg-resize-shared-buffers-flag-set',
+ },
+ {
+ name => 'new buffer alloc barrier complete',
+ injection_point => 'pgrsb-new-buffer-alloc-barrier-sent',
+ },
+ {
+ name => 'buffer pool size barrier complete',
+ injection_point => 'pgrsb-buffer-pool-size-barrier-sent',
+ },
+ {
+ name => 'buffer pool resize barrier complete',
+ injection_point => 'pgrsb-buffer-pool-resize-barrier-sent',
+ },);
+
+foreach my $test (@injection_tests)
+{
+ # Test shrinking scenario
+ run_injection_point_test($test->{name}, $test->{injection_point},
+ '272kB', 'shrinking');
+
+ # Test expanding scenario
+ run_injection_point_test($test->{name}, $test->{injection_point},
+ '400kB', 'expanding');
+}
+
+$injection_session->quit;
+$query_session->quit;
+$resize_session->quit;
+
+done_testing();
diff --git a/src/test/buffermgr/t/003_resize_failures.pl b/src/test/buffermgr/t/003_resize_failures.pl
new file mode 100644
index 00000000000..9b92df66a59
--- /dev/null
+++ b/src/test/buffermgr/t/003_resize_failures.pl
@@ -0,0 +1,202 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Test that pg_resize_shared_buffers() rolls back cleanly when resize fails.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $have_injection_points = ($ENV{enable_injection_points} eq 'yes');
+
+# Start the pool large enough that there is room to shrink below a pinned
+# buffer while still satisfying the shared_buffers GUC minimum.
+my $initial_nbuffers = 24;
+my $max_nbuffers = 32;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+if ($have_injection_points)
+{
+ $node->append_conf('postgresql.conf',
+ 'shared_preload_libraries = injection_points');
+}
+$node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
+$node->append_conf('postgresql.conf', "max_shared_buffers = $max_nbuffers");
+$node->start;
+
+# Bail out if this build does not support resizable shared memory, which
+# also means that resizing buffer pool is not supported.
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
+# pg_buffercache lets us locate the bufferid holding a given page.
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+if ($have_injection_points)
+{
+ $node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+}
+
+# ---------------------------------------------------------------------------
+# Test the case when shrinking is aborted by a pinned buffer
+# ---------------------------------------------------------------------------
+
+my $min_nbuffers = $node->safe_psql('postgres',
+ "SELECT min_val::int FROM pg_settings WHERE name = 'shared_buffers'");
+
+# In order to reliably pin a buffer above $min_nbuffers, we create as many
+# tables $min_nbuffers + 1, open a cursor on the tables and fetch one row from
+# each cursor one at a time. This will pin one buffer per table, guaranteeing
+# that at least one of the pinned buffers will be above $min_nbuffers.
+my $ntables = $min_nbuffers + 1;
+for my $i (1 .. $ntables)
+{
+ $node->safe_psql('postgres',
+ "CREATE TABLE evict_target_$i AS SELECT generate_series(1, 2) AS i");
+}
+my $pinner = $node->background_psql('postgres', on_error_stop => 0);
+$pinner->query_safe("BEGIN", verbose => 0);
+my $pinned_buf = 0;
+for my $i (1 .. $ntables)
+{
+ $pinner->query_safe(
+ "DECLARE c_$i CURSOR FOR SELECT * FROM evict_target_$i",
+ verbose => 0);
+ $pinner->query_safe("FETCH 1 FROM c_$i", verbose => 0);
+
+ my $buf = $node->safe_psql('postgres',
+ "SELECT min(bufferid) FROM pg_buffercache WHERE pinning_backends > 0 AND bufferid > $min_nbuffers"
+ );
+
+ if ($buf =~ /^\d+$/)
+ {
+ $pinned_buf = $buf;
+ last;
+ }
+}
+cmp_ok($pinned_buf, '>', $min_nbuffers,
+ "pinned a buffer above $min_nbuffers");
+
+# Set the target so that the pinned buffer is in the range of buffers to be evicted.
+my $shrink_target = $pinned_buf - 1;
+$node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$shrink_target'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+my $log_offset = -s $node->logfile;
+
+is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 'f', "shrink returns false when a buffer to be evicted is pinned");
+ok( $node->log_contains(
+ qr/could not remove buffer $pinned_buf, it is pinned/, $log_offset),
+ "log reports the pinned buffer that blocked eviction");
+ok( $node->log_contains(
+ qr/failed to evict extra buffers during shrinking/, $log_offset),
+ "log reports the eviction failure");
+
+is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$initial_nbuffers|$initial_nbuffers|$initial_nbuffers|0",
+ "pool unchanged after eviction failure");
+
+is( $node->safe_psql(
+ 'postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$initial_nbuffers (pending: $shrink_target)",
+ "pg_settings reports pending shrink target");
+
+# Releasing all pins lets the retry succeed.
+$pinner->quit;
+
+is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't', "shrink succeeds after pins released");
+
+is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$shrink_target|$shrink_target|$shrink_target|0",
+ "pool shrunk to $shrink_target after pins released");
+
+# ---------------------------------------------------------------------------
+# Test the case when memory allocation fails when expanding the buffer pool.
+# Uses an injection point to simulate the failure without exhausting real
+# memory.
+# ---------------------------------------------------------------------------
+
+SKIP:
+{
+ skip "injection points not supported by this build"
+ unless $have_injection_points;
+
+ # The buffer manager's resizable structures whose sizes must be rolled
+ # back if any one of them fails to grow.
+ my $resizable_structs =
+ q{('Buffer Descriptors', 'Buffer Blocks', 'Buffer IO Condition Variables', 'Checkpoint BufferIds')};
+ my $sizes_query =
+ "SELECT name, size FROM pg_shmem_allocations WHERE name IN $resizable_structs ORDER BY name";
+ my $sizes_before = $node->safe_psql('postgres', $sizes_query);
+
+ my $resizer = $node->background_psql('postgres');
+ $resizer->query_safe("SELECT injection_points_set_local()", verbose => 0);
+ $resizer->query_safe(
+ "SELECT injection_points_attach('buffer-mgr-resize-struct-fail', 'notice')",
+ verbose => 0);
+
+ my $expand_target = $max_nbuffers;
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$expand_target'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ my $expand_log_offset = -s $node->logfile;
+
+ is($resizer->query("SELECT pg_resize_shared_buffers()"),
+ 'f', "expansion fails when a structure can not be expanded");
+
+ # Discard the expected WARNINGs so later query_safe calls do not die.
+ $resizer->{stderr} = '';
+
+ ok( $node->log_contains(
+ qr/failed to expand buffer pool structures/,
+ $expand_log_offset),
+ "log reports the expansion failure");
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$shrink_target|$shrink_target|$shrink_target|0",
+ "buffer pool status after expansion failure");
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$shrink_target (pending: $expand_target)",
+ "pg_settings reports pending expand target");
+
+ is($node->safe_psql('postgres', $sizes_query),
+ $sizes_before,
+ "resizable buffer manager structures rolled back to previous sizes");
+
+ # Detach the injection point, to retry again. The retry should succeed.
+ $resizer->query_safe(
+ "SELECT injection_points_detach('buffer-mgr-resize-struct-fail')",
+ verbose => 0);
+ is($resizer->query("SELECT pg_resize_shared_buffers()"),
+ 't', "expand succeeds after the injection point is detached");
+
+ $resizer->quit;
+
+ is( $node->safe_psql(
+ 'postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"
+ ),
+ "$expand_target|$expand_target|$expand_target|0",
+ "pool expanded to $expand_target after detach");
+}
+
+done_testing();
diff --git a/src/test/buffermgr/t/004_resize_with_syslogger.pl b/src/test/buffermgr/t/004_resize_with_syslogger.pl
new file mode 100644
index 00000000000..5533931bdd8
--- /dev/null
+++ b/src/test/buffermgr/t/004_resize_with_syslogger.pl
@@ -0,0 +1,68 @@
+# Copyright (c) 2026-2026, PostgreSQL Global Development Group
+#
+# Test that pg_resize_shared_buffers() works when a backend that never
+# attaches to shared memory is running.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $initial_nbuffers = 16;
+my $expanded_nbuffers = 24;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# When logging_collector is on, the server starts a syslogger process that never
+# attaches to the shared memory. We use that as a proxy for a backend that never
+# attaches to shared memory.
+$node->append_conf(
+ 'postgresql.conf', qq{
+shared_buffers = $initial_nbuffers
+max_shared_buffers = $expanded_nbuffers
+logging_collector = on
+});
+$node->start;
+
+# Bail out if this build does not support resizable shared memory, which
+# also means that resizing buffer pool is not supported.
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
+# Check that the syslogger is running by writing a log marker and waiting for it
+# to appear in the log file.
+sub check_syslogger_running
+{
+ my ($marker) = @_;
+
+ $node->safe_psql('postgres',
+ "DO \$\$ BEGIN RAISE LOG '$marker'; END \$\$");
+ return $node->poll_query_until('postgres',
+ "SELECT pg_read_file(pg_current_logfile()) ~ '$marker'");
+}
+
+check_syslogger_running('syslogger_marker_before_resize')
+ or die "syslogger is not running";
+
+# Resize the buffer pool, and check that the syslogger continues to run while
+# the resize is in progress.
+# TODO: Instead of custom markers we could use the log line that is emitted when
+# the resize is complete, when we have frozen those.
+for
+ my $dir ([ 'expand', $expanded_nbuffers ], [ 'shrink', $initial_nbuffers ])
+{
+ my ($name, $target) = @$dir;
+
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$target'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't', "$name to $target succeeds with syslogger running");
+ ok(check_syslogger_running("syslogger_marker_after_$name"),
+ "syslogger drains logs after $name");
+}
+
+done_testing();
diff --git a/src/test/buffermgr/t/005_resize_unsupported.pl b/src/test/buffermgr/t/005_resize_unsupported.pl
new file mode 100644
index 00000000000..1ee90c1cafd
--- /dev/null
+++ b/src/test/buffermgr/t/005_resize_unsupported.pl
@@ -0,0 +1,50 @@
+# Copyright (c) 2026-2026, PostgreSQL Global Development Group
+#
+# Test that pg_resize_shared_buffers() errors out when resizable shared
+# memory is not supported.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $initial_nbuffers = 256;
+my $max_nbuffers = 512;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf(
+ 'postgresql.conf', qq{
+shared_buffers = $initial_nbuffers
+max_shared_buffers = $max_nbuffers
+});
+$node->start;
+
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') eq 'on')
+{
+ # The builds that support resizable shared memory, usually, will not support
+ # the feature when SysV shared memory is used.
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_memory_type = 'sysv'");
+ $node->restart;
+}
+
+is($node->safe_psql('postgres', 'SHOW have_resizable_shmem'),
+ 'off', 'have_resizable_shmem reports off');
+
+my $target_nbuffers = $initial_nbuffers / 2;
+$node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$target_nbuffers'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', "SELECT pg_resize_shared_buffers()");
+isnt($ret, 0,
+ 'pg_resize_shared_buffers fails when resizable shared memory is unsupported'
+);
+like(
+ $stderr,
+ qr/resizing shared buffer pool is not supported on this platform/,
+ 'error message reports that resizing shared buffer pool is unsupported');
+
+done_testing();
diff --git a/src/test/buffermgr/t/010_stress_resize_buffer.pl b/src/test/buffermgr/t/010_stress_resize_buffer.pl
new file mode 100644
index 00000000000..70b33f83e97
--- /dev/null
+++ b/src/test/buffermgr/t/010_stress_resize_buffer.pl
@@ -0,0 +1,34 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Minimal stress test: resize shared_buffers repeatedly against regular pgbench
+# workload.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios.
+my @buffer_sizes =
+ (128, 28, 16 * 1024, 32 * 1024, 1024, 512, 16, 24, 256, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_buffer_resize_test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+
+$stress->setup;
+
+$stress->run;
+
+done_testing();
diff --git a/src/test/buffermgr/t/011_stress_drop_relation_buffers.pl b/src/test/buffermgr/t/011_stress_drop_relation_buffers.pl
new file mode 100644
index 00000000000..a3053a87d69
--- /dev/null
+++ b/src/test/buffermgr/t/011_stress_drop_relation_buffers.pl
@@ -0,0 +1,99 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Stress test execution of DropRelationBuffers(), DropRelationsAllBuffers() and
+# FlushRelationsAllBuffers() concurrently with shared_buffers resizing.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use List::Util qw(max);
+use PostgreSQL::Test::Utils;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios. At any time during the run buffer pool should be large enough to
+# let a new backend join while other backends are performing COPY, which seems
+# to pin many buffers at a time; avoid too small sizes.
+my @buffer_sizes = (512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+# The injection points verify that the buffer pool scan is exercised as expected
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_drop_relation_buffers_test',
+ injection_points => [
+ 'drop-relation-buffers-scan',
+ 'drop-relations-all-buffers-scan',
+ 'flush-relations-all-buffers-scan',
+ ],
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+$stress->setup;
+
+# Force execution of FlushRelationsAllBuffers() by skipping WAL logging DMLs to
+# a newly created table.
+my $node = $stress->node;
+$node->append_conf(
+ 'postgresql.conf', qq{
+wal_level = minimal
+max_wal_senders = 0
+wal_skip_threshold = 0
+});
+$node->restart;
+
+# Test specific load preparation.
+#
+# DropRelationBuffers() and DropRelationsAllBuffers() scan the buffer pool
+# only when the size of the relation exceeds NBuffers/32. Create a
+# relation larger than max(@buffer_sizes)/32 so a scan always runs. Dump
+# it once so pgbench clients can COPY it back in each iteration instead of
+# regenerating the data.
+my $tempdir = PostgreSQL::Test::Utils::tempdir;
+my $refdata_path = "$tempdir/refdata.bin";
+$node->safe_psql(
+ 'postgres', qq{
+ CREATE UNLOGGED TABLE refdata_source AS
+ SELECT g AS a, repeat('x', 1900)::bytea AS b
+ FROM generate_series(1, 16800) g;
+});
+$node->safe_psql('postgres',
+ "COPY refdata_source TO '$refdata_path' WITH (FORMAT binary)");
+my $max_nbuffers = max @buffer_sizes;
+my $required_pages = int($max_nbuffers / 32) + 1;
+my $refdata_pages = $node->safe_psql('postgres',
+ "SELECT (pg_relation_size('refdata_source') / current_setting('block_size')::bigint)::int"
+);
+cmp_ok($refdata_pages, '>', $required_pages,
+ "refdata_source spans more than NBuffers/32 pages at the largest tested pool size"
+);
+
+# Workload script fed to pgbench.
+#
+# Each client picks table names from a disjoint numeric range so table
+# names from concurrent clients do not collide. The :pgbench_id prefix
+# further disambiguates across the two pgbench flavors (persistent /
+# per_transaction) that StressUtil runs in parallel.
+my $workload_sql = qq{
+\\set tid :pgbench_id * 1000000000 + :client_id * 100000000 + random(1, 100000000)
+BEGIN;
+CREATE TABLE t_:tid (a int, b bytea);
+COPY t_:tid FROM '$refdata_path' WITH (FORMAT binary);
+TRUNCATE t_:tid; -- hit DropRelationBuffers()
+COPY t_:tid FROM '$refdata_path' WITH (FORMAT binary);
+COMMIT; -- hit FlushRelationsAllBuffers()
+DROP TABLE t_:tid; -- hit DropRelationsAllBuffers()
+};
+
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/012_stress_drop_database_buffers.pl b/src/test/buffermgr/t/012_stress_drop_database_buffers.pl
new file mode 100644
index 00000000000..77afb1f950d
--- /dev/null
+++ b/src/test/buffermgr/t/012_stress_drop_database_buffers.pl
@@ -0,0 +1,82 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Exercise the buffer pool scan that drops buffers belonging to a given
+# database (DropDatabaseBuffers()) concurrently with shared_buffers
+# resizing.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use List::Util qw(max);
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety of
+# scenarios. Avoid very small sizes because concurrent CREATE DATABASE clones
+# can pin more buffers than a very small pool provides, which can cause
+# unrelated failures (in particular, per-transaction pgbench connections
+# get "no unpinned buffers available" when opening a fresh backend).
+my @buffer_sizes = (256, 512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_drop_database_buffers_test',
+ injection_points => ['drop-database-buffers-scan'],
+
+ # Concurrent CREATE DATABASE clones pin many buffers per client; raising
+ # this can exhaust the small pool sizes resulting in the
+ # pg_resize_shared_buffers() query failing with and fail with error "no
+ # unpinned buffers available".
+ pgbench_clients => 4,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+
+$stress->setup;
+
+# Populate a template database with enough data. Size the seed table to roughly
+# 1/8 of the largest buffer pool we exercise. At the largest pool it still
+# occupies ~12% so that DropDatabaseBuffers() finds enough fraction of buffers
+# to drop. At smaller pool sizes the template exceeds the pool, thus covering
+# all the combinations of scanning buffer pool and dropping buffers.
+#
+# The 1900-byte payload makes sure that each row remains in the heap.
+my $node = $stress->node;
+my $seed_template = 'seedtemplate';
+my $max_buffer_pool = max @buffer_sizes;
+my $seed_row_count = 4 * $max_buffer_pool / 8;
+$node->safe_psql('postgres', "CREATE DATABASE $seed_template");
+$node->safe_psql(
+ $seed_template, qq{
+ CREATE TABLE seedtab (a int, b bytea);
+ INSERT INTO seedtab
+ SELECT g, repeat('x', 1900)::bytea
+ FROM generate_series(1, $seed_row_count) g;
+});
+$node->safe_psql('postgres',
+ "UPDATE pg_database SET datistemplate = true, datallowconn = false "
+ . "WHERE datname = '$seed_template'");
+
+# Workload script fed to pgbench.
+#
+# Each client picks database names from a disjoint numeric range so
+# database names from concurrent clients do not collide. The :pgbench_id
+# prefix further disambiguates across the two pgbench flavors
+# (persistent / per_transaction) that StressUtil runs in parallel.
+my $workload_sql = qq{
+\\set tid :pgbench_id * 1000000000 + :client_id * 100000000 + random(1, 100000000)
+CREATE DATABASE d_:tid TEMPLATE $seed_template;
+DROP DATABASE d_:tid;
+};
+
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/013_stress_checkpoint.pl b/src/test/buffermgr/t/013_stress_checkpoint.pl
new file mode 100644
index 00000000000..8c7d6bb5841
--- /dev/null
+++ b/src/test/buffermgr/t/013_stress_checkpoint.pl
@@ -0,0 +1,47 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Test synchronization between shared_buffers resize and CHECKPOINT. The test
+# issues explicit CHECKPOINTs so that the frequency of heckpoints can be
+# controlled so as increase the likelihood of concurrent checkpoint with every
+# resize.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios.
+my @buffer_sizes =
+ (128, 28, 16 * 1024, 32 * 1024, 1024, 512, 16, 24, 256, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_checkpoint_stress_test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+
+$stress->setup;
+
+# Disable implicit checkpoints
+$stress->node->append_conf(
+ 'postgresql.conf', qq{
+checkpoint_timeout = 1h
+max_wal_size = 100GB
+});
+$stress->node->reload;
+
+$stress->run(
+ workload_sql => "CHECKPOINT;\n",
+ workload_weight => 1,
+ default_load_weight => 10,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/014_stress_flush_relation_buffers.pl b/src/test/buffermgr/t/014_stress_flush_relation_buffers.pl
new file mode 100644
index 00000000000..52e326ef391
--- /dev/null
+++ b/src/test/buffermgr/t/014_stress_flush_relation_buffers.pl
@@ -0,0 +1,68 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Stress FlushRelationBuffers() concurrently with shared_buffers
+# resizing.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios. At any time during the run the buffer pool must be large
+# enough to let a new backend join while other backends are CLUSTERing a
+# small table, which pins several buffers at once; avoid too small sizes.
+my @buffer_sizes = (512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+# The injection point verifies that the buffer pool scan is exercised in
+# FlushRelationBuffers(). If the function stops scanning the buffer pool
+# the test is useless, so we assert that it fires at least once per
+# workload transaction.
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_flush_relation_buffers_test',
+ injection_points => ['flush-relation-buffers-scan'],
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+
+$stress->setup;
+
+my $node = $stress->node;
+my $ts2_dir = $node->basedir . '/tblsp2';
+mkdir $ts2_dir or die "mkdir $ts2_dir: $!";
+$node->safe_psql('postgres', "CREATE TABLESPACE ts2 LOCATION '$ts2_dir'");
+
+# Workload script fed to pgbench.
+#
+# Each client picks table names from a disjoint numeric range so table
+# names from concurrent clients do not collide. The :pgbench_id prefix
+# further disambiguates across the two pgbench flavors (persistent /
+# per_transaction) that StressUtil runs in parallel.
+my $workload_sql = qq{
+\\set tid :pgbench_id * 1000000000 + :client_id * 100000000 + random(1, 100000000)
+BEGIN;
+CREATE TABLE t_:tid (a int PRIMARY KEY, b bytea);
+INSERT INTO t_:tid
+ SELECT g, repeat('x', 1900)::bytea
+ FROM generate_series(1, 100) g;
+-- hit FlushRelationBuffers()
+ALTER TABLE t_:tid SET TABLESPACE ts2;
+ALTER TABLE t_:tid SET TABLESPACE pg_default;
+DROP TABLE t_:tid;
+COMMIT;
+};
+
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/015_stress_pg_buffercache.pl b/src/test/buffermgr/t/015_stress_pg_buffercache.pl
new file mode 100644
index 00000000000..7166e343590
--- /dev/null
+++ b/src/test/buffermgr/t/015_stress_pg_buffercache.pl
@@ -0,0 +1,60 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Stress the pg_buffercache diagnostic and monitoring functions
+# concurrently with shared_buffers resizing.
+#
+# Destructive functions like pg_buffercache_evict_* and
+# pg_buffercache_mark_dirty_* are excluded because they might cause failures
+# unrelated to the test.
+#
+# TODO: This test fails because pg_buffercache_os_pages_internal() expects the
+# buffer pool size to be constant. Fix is on the way.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios. At any time during the run buffer pool should be large
+# enough to let a new backend join while other backends are scanning the
+# buffer pool via pg_buffercache; avoid too small sizes.
+my @buffer_sizes = (512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_pg_buffercache_test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+$stress->setup;
+
+$stress->node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache');
+
+# Test specific workload. Include NUMA view only if the server supports NUMA.
+my $workload_sql = qq{
+SELECT count(*) FROM pg_buffercache;
+SELECT * FROM pg_buffercache_summary();
+SELECT count(*) FROM pg_buffercache_usage_counts();
+SELECT count(*) FROM pg_buffercache_os_pages;
+};
+if ($stress->node->safe_psql('postgres', 'SELECT pg_numa_available()') eq 't')
+{
+ $workload_sql .= "SELECT count(*) FROM pg_buffercache_numa;\n";
+}
+
+# Use the default workload only to populate the buffer pool, but main workload
+# is the pg_buffercache queries.
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/016_stress_pg_prewarm.pl b/src/test/buffermgr/t/016_stress_pg_prewarm.pl
new file mode 100644
index 00000000000..39b44191595
--- /dev/null
+++ b/src/test/buffermgr/t/016_stress_pg_prewarm.pl
@@ -0,0 +1,70 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Stress test the autoprewarm concurrently with shared_buffers resizing.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+my @buffer_sizes = (512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_pg_prewarm_test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+$stress->setup;
+
+my $node = $stress->node;
+
+# Configure the autoprewarm background worker to run as frequently as possible.
+$node->append_conf(
+ 'postgresql.conf', qq{
+shared_preload_libraries = 'pg_prewarm'
+pg_prewarm.autoprewarm = on
+pg_prewarm.autoprewarm_interval = 1s
+});
+$node->restart;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_prewarm');
+
+# Dump the buffer pool through additional load to increase the likelihood of it
+# happening concurrently with a resize. Wrap the call in an exception block to
+# swallow "dump file is being used by PID N" errors.
+my $workload_sql = qq{
+DO \$\$
+BEGIN
+ PERFORM autoprewarm_dump_now();
+EXCEPTION WHEN OTHERS THEN
+ NULL;
+END
+\$\$;
+};
+
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+my $dumpfile = $node->data_dir . '/autoprewarm.blocks';
+ok(-s $dumpfile, "buffer pool was dumped at least once");
+
+# Restart to confirm that the dump file can be read and the buffer pool can be
+# prewarmed.
+my $log_offset = -s $node->logfile;
+$node->restart;
+$node->wait_for_log(
+ qr/autoprewarm successfully prewarmed \d+ of \d+ previously-loaded blocks/,
+ $log_offset);
+pass("autoprewarm prewarmed shared buffers after restart");
+
+done_testing();
diff --git a/src/test/buffermgr/t/StressUtil.pm b/src/test/buffermgr/t/StressUtil.pm
new file mode 100644
index 00000000000..e804226dcfa
--- /dev/null
+++ b/src/test/buffermgr/t/StressUtil.pm
@@ -0,0 +1,628 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+
+=pod
+
+=head1 NAME
+
+StressUtil - shared driver for the buffermgr shared_buffers resize stress tests
+
+=head1 SYNOPSIS
+
+ use StressUtil;
+
+ # Configure a stress test driver
+ my $stress = StressUtil->new(
+ buffer_sizes => [128, 28, ...],
+ application_name => 'pgbench_..._test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,
+ injection_points => [...], # optional
+ );
+
+ # Setup the cluster and pgbench workload
+ $stress->setup;
+
+ # -- per-test prep goes here; may use $stress->node --
+
+ # Run the stress test with optional custom workload and perform post-stress
+ # checks
+ $stress->run(
+ default_load_weight => 1, # required if workload_sql set
+ workload_sql => $sql, # optional
+ workload_weight => 10, # required if workload_sql set
+ );
+
+=head1 DESCRIPTION
+
+StressUtil provides common routines for the shared_buffers resize stress tests.
+These are the routines for setting up the cluster and pgbench database, resizing
+shared_buffers in a tight loop while a pgbench workload runs concurrently, and
+performing post-stress checks.
+
+=cut
+
+package StressUtil;
+
+use strict;
+use warnings FATAL => 'all';
+
+use IPC::Run;
+use List::Util qw(max min shuffle);
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+=pod
+
+=head1 METHODS
+
+=over
+
+=item StressUtil->new(%opts)
+
+Construct a stress-test object which can be used to run the stress test with the
+given specifications. Named options:
+
+=over
+
+=item buffer_sizes
+
+Array reference of shared_buffers values (in number of buffers) that the resize
+loop cycles through. Required.
+
+=item application_name
+
+application_name string set on the pgbench connection; the resize loop
+polls pg_stat_activity for this value to detect when pgbench has
+exited. Required.
+
+=item injection_points
+
+Array reference of injection point names used to detect whether a code path is
+hit during stress test. They are attached with the B<notice> action. Number of
+times the notice message appears in the server error log indicates the number of
+times a certain code path is hit during the stress test. Defaults to the empty
+list. Used only when the build supports injection points.
+
+=item pgbench_clients
+
+Number of pgbench client connections. Split evenly across the persistent and
+per-transaction pgbench processes, so must be at least B<2>. May be overridden
+at run time by the C<PG_TEST_RESIZE_STRESS_CLIENTS> environment variable.
+
+=item pgbench_scale
+
+pgbench scale factor. May be overridden at run time by the
+C<PG_TEST_RESIZE_STRESS_SCALE> environment variable.
+
+=item pgbench_duration
+
+pgbench duration in seconds. May be overridden at run time by the
+C<PG_TEST_RESIZE_STRESS_SECONDS> environment variable so that the test never
+hits the timeout in a successful run.
+
+=back
+
+=cut
+
+sub new
+{
+ my ($class, %opts) = @_;
+ for my $required (
+ qw(buffer_sizes
+ application_name
+ pgbench_clients
+ pgbench_scale
+ pgbench_duration))
+ {
+ defined $opts{$required} or die "$required required";
+ }
+
+ # Env vars override the test-specified values.
+ my $duration =
+ int($ENV{PG_TEST_RESIZE_STRESS_SECONDS} // $opts{pgbench_duration});
+ my $clients =
+ int($ENV{PG_TEST_RESIZE_STRESS_CLIENTS} // $opts{pgbench_clients});
+ my $scale =
+ int($ENV{PG_TEST_RESIZE_STRESS_SCALE} // $opts{pgbench_scale});
+
+ my $timeout_cap = int(($ENV{PG_TEST_TIMEOUT_DEFAULT} // 0) * 0.8);
+ if ($timeout_cap > 0 && $duration > $timeout_cap)
+ {
+ note "clamping pgbench duration from $duration to $timeout_cap";
+ $duration = $timeout_cap;
+ }
+
+ $clients >= 2
+ or die "pgbench_clients must be at least 2 "
+ . "(split between persistent and per-transaction pgbench)";
+
+ my $self = {
+ buffer_sizes => $opts{buffer_sizes},
+ application_name => $opts{application_name},
+ pgbench_clients => $clients,
+ pgbench_scale => $scale,
+ pgbench_duration => $duration,
+ injection_points => $opts{injection_points} // [],
+ node => undef,
+ injection_points_supported => undef,
+ };
+ return bless $self, $class;
+}
+
+=pod
+
+=item $stress->node
+
+Return the underlying C<PostgreSQL::Test::Cluster> node. Valid only
+after setup().
+
+=cut
+
+sub node { return $_[0]->{node}; }
+
+=pod
+
+=item $stress->setup
+
+Create and initialize PostgreSQL cluster and other necessary objects required
+for the stress test.
+
+=cut
+
+sub setup
+{
+ my ($self) = @_;
+
+ my $node = PostgreSQL::Test::Cluster->new('main');
+ $node->init;
+ $self->{node} = $node;
+
+ my $max_buffer_pool = max @{ $self->{buffer_sizes} };
+ my $initial_buffers = min @{ $self->{buffer_sizes} };
+ my $ips_supported = ($ENV{enable_injection_points} // 'no') eq 'yes';
+ my $use_ips = $ips_supported && @{ $self->{injection_points} };
+ $self->{injection_points_supported} = $ips_supported;
+
+ $node->append_conf(
+ 'postgresql.conf', qq{
+max_shared_buffers = $max_buffer_pool
+shared_buffers = $initial_buffers
+log_statement = none
+restart_after_crash = off
+});
+
+ # Route injection-point NOTICEs to the server log, not to the pgbench
+ # client which does not expect them.
+ if ($use_ips)
+ {
+ $node->append_conf(
+ 'postgresql.conf', qq{
+shared_preload_libraries = injection_points
+log_min_messages = notice
+client_min_messages = warning
+});
+ }
+
+ $node->start;
+
+ # Bail out if this build does not support resizable shared memory, which
+ # also means that resizing buffer pool is not supported.
+ if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+ {
+ plan skip_all =>
+ "resizable shared memory not supported by this build";
+ }
+
+ $node->safe_psql('postgres', "CREATE EXTENSION buffermgr_test");
+ $node->safe_psql('postgres', "CREATE EXTENSION amcheck");
+
+ if ($use_ips)
+ {
+ $node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+ }
+
+ # Create a table to capture the history of resizes
+ $node->safe_psql(
+ 'postgres', qq{
+CREATE TABLE resize_log(
+ size int NOT NULL,
+ started_at timestamptz NOT NULL,
+ ended_at timestamptz NOT NULL,
+ num_tries int NOT NULL);
+});
+
+ # Reset the bgwriter stats so we can assert that it ran during the test.
+ $node->safe_psql('postgres', "SELECT pg_stat_reset_shared('bgwriter')");
+}
+
+=pod
+
+=item $stress->run(%opts)
+
+Run the stress test. This resizes shared_buffers repeatedly while a pgbench
+workload runs concurrently. Perform post-stress checks.
+
+Two pgbench processes run in parallel: one keeps its connections open for the
+whole run (persistent), the other reconnects for every transaction. The
+C<pgbench_clients> count is split evenly between them. Each pgbench is passed a
+distinct C<pgbench_id> variable and gets a distinct application_name so custom
+workloads can build names that do not collide across the two.
+
+pgbench always runs its built-in tpcb-like workload; an optional custom
+workload is added if requested.
+
+Named options:
+
+=over
+
+=item workload_sql
+
+Contents of a custom pgbench workload script. Optional.
+
+=item workload_weight
+
+Weight of the custom workload script relative to the built-in
+tpcb-like script. Required when C<workload_sql> is set; must not be
+set otherwise.
+
+=item default_load_weight
+
+Weight of the built-in tpcb-like workload relative to the custom
+workload script. Required when C<workload_sql> is set; must not be
+set otherwise.
+
+=back
+
+=cut
+
+sub run
+{
+ my ($self, %opts) = @_;
+ my $node = $self->{node} or die "setup() must be called first";
+ my $ips = $self->{injection_points};
+ my $use_ips = $self->{injection_points_supported} && @$ips;
+
+ my @pgbench_args;
+
+ my $workload_path;
+ if (defined $opts{workload_sql})
+ {
+ my $default_weight = $opts{default_load_weight}
+ // die "default_load_weight required when workload_sql is set";
+ my $workload_weight = $opts{workload_weight}
+ // die "workload_weight required when workload_sql is set";
+
+ push @pgbench_args, '-b', "tpcb-like\@$default_weight";
+
+ $workload_path = $node->basedir . '/workload.sql';
+ open(my $wfh, '>', $workload_path)
+ or die "cannot write $workload_path: $!";
+ print $wfh $opts{workload_sql};
+ close($wfh);
+ push @pgbench_args, '-f', "$workload_path\@$workload_weight";
+ }
+ elsif (defined $opts{default_load_weight}
+ || defined $opts{workload_weight})
+ {
+ die "default_load_weight and workload_weight require workload_sql";
+ }
+
+ # Attach injection points just before starting the workload.
+ if ($use_ips)
+ {
+ for my $ip (@$ips)
+ {
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('$ip', 'notice')");
+ }
+ }
+
+ my $log_offset = -s $node->logfile;
+
+ my @procs = _start_pgbench_workloads($self, \@pgbench_args);
+
+ _wait_for_pgbench_ready($node, $self->{application_name});
+
+ my $tests_completed = _run_resize_loop($self);
+
+ _run_post_checks($self, \@procs, $log_offset, $workload_path,
+ $tests_completed);
+}
+
+=back
+
+=cut
+
+# Resize the buffer pool and log the outcome to resize_log.
+sub _apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ # Start a new backend so that it inherits the reloaded GUC directly from the
+ # postmaster.
+ $node->safe_psql(
+ 'postgres', qq{
+ INSERT INTO resize_log(size, started_at, ended_at, num_tries)
+ SELECT $new_size, started_at, ended_at, num_tries
+ FROM pg_resize_shared_buffers_sql($new_size)
+ });
+
+ # A resize failure causes the test to time out, so reaching here means
+ # success.
+ ok(1, "buffer pool resized to $new_size");
+}
+
+# Return true if either pgbench workload is still running, false otherwise.
+#
+# IPC::Run's pumpable status is unreliable; check pg_stat_activity instead.
+sub _pgbench_processes_active
+{
+ my ($node, $application_name) = @_;
+
+ my $result = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_stat_activity "
+ . "WHERE application_name LIKE '${application_name}%'");
+ return int($result) > 0;
+}
+
+# Wait until at least one pgbench workload has registered in
+# pg_stat_activity, so the resize loop's first _pgbench_processes_active
+# check does not race pgbench startup and exit immediately.
+sub _wait_for_pgbench_ready
+{
+ my ($node, $application_name) = @_;
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) >= 1 FROM pg_stat_activity "
+ . "WHERE application_name LIKE '${application_name}%'")
+ or die "timed out waiting for pgbench workloads to connect";
+}
+
+# Initialize pgbench, then start two pgbench workloads in parallel: one
+# persistent, one with -C (per_transaction). Clients are split evenly. Each
+# pgbench gets a distinct application_name and a distinct :pgbench_id script
+# variable so custom workloads can build non-colliding names across the two.
+#
+# Returns a list of per-process hashes with keys process, stdout_ref,
+# stderr_ref, pgbench_id.
+sub _start_pgbench_workloads
+{
+ my ($self, $extra_args) = @_;
+ my $node = $self->{node};
+ my $scale = $self->{pgbench_scale};
+ my $app_name = $self->{application_name};
+ my $total_clients = $self->{pgbench_clients};
+
+ $node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$scale --quiet",
+ 0,
+ [qr{^$}],
+ [
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$scale)");
+
+ my @flavors = (
+ { pgbench_id => 1, extra => [] },
+ { pgbench_id => 2, extra => ['-C'] },);
+ my $clients_1 = int($total_clients / 2);
+ my $clients_2 = $total_clients - $clients_1;
+ $flavors[0]->{clients} = $clients_1;
+ $flavors[1]->{clients} = $clients_2;
+
+ my @procs;
+ for my $f (@flavors)
+ {
+ my $id = $f->{pgbench_id};
+ my $flavor_app = "${app_name}_${id}";
+ my ($stdin, $stdout, $stderr) = ('', '', '');
+ my $process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-h', $node->host,
+ '-T', $self->{pgbench_duration},
+ '-c', $f->{clients},
+ '-D', "pgbench_id=$id",
+ # stop on first server crash, so that conditions at the time of
+ # crash are preserved for diagnosis.
+ '--exit-on-abort',
+ '--continue-on-error',
+ @{ $f->{extra} },
+ @$extra_args,
+ "dbname=postgres application_name=$flavor_app"
+ ],
+ '<' => \$stdin,
+ '>' => \$stdout,
+ '2>' => \$stderr);
+
+ ok($process, "pgbench started successfully (pgbench_id=$id)");
+ push @procs,
+ {
+ process => $process,
+ stdout_ref => \$stdout,
+ stderr_ref => \$stderr,
+ pgbench_id => $id,
+ };
+ }
+ return @procs;
+}
+
+# Resize as many times as possible while pgbench is running, cycling
+# through $self->{buffer_sizes} in a shuffled order without ever picking
+# the same size twice in a row. Returns the number of resizes performed.
+sub _run_resize_loop
+{
+ my ($self) = @_;
+ my $node = $self->{node};
+ my $app_name = $self->{application_name};
+ my $buffer_sizes = $self->{buffer_sizes};
+ my @queue;
+ my $last_picked;
+ my $tests_completed = 0;
+
+ while (_pgbench_processes_active($node, $app_name))
+ {
+ if (!@queue)
+ {
+ @queue = shuffle(@$buffer_sizes);
+ if ( defined $last_picked
+ && @queue > 1
+ && $queue[0] == $last_picked)
+ {
+ @queue[ 0, 1 ] = @queue[ 1, 0 ];
+ }
+ }
+ $last_picked = shift @queue;
+ _apply_and_verify_buffer_change($node, $last_picked);
+ $tests_completed++;
+ }
+ return $tests_completed;
+}
+
+# Assert the resize loop ran through at least one full sequence and
+# every size in @$buffer_sizes was picked at least once.
+sub _assert_all_sizes_used
+{
+ my ($node, $buffer_sizes, $tests_completed) = @_;
+
+ cmp_ok($tests_completed, '>', scalar(@$buffer_sizes),
+ "all buffer size transitions were tested");
+ note
+ "completed $tests_completed buffer resize operations while pgbench was running";
+
+ my $ndistinct = $node->safe_psql('postgres',
+ "SELECT count(DISTINCT size) FROM resize_log");
+ is($ndistinct, scalar(@$buffer_sizes),
+ "every buffer size was exercised at least once");
+}
+
+# Make sure that the pgbench workloads have ended and perform post-stress
+# checks.
+sub _run_post_checks
+{
+ my ($self, $procs, $log_offset, $workload_path, $tests_completed) = @_;
+
+ my $node = $self->{node};
+ my $ips = $self->{injection_points};
+
+ for my $p (@$procs)
+ {
+ my $id = $p->{pgbench_id};
+ $p->{process}->signal('TERM');
+ ok((IPC::Run::finish $p->{process}),
+ "pgbench finished successfully (pgbench_id=$id)");
+ note("pgbench_id=$id stderr:\n" . ${ $p->{stderr_ref} })
+ if ${ $p->{stderr_ref} } ne '';
+ note("pgbench_id=$id stdout:\n" . ${ $p->{stdout_ref} });
+ }
+
+ _assert_all_sizes_used($node, $self->{buffer_sizes}, $tests_completed);
+
+ # Log resize latency distribution and max retry count, for post-mortem
+ note $node->safe_psql(
+ 'postgres',
+ q{SELECT format('resize stats: n=%s, min=%s, avg=%s, max=%s, max_tries=%s',
+ count(*),
+ min(ended_at - started_at),
+ avg(ended_at - started_at),
+ max(ended_at - started_at),
+ max(num_tries))
+ FROM resize_log});
+
+ # Checkpointer activity, for post-mortem.
+ note $node->safe_psql(
+ 'postgres',
+ q{SELECT format('checkpointer stats: timed=%s, requested=%s, buffers_written=%s',
+ num_timed, num_requested, buffers_written)
+ FROM pg_stat_checkpointer});
+
+ is( $node->safe_psql(
+ 'postgres', "SELECT buffers_clean > 0 FROM pg_stat_bgwriter"),
+ 't',
+ "background writer ran during resize cycle");
+
+ # Server error log is expected to be crash free
+ $node->log_check("no PANIC or SIGBUS during stress run",
+ $log_offset, log_unlike => [ qr/PANIC/, qr/signal 7/ ]);
+
+ # pg_dumpall reads every table and catalog in every database; An error free
+ # dump indicates that the database remained non-corrupt after the stress
+ # run. We are not interested in the dump output, so discard it to /dev/null.
+ $node->command_ok(
+ [ 'pg_dumpall', '--no-sync', '-f', '/dev/null' ],
+ "pg_dumpall succeeds after stress run");
+
+ # pg_dumpall does not scan indexes; run bt_index_parent_check over every
+ # btree index to catch index-level corruption.
+ $node->safe_psql(
+ 'postgres', q{
+ SELECT bt_index_parent_check(c.oid, true, true)
+ FROM pg_class c
+ JOIN pg_index i ON i.indexrelid = c.oid
+ WHERE c.relkind = 'i'
+ AND c.relam = (SELECT oid FROM pg_am WHERE amname = 'btree')
+ });
+ ok(1, "all btree indexes verified");
+
+ # Verify tables
+ my $heap_findings = $node->safe_psql(
+ 'postgres', q{
+ SELECT count(*)
+ FROM (SELECT c.oid AS rel
+ FROM pg_class c
+ WHERE c.relkind IN ('r', 'S')
+ AND c.relpersistence = 'p') r,
+ LATERAL verify_heapam(r.rel, check_toast => true) v
+ });
+ is($heap_findings, '0', "verify_heapam found no corruption");
+
+ if (@$ips)
+ {
+ SKIP:
+ {
+ skip "injection points not supported by this build"
+ unless $self->{injection_points_supported};
+
+ my $workload_txns = 0;
+ for my $p (@$procs)
+ {
+ my $n;
+ if (defined $workload_path)
+ {
+ ($n) = ${ $p->{stdout_ref} } =~
+ m{SQL script \d+:\s+\Q$workload_path\E.*?number of transactions actually processed:\s*(\d+)}s;
+ }
+ else
+ {
+ ($n) = ${ $p->{stdout_ref} } =~
+ m{number of transactions actually processed:\s*(\d+)};
+ }
+ ok( defined $n,
+ "transaction count found in pgbench stdout (pgbench_id="
+ . $p->{pgbench_id} . ")");
+ $workload_txns += $n if defined $n;
+ }
+
+ my $log_content = slurp_file($node->logfile, $log_offset);
+ for my $ip (@$ips)
+ {
+ my $count = () = $log_content =~
+ /notice triggered for injection point $ip\b/g;
+ cmp_ok($count, '>=', $workload_txns,
+ "injection point $ip fired at least $workload_txns times"
+ );
+ }
+ }
+ }
+}
+
+1;
diff --git a/src/test/meson.build b/src/test/meson.build
index cd45cbf57fb..e9550933063 100644
--- a/src/test/meson.build
+++ b/src/test/meson.build
@@ -4,6 +4,7 @@ subdir('regress')
subdir('isolation')
subdir('authentication')
+subdir('buffermgr')
subdir('postmaster')
subdir('recovery')
subdir('subscription')
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index 699334320d9..9e7b3222297 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -61,6 +61,7 @@ use Config;
use IPC::Run;
use PostgreSQL::Test::Utils qw(pump_until);
use Test::More;
+use Time::HiRes qw(usleep);
=pod
@@ -403,4 +404,80 @@ sub set_query_timer_restart
return $self->{query_timer_restart};
}
+=pod
+
+=item $session->poll_query_until($query [, $expected ])
+
+Run B<$query> repeatedly in this background session, until it returns the
+B<$expected> result ('t', or SQL boolean true, by default).
+Continues polling if the query returns an error result.
+Times out after a reasonable number of attempts.
+Returns 1 if successful, 0 if timed out.
+
+=cut
+
+sub poll_query_until
+{
+ my ($self, $query, $expected, %params) = @_;
+
+ $expected = 't' unless defined($expected); # default value
+
+ my $max_attempts = 10 * $PostgreSQL::Test::Utils::timeout_default;
+ my $attempts = 0;
+ my ($stdout, $stderr_flag);
+
+ while ($attempts < $max_attempts)
+ {
+ ($stdout, $stderr_flag) = $self->query($query, %params);
+
+ chomp($stdout);
+
+ # If query succeeded and returned expected result
+ if (!$stderr_flag && $stdout eq $expected)
+ {
+ return 1;
+ }
+
+ # Wait 0.1 second before retrying.
+ usleep(100_000);
+
+ $attempts++;
+ }
+
+ # Give up. Print the output from the last attempt, hopefully that's useful
+ # for debugging.
+ my $stderr_output = $stderr_flag ? $self->{stderr} : '';
+ diag qq(poll_query_until timed out executing this query:
+$query
+expecting this output:
+$expected
+last actual query output:
+$stdout
+with stderr:
+$stderr_output);
+ return 0;
+}
+
+=item $session->wait_for_event(backend_type, wait_event_name)
+
+Poll pg_stat_activity until backend_type reaches wait_event_name using this
+background session.
+
+=cut
+
+sub wait_for_event
+{
+ my ($self, $backend_type, $wait_event_name, %params) = @_;
+
+ $self->poll_query_until(
+ qq[
+ SELECT count(*) > 0 FROM pg_stat_activity
+ WHERE backend_type = '$backend_type' AND wait_event = '$wait_event_name'
+ ], undef, %params)
+ or die
+ qq(timed out when waiting for $backend_type to reach wait event '$wait_event_name');
+
+ return;
+}
+
1;
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 0ff41ec129c..17447c37f77 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -359,6 +359,7 @@ BufferAccessStrategy
BufferAccessStrategyType
BufferCacheOsPagesContext
BufferCacheOsPagesRec
+BufferControlBlock
BufferDesc
BufferDescPadded
BufferHeapTupleTableSlot
--
2.34.1
[text/x-patch] v20260817-0001-Add-BgBufferSync-sanity-Asserts.patch (2.1K, ../../CAExHW5ts93Rnof7pjFFYrY9aTOmCo+xsdQazqeqN1BaKozTBvA@mail.gmail.com/7-v20260817-0001-Add-BgBufferSync-sanity-Asserts.patch)
download | inline diff:
From 64454cbddfb02bf6483970a343d0ca1f9f4272b4 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 21 Jul 2026 16:16:48 +0530
Subject: [PATCH v20260817 1/7] Add BgBufferSync sanity Asserts
Assert that BgBufferSync() is only ever invoked from the background
writer process, matching the actual caller in BackgroundWriterMain.
Also move the existing Assert(strategy_delta >= 0) to the end of the
enclosing block so that the surrounding elog(DEBUG2) messages get a
chance to reach the server log before the Assert fires, easing diagnosis
if the Assert ever fails.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
src/backend/storage/buffer/bufmgr.c | 15 +++++++++++++--
1 file changed, 13 insertions(+), 2 deletions(-)
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 17f142e4c5b..5eadaabf6ca 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3894,6 +3894,8 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ Assert(AmBackgroundWriterProcess());
+
/*
* Find out where the clock-sweep currently is, and how many buffer
* allocations have happened since our last call.
@@ -3929,8 +3931,6 @@ BgBufferSync(WritebackContext *wb_context)
strategy_delta = strategy_buf_id - prev_strategy_buf_id;
strategy_delta += (long) passes_delta * NBuffers;
- Assert(strategy_delta >= 0);
-
if ((int32) (next_passes - strategy_passes) > 0)
{
/* we're one pass ahead of the strategy point */
@@ -3970,6 +3970,17 @@ BgBufferSync(WritebackContext *wb_context)
next_passes = strategy_passes;
bufs_to_lap = NBuffers;
}
+
+ /*
+ * We do not expect the current strategy point to be behind the
+ * previous one.
+ *
+ * If this Assert fails, we would have the debug messages printed in
+ * the server error log at appropriate debug level. Hence Asserting
+ * here provides minor convenience compared to Asserting right after
+ * calculating the difference.
+ */
+ Assert(strategy_delta >= 0);
}
else
{
base-commit: 51c43a5dbd86ad8461544e54c00bd9f487abfacd
--
2.34.1
[text/x-patch] v20260817-0002-Decouple-GUC-shared_buffers-and-size-of-th.patch (19.2K, ../../CAExHW5ts93Rnof7pjFFYrY9aTOmCo+xsdQazqeqN1BaKozTBvA@mail.gmail.com/8-v20260817-0002-Decouple-GUC-shared_buffers-and-size-of-th.patch)
download | inline diff:
From 02dc25ec4b5995623c0ebfd3704c3d701f29ec5c Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 4 Jun 2026 15:10:56 +0530
Subject: [PATCH v20260817 2/7] Decouple GUC shared_buffers and size of the
buffer pool
A single variable NBuffers holds the value of the GUC 'shared_buffers' and the
size of the buffer pool in number of buffers. With the introduction of resizable
shared buffer pool feature, the value of 'shared_buffers' GUC and the size of
the buffer pool can be different. This commit prepares for the same by
decoupling the two. A new variable NBuffersGUC holds the value of the GUC
'shared_buffers'. The variable NBuffers continues to hold the size of the
buffer pool. It is set to the value of NBuffersGUC during initialization.
Because of this decoupling, the code which references NBuffers as the value of
the GUC 'shared_buffers' is clearly differentiated from the code which
references NBuffers as the size of the buffer pool.
Also change comments to use term "size of the buffer pool" or "number of
buffers" instead of referencing variable NBuffers.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
contrib/pg_prewarm/autoprewarm.c | 2 +-
src/backend/access/heap/heapam.c | 15 ++++++++-------
src/backend/access/transam/slru.c | 2 +-
src/backend/access/transam/xlog.c | 8 ++++----
src/backend/optimizer/path/costsize.c | 8 ++++----
src/backend/postmaster/checkpointer.c | 11 ++++++-----
src/backend/storage/aio/aio_init.c | 2 +-
src/backend/storage/buffer/buf_init.c | 17 +++++++++++++----
src/backend/storage/buffer/buf_table.c | 14 ++++++++------
src/backend/storage/buffer/bufmgr.c | 13 +++++++------
src/backend/storage/buffer/freelist.c | 7 ++++---
src/backend/utils/init/globals.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 2 +-
src/include/miscadmin.h | 2 +-
src/include/storage/buf.h | 2 +-
src/include/storage/buf_internals.h | 4 ++--
src/include/storage/bufmgr.h | 3 ++-
17 files changed, 65 insertions(+), 49 deletions(-)
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index deb4c2671b5..87959af3c54 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -850,7 +850,7 @@ autoprewarm_start_worker(PG_FUNCTION_ARGS)
* SQL-callable function to perform an immediate block dump.
*
* Note: this is declared to return int8, as insurance against some
- * very distant day when we might make NBuffers wider than int.
+ * very distant day when we might make the size of buffer pool wider than int.
*/
Datum
autoprewarm_dump_now(PG_FUNCTION_ARGS)
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 72d6541734c..f1fade00aa2 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -383,13 +383,14 @@ initscan(HeapScanDesc scan, ScanKey key, bool keep_startblock)
scan->rs_nblocks = RelationGetNumberOfBlocks(scan->rs_base.rs_rd);
/*
- * If the table is large relative to NBuffers, use a bulk-read access
- * strategy and enable synchronized scanning (see syncscan.c). Although
- * the thresholds for these features could be different, we make them the
- * same so that there are only two behaviors to tune rather than four.
- * (However, some callers need to be able to disable one or both of these
- * behaviors, independently of the size of the table; also there is a GUC
- * variable that can disable synchronized scanning.)
+ * If the table is large relative to the size of the buffer pool, use a
+ * bulk-read access strategy and enable synchronized scanning (see
+ * syncscan.c). Although the thresholds for these features could be
+ * different, we make them the same so that there are only two behaviors
+ * to tune rather than four. (However, some callers need to be able to
+ * disable one or both of these behaviors, independently of the size of
+ * the table; also there is a GUC variable that can disable synchronized
+ * scanning.)
*
* Note that table_block_parallelscan_initialize has a very similar test;
* if you change this, consider changing that one, too.
diff --git a/src/backend/access/transam/slru.c b/src/backend/access/transam/slru.c
index 47dd52d6749..b47ace37974 100644
--- a/src/backend/access/transam/slru.c
+++ b/src/backend/access/transam/slru.c
@@ -236,7 +236,7 @@ SimpleLruAutotuneBuffers(int divisor, int max)
{
return Min(max - (max % SLRU_BANK_SIZE),
Max(SLRU_BANK_SIZE,
- NBuffers / divisor - (NBuffers / divisor) % SLRU_BANK_SIZE));
+ NBuffersGUC / divisor - (NBuffersGUC / divisor) % SLRU_BANK_SIZE));
}
/*
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index c3baca5193b..10650d01b9c 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -5017,14 +5017,14 @@ GetFakeLSNForUnloggedRel(void)
* and a minimum of 8 blocks (which was the default value prior to PostgreSQL
* 9.1, when auto-tuning was added).
*
- * This should not be called until NBuffers has received its final value.
+ * This should not be called until NBuffersGUC has received its final value.
*/
static int
XLOGChooseNumBuffers(void)
{
int xbuffers;
- xbuffers = NBuffers / 32;
+ xbuffers = NBuffersGUC / 32;
if (xbuffers > (wal_segment_size / XLOG_BLCKSZ))
xbuffers = (wal_segment_size / XLOG_BLCKSZ);
if (xbuffers < 8)
@@ -5297,8 +5297,8 @@ XLOGShmemRequest(void *arg)
/*
* If the value of wal_buffers is -1, use the preferred auto-tune value.
* This isn't an amazingly clean place to do this, but we must wait till
- * NBuffers has received its final value, and must do it before using the
- * value of XLOGbuffers to do anything important.
+ * NBuffersGUC has received its final value, and must do it before using
+ * the value of XLOGbuffers to do anything important.
*
* We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
* However, if the DBA explicitly set wal_buffers = -1 in the config file,
diff --git a/src/backend/optimizer/path/costsize.c b/src/backend/optimizer/path/costsize.c
index fd794c946ab..70c78cfb7e2 100644
--- a/src/backend/optimizer/path/costsize.c
+++ b/src/backend/optimizer/path/costsize.c
@@ -19,10 +19,10 @@
* is normally considerably less than random_page_cost. (However, if the
* database is fully cached in RAM, it is reasonable to set them equal.)
*
- * We also use a rough estimate "effective_cache_size" of the number of
- * disk pages in Postgres + OS-level disk cache. (We can't simply use
- * NBuffers for this purpose because that would ignore the effects of
- * the kernel's disk cache.)
+ * We also use a rough estimate "effective_cache_size" of the number of disk
+ * pages in Postgres + OS-level disk cache. (We can't simply use size of the
+ * buffer pool for this purpose because that would ignore the effects of the
+ * kernel's disk cache.)
*
* Obviously, taking constants for these values is an oversimplification,
* but it's tough enough to get any useful estimates even at this level of
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index 580c7944119..6640a9f8ea7 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -963,12 +963,13 @@ CheckpointerShmemRequest(void *arg)
Size size;
/*
- * The size of the requests[] array is arbitrarily set equal to NBuffers.
- * But there is a cap of MAX_CHECKPOINT_REQUESTS to prevent accumulating
- * too many checkpoint requests in the ring buffer.
+ * The size of the requests[] array is arbitrarily set equal to the
+ * initial size of buffer pool. But there is a cap of
+ * MAX_CHECKPOINT_REQUESTS to prevent accumulating too many checkpoint
+ * requests in the ring buffer.
*/
size = offsetof(CheckpointerShmemStruct, requests);
- size = add_size(size, mul_size(Min(NBuffers,
+ size = add_size(size, mul_size(Min(NBuffersGUC,
MAX_CHECKPOINT_REQUESTS),
sizeof(CheckpointerRequest)));
ShmemRequestStruct(.name = "Checkpointer Data",
@@ -985,7 +986,7 @@ static void
CheckpointerShmemInit(void *arg)
{
SpinLockInit(&CheckpointerShmem->ckpt_lck);
- CheckpointerShmem->max_requests = Min(NBuffers, MAX_CHECKPOINT_REQUESTS);
+ CheckpointerShmem->max_requests = Min(NBuffersGUC, MAX_CHECKPOINT_REQUESTS);
CheckpointerShmem->head = CheckpointerShmem->tail = 0;
ConditionVariableInit(&CheckpointerShmem->start_cv);
ConditionVariableInit(&CheckpointerShmem->done_cv);
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index de50e6a8a31..81825dd1b73 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -109,7 +109,7 @@ AioChooseMaxConcurrency(void)
/* Similar logic to LimitAdditionalPins() */
max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
- max_proportional_pins = NBuffers / max_backends;
+ max_proportional_pins = NBuffersGUC / max_backends;
max_proportional_pins = Max(max_proportional_pins, 1);
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 1407c930c56..9ddf6551fcd 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -77,21 +77,21 @@ static void
BufferManagerShmemRequest(void *arg)
{
ShmemRequestStruct(.name = "Buffer Descriptors",
- .size = NBuffers * sizeof(BufferDescPadded),
+ .size = NBuffersGUC * sizeof(BufferDescPadded),
/* Align descriptors to a cacheline boundary. */
.alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &BufferDescriptors,
);
ShmemRequestStruct(.name = "Buffer Blocks",
- .size = NBuffers * (Size) BLCKSZ,
+ .size = NBuffersGUC * (Size) BLCKSZ,
/* Align buffer pool on IO page size boundary. */
.alignment = PG_IO_ALIGN_SIZE,
.ptr = (void **) &BufferBlocks,
);
ShmemRequestStruct(.name = "Buffer IO Condition Variables",
- .size = NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ .size = NBuffersGUC * sizeof(ConditionVariableMinimallyPadded),
/* Align descriptors to a cacheline boundary. */
.alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &BufferIOCVArray,
@@ -105,7 +105,7 @@ BufferManagerShmemRequest(void *arg)
* painful.
*/
ShmemRequestStruct(.name = "Checkpoint BufferIds",
- .size = NBuffers * sizeof(CkptSortItem),
+ .size = NBuffersGUC * sizeof(CkptSortItem),
.ptr = (void **) &CkptBufferIds,
);
}
@@ -119,6 +119,12 @@ BufferManagerShmemRequest(void *arg)
static void
BufferManagerShmemInit(void *arg)
{
+ /*
+ * Set the size of the buffer pool, now that it's allocated and ready to
+ * be initialized.
+ */
+ NBuffers = NBuffersGUC;
+
/*
* Initialize all the buffer headers.
*/
@@ -147,6 +153,9 @@ BufferManagerShmemInit(void *arg)
static void
BufferManagerShmemAttach(void *arg)
{
+ /* Update the size of the buffer pool. */
+ NBuffers = NBuffersGUC;
+
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
&backend_flush_after);
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 347bf267d73..5c8ccbee13f 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -42,7 +42,7 @@ const ShmemCallbacks BufTableShmemCallbacks = {
/*
* Register shmem hash table for mapping buffers.
- * size is the desired hash table size (possibly more than NBuffers)
+ * size is the desired hash table size (possibly more than the size of the buffer pool).
*/
void
BufTableShmemRequest(void *arg)
@@ -54,12 +54,14 @@ BufTableShmemRequest(void *arg)
*
* Since we can't tolerate running out of lookup table entries, we must be
* sure to specify an adequate table size here. The maximum steady-state
- * usage is of course NBuffers entries, but BufferAlloc() tries to insert
- * a new entry before deleting the old. In principle this could be
- * happening in each partition concurrently, so we could need as many as
- * NBuffers + NUM_BUFFER_PARTITIONS entries.
+ * usage is of course as many entries as the number of buffers in the
+ * pool, but BufferAlloc() tries to insert a new entry before deleting the
+ * old. In principle this could be happening in each partition
+ * concurrently, so we could need as many as (number of buffers in the
+ * pool) + NUM_BUFFER_PARTITIONS entries. Since we are still requesting
+ * shared memory, use the GUC value instead of the actual size.
*/
- size = NBuffers + NUM_BUFFER_PARTITIONS;
+ size = NBuffersGUC + NUM_BUFFER_PARTITIONS;
ShmemRequestHash(.name = "Shared Buffer Lookup Table",
.nelems = size,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 5eadaabf6ca..bb6902777cd 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -223,6 +223,7 @@ int io_max_combine_limit = DEFAULT_IO_COMBINE_LIMIT;
int checkpoint_flush_after = DEFAULT_CHECKPOINT_FLUSH_AFTER;
int bgwriter_flush_after = DEFAULT_BGWRITER_FLUSH_AFTER;
int backend_flush_after = DEFAULT_BACKEND_FLUSH_AFTER;
+int NBuffers = 0; /* number of buffers in the buffer pool */
/* local state for LockBufferForCleanup */
static BufferDesc *PinCountWaitBuf = NULL;
@@ -239,11 +240,11 @@ static BufferDesc *PinCountWaitBuf = NULL;
* and, if so, in what mode.
*
*
- * To avoid - as we used to - requiring an array with NBuffers entries to keep
- * track of local buffers, we use a small sequentially searched array
- * (PrivateRefCountArrayKeys, with the corresponding data stored in
- * PrivateRefCountArray) and an overflow hash table (PrivateRefCountHash) to
- * keep track of backend local pins.
+ * To avoid - as we used to - requiring an array, with as many entries as the
+ * size of buffer pool, to keep track of local buffers, we use a small
+ * sequentially searched array (PrivateRefCountArrayKeys, with the corresponding
+ * data stored in PrivateRefCountArray) and an overflow hash table
+ * (PrivateRefCountHash) to keep track of backend local pins.
*
* Until no more than REFCOUNT_ARRAY_ENTRIES buffers are pinned at once, all
* refcounts are kept track of in the array; after that, new array entries
@@ -3642,7 +3643,7 @@ BufferSync(int flags)
set_bits, 0,
0);
- /* Check for barrier events in case NBuffers is large. */
+ /* Check for barrier events in case the buffer pool is large. */
if (ProcSignalBarrierPending)
ProcessProcSignalBarrier();
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index fdb5bad7910..4d5ee52ddc0 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -37,7 +37,8 @@ typedef struct
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo NBuffers.
+ * get an actual buffer, it needs to be used modulo size of the buffer
+ * pool.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -522,10 +523,10 @@ GetAccessStrategyWithSize(BufferAccessStrategyType btype, int ring_size_kb)
if (ring_buffers == 0)
return NULL;
- /* Cap to 1/8th of shared_buffers */
+ /* Cap to 1/8th of number of buffers in the buffer pool. */
ring_buffers = Min(NBuffers / 8, ring_buffers);
- /* NBuffers should never be less than 16, so this shouldn't happen */
+ /* Buffer pool should always have more than 16 buffers. */
Assert(ring_buffers > 0);
/* Allocate the object and initialize all elements to zeroes */
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index bbd28d14d99..ccf845e87b9 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -141,7 +141,7 @@ int max_parallel_maintenance_workers = 2;
* MaxBackends is computed by PostmasterMain after modules have had a chance to
* register background workers.
*/
-int NBuffers = 16384;
+int NBuffersGUC = 16384;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index 3c5e16ad1e7..5afa3e47552 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2725,7 +2725,7 @@
{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
- variable => 'NBuffers',
+ variable => 'NBuffersGUC',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 0fc59af02b9..e2df51d275a 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -175,7 +175,7 @@ extern PGDLLIMPORT bool ExitOnAnyError;
extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
-extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersGUC;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/storage/buf.h b/src/include/storage/buf.h
index b21445522b1..a12d6b9082a 100644
--- a/src/include/storage/buf.h
+++ b/src/include/storage/buf.h
@@ -17,7 +17,7 @@
/*
* Buffer identifiers.
*
- * Zero is invalid, positive is the index of a shared buffer (1..NBuffers),
+ * Zero is invalid, positive is the index of a shared buffer (1..{size of shared buffer pool}),
* negative is the index of a local buffer (-1 .. -NLocBuffer).
*/
typedef int Buffer;
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index e4ff5619b79..a7606c6e92b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -135,8 +135,8 @@ StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2),
/*
* The maximum allowed value of usage_count represents a tradeoff between
- * accuracy and speed of the clock-sweep buffer management algorithm. A
- * large value (comparable to NBuffers) would approximate LRU semantics.
+ * accuracy and speed of the clock-sweep buffer management algorithm. A large
+ * value (comparable to the size of buffer pool) would approximate LRU semantics.
* But it can take as many as BM_MAX_USAGE_COUNT+1 complete cycles of the
* clock-sweep hand to find a free buffer, so in practice we don't want the
* value to be very large.
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 6837b35fc6d..f1f6e601f51 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -159,13 +159,14 @@ typedef struct ReadBuffersOperation ReadBuffersOperation;
typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
-extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersGUC;
/* in bufmgr.c */
extern PGDLLIMPORT bool zero_damaged_pages;
extern PGDLLIMPORT int bgwriter_lru_maxpages;
extern PGDLLIMPORT double bgwriter_lru_multiplier;
extern PGDLLIMPORT bool track_io_timing;
+extern PGDLLIMPORT int NBuffers;
#define DEFAULT_EFFECTIVE_IO_CONCURRENCY 16
#define DEFAULT_MAINTENANCE_IO_CONCURRENCY 16
--
2.34.1
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-08-17 14:26 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-20 09:11 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Palak Chaturvedi @ 2026-08-17 14:26 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Ashutosh,
On Mon, 17 Aug 2026 at 17:27, Ashutosh Bapat
<ashutosh.bapat.oss@gmail.com> wrote:
>
> On Fri, Jul 24, 2026 at 6:26 PM Ashutosh Bapat
> <ashutosh.bapat.oss@gmail.com> wrote:
> >
> > Hi,
> >
> > On Fri, Feb 13, 2026 at 5:22 PM Ashutosh Bapat
> > <ashutosh.bapat.oss@gmail.com> wrote:
> > >
> > > On Thu, Feb 12, 2026 at 7:43 PM Jakub Wartak
> > > <jakub.wartak@enterprisedb.com> wrote:
> > > >
> > > >
> > > > TBH, I haven't really looked at the code outside of that region, I'm just
> > > > trespasser that was interested in memfd ;)
> > >
> > > Your trespassing has been very helpful. I have started a separate
> > > thread to discuss resizable shared structures at [1]. Once the
> > > implementation there is somewhat finalized, it will be good to try
> > > your huge page tests again.
> >
> > Here's the next version of the patch implementing shared buffer pool
> > resizing. The patch is based on the latest master. Here's the summary
> > of changes since the last version:
> >
> > 1. The patch now uses the new shared memory infrastructure that was
> > introduced in PG 19. Patch 0006 enhances that infrastructure to
> > support resizable shared structures. I will also post the same patch
> > to [1]. I am fine to discuss the patch in that thread or here. The
> > APIs for registering and resizing the structures are documented in the
> > programming interface documentation.
>
> Palak has provided an incremental patch fixing CI failures. It needs
> to be reviewed, hence not a part of this patchset.
>
> >
> > 2. Patch 0007 implements the shared buffer pool resizing using the
> > resizable shared structures infrastructure. It has a lot of code
> > improvements, including better documentation in comments, READMEs,
> > user-facing documentation and more TAP tests. The
> > storage/buffer/README has a section on buffer resizing. The buffer
> > resizing is implemented in buf_resize.c, which also has detailed
> > comments about the implementation. I suggest starting the review with
> > the user documentation, README and buf_resize.c.
> >
>
> The attached patches have a major change in this patch: Stress tests.
> I have added one stress test for every part of the code which scans
> the buffer to stress exercise that portion of code again buffer pool
> resizing. The patch also contains fixes for the crashes or bugs
> revealed by the stress tests. Specifically the stress tests cover
> synchronization between buffer pool resizing and
> DropRelationBuffers(), DropRelationsAllBuffers(),
> DropDatabaseBuffers(), CHECKPOINT, FlushRelationBuffers(),
> FlushRelationsAllBuffers(), pg_prewarm, monitoring and diagnostic
> functions in pg_buffercache. All these tests share common utility code
> in test/buffermgr/StressUtility.pm. At the end of each stress run, it
> carries out sanity checks to make sure that the database is not
> corrupted, the shared buffer pool state is in a sane state etc.
>
> FlushDatabaseBuffers() is not covered by any stress test since it is
> only called during WAL replay of xl_dbase_create_file_copy_rec. The
> primary CREATE DATABASE and ALTER DATABASE SET TABLESPACE paths use
> RequestCheckpoint() instead. I could not find a way to reach
> FlushDatabaseBuffers through normal SQL workload. But the function
> should be able to cope with the resized buffer pool just like other
> functions which scan the buffer pool.
>
> These tests are run only when PG_TEST_EXTRA has bufmgr_stress in it
> since these tests run longer (2 minutes each) and use many resources.
> We may not want to accept all these stress tests necessarily. We may
> want to pack all of them into a single test or just not accept any of
> them. They are pretty useful to build confidence that the reisizing
> protocol, shadow variables and barriers are working correctly and are
> hazard free. I would like to keep these tests in the patchset as long
> as possible and remove them just before the final commit to keep that
> confidence as we change the code and protocol while responding to the
> review comments. We may add more deterministic white box tests, like
> 001 and 002, using injection points for specific hazardous scenarios.
>
> Following bugs/crashes were revealed by the stress tests and their
> fixes (except one) are included in the patch.
> 1. BufferSync() is fixed to clean up a buffer-invalidated-by-resize
> properly from the checkpointer datastructures.
> 2. Most of the loops scanning the buffer pool invoked CFI at the
> beginning of the loop, which meant that the buffer being processed can
> be invalidated right at the beginning of each iteration. Instead moved
> CFI calls to the end of the loop.
> 3. Fixed pg_prewarm to not rely on NBuffers being static always,
> instead it adapts to the new size after CFI. But possibly we could
> change the function to process and write one buffer's tag at a time. I
> think we need a separate discussion for this.
>
> pg_buffercache_os_pages() still crashes when it hits a concurrent
> resize. But the fix is already being written.
Attached is the fix for this. It applies on top of v20260817-0007.
Rewrites pg_buffercache_os_pages_internal() to walk one buffer at a time
with stack-local scratch arrays, so it no longer sizes anything from
NBuffers upfront. Removes the multi-call SRF machinery and switches to
InitMaterializedSRF, matching the sibling functions in the same file.
The TODO comment in 0007 ("This allocates memory using NBuffers which
may change") is addressed and removed by this patch.
015_stress_pg_buffercache now passes cleanly (653 subtests, 138s) with
concurrent resize across the six pool sizes in StressUtil.
Performance (median of 3 runs, psql \timing, no concurrent load):
pool NBuffers base numa patched numa base os_pages patched os_pages
128M 16384 34.3 37.7 27.2 19.6
512M 65536 155 170 132 98.8
2G 262144 605 658 517 385
8G 1048576 2403 2590 2076 1538
NUMA path is 8 to 10% slower due to one move_pages(2) syscall per buffer
instead of one big call for the whole pool. Non-NUMA path is 25 to 28%
faster because streaming into the tuplestore is cheaper than the old
multi-call SRF trampoline.
Thanks,
Palak
Attachments:
[application/octet-stream] v20260817-0008-pg_buffercache-process-one-buffer-at-a-time.patch (15.1K, ../../CALfch19L-yVE+=quTraSu+KrGRGAP+AqY721hVXOcEYMPtaoOw@mail.gmail.com/2-v20260817-0008-pg_buffercache-process-one-buffer-at-a-time.patch)
download | inline diff:
From 2f5b75cf5eac2d751c29e02f3c66913993394b34 Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palak.chaturvedi@example.com>
Date: Fri, 14 Aug 2026 04:30:51 +0000
Subject: [PATCH v20260817] pg_buffercache: process one buffer at a time in
pg_buffercache_os_pages()
Following Ashutosh's suggestion, walk the buffer pool one buffer at a
time. For each buffer we compute the OS pages it overlaps, query the
NUMA state of just those pages via pg_numa_query_pages() with a small
stack-allocated scratch array, and stream one row per OS page directly
into the tuplestore.
This eliminates the fixed-size upfront allocation of fctx->record[] and
os_page_status[] that the old code sized from NBuffers, which is the
value pg_resize_shared_buffers() may change while this function runs.
No snapshot, no upfront allocation tied to NBuffers, no min-cap loop
bound, no assertions guarding invariants of a fixed-size array.
The loop is bounded by the live shadow NBuffers, matching the shape of
pg_buffercache_pages(): a concurrent shrink exits early, a concurrent
expand may or may not include the newly added tail depending on
timing, both cases return a partial-but-consistent snapshot.
Trade-offs vs the old batch approach:
- One pg_numa_query_pages syscall per buffer instead of one big call
covering the whole pool. Measured overhead at default 128MB shared
buffers is ~3 ms out of ~36 ms for pg_buffercache_numa (+7-10%).
For include_numa=false there is no syscall path at all.
- Per-invocation memory drops from O(NBuffers) records (megabytes at
typical sizes, ~750 MB at 128 GB shared_buffers) to O(1) stack.
- Non-NUMA path becomes ~30 % faster because streaming rows into the
tuplestore is cheaper than building fctx->record[] and returning
per-call.
Switches the SRF from multi-call to materialized, matching the sibling
functions pg_buffercache_pages() and pg_buffercache_usage_counts() in
the same file. Deletes BufferCacheOsPagesRec and
BufferCacheOsPagesContext typedefs and all SRF_IS_FIRSTCALL /
SRF_RETURN_NEXT scaffolding.
Verified: 015_stress_pg_buffercache passes (653 subtests, 138s) on
v20260817-0007 with concurrent resize across six pool sizes.
---
contrib/pg_buffercache/pg_buffercache_pages.c | 341 ++++--------------
1 file changed, 77 insertions(+), 264 deletions(-)
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index a1e72d79108..91f5c126520 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -32,33 +32,18 @@
#define NUM_BUFFERCACHE_OS_PAGES_ELEM 3
+/*
+ * Upper bound on OS pages a single BLCKSZ buffer can overlap. With BLCKSZ
+ * up to 32 KB and os_page_size at least 4 KB that's at most 9; 16 is a safe
+ * cap so the per-iteration scratch arrays fit on the stack.
+ */
+#define MAX_PAGES_PER_BUFFER 16
+
PG_MODULE_MAGIC_EXT(
.name = "pg_buffercache",
.version = PG_VERSION
);
-/*
- * Record structure holding the to be exposed cache data for OS pages. This
- * structure is used by pg_buffercache_os_pages(), where NUMA information may
- * or may not be included.
- */
-typedef struct
-{
- uint32 bufferid;
- int64 page_num;
- int32 numa_node;
-} BufferCacheOsPagesRec;
-
-/*
- * Function context for data persisting over repeated calls.
- */
-typedef struct
-{
- TupleDesc tupdesc;
- bool include_numa;
- BufferCacheOsPagesRec *record;
-} BufferCacheOsPagesContext;
-
static TupleDesc build_buffercache_pages_tupledesc(int natts);
@@ -285,279 +270,107 @@ build_buffercache_pages_tupledesc(int natts)
static Datum
pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
{
- FuncCallContext *funcctx;
- MemoryContext oldcontext;
- BufferCacheOsPagesContext *fctx; /* User function context. */
- TupleDesc tupledesc;
- TupleDesc expected_tupledesc;
- HeapTuple tuple;
- Datum result;
-
- /*
- * TODO: This allocates memory using NBuffers which may change while this
- * function is executed. We need to change this function so that it
- * doesn't rely on NBuffers being static throughout the execution of this
- * function.
- */
-
- if (SRF_IS_FIRSTCALL())
- {
- int i,
- idx;
- Size os_page_size;
- int pages_per_buffer;
- int *os_page_status = NULL;
- uint64 os_page_count = 0;
- int max_entries;
- char *startptr,
- *endptr;
-
- /* If NUMA information is requested, initialize NUMA support. */
- if (include_numa && pg_numa_init() == -1)
- elog(ERROR, "libnuma initialization failed or NUMA is not supported on this platform");
-
- /*
- * The database block size and OS memory page size are unlikely to be
- * the same. The block size is 1-32KB, the memory page size depends on
- * platform. On x86 it's usually 4KB, on ARM it's 4KB or 64KB, but
- * there are also features like THP etc. Moreover, we don't quite know
- * how the pages and buffers "align" in memory - the buffers may be
- * shifted in some way, using more memory pages than necessary.
- *
- * So we need to be careful about mapping buffers to memory pages. We
- * calculate the maximum number of pages a buffer might use, so that
- * we allocate enough space for the entries. And then we count the
- * actual number of entries as we scan the buffers.
- *
- * This information is needed before calling move_pages() for NUMA
- * node id inquiry.
- */
- os_page_size = pg_get_shmem_pagesize();
-
- /*
- * The pages and block size is expected to be 2^k, so one divides the
- * other (we don't know in which direction). This does not say
- * anything about relative alignment of pages/buffers.
- */
- Assert((os_page_size % BLCKSZ == 0) || (BLCKSZ % os_page_size == 0));
-
- if (include_numa)
- {
- void **os_page_ptrs = NULL;
-
- /*
- * How many addresses we are going to query? Simply get the page
- * for the first buffer, and first page after the last buffer, and
- * count the pages from that.
- */
- startptr = (char *) TYPEALIGN_DOWN(os_page_size,
- BufferGetBlock(1));
- endptr = (char *) TYPEALIGN(os_page_size,
- (char *) BufferGetBlock(NBuffers) + BLCKSZ);
- os_page_count = (endptr - startptr) / os_page_size;
-
- /* Used to determine the NUMA node for all OS pages at once */
- os_page_ptrs = palloc0_array(void *, os_page_count);
- os_page_status = palloc_array(int, os_page_count);
-
- /*
- * Fill pointers for all the memory pages. This loop stores and
- * touches (if needed) addresses into os_page_ptrs[] as input to
- * one big move_pages(2) inquiry system call, as done in
- * pg_numa_query_pages().
- */
- idx = 0;
- for (char *ptr = startptr; ptr < endptr; ptr += os_page_size)
- {
- os_page_ptrs[idx++] = ptr;
-
- /* Only need to touch memory once per backend process lifetime */
- if (firstNumaTouch)
- pg_numa_touch_mem_if_required(ptr);
- }
-
- Assert(idx == os_page_count);
-
- elog(DEBUG1, "NUMA: NBuffers=%d os_page_count=" UINT64_FORMAT " "
- "os_page_size=%zu", NBuffers, os_page_count, os_page_size);
-
- /*
- * If we ever get 0xff back from kernel inquiry, then we probably
- * have bug in our buffers to OS page mapping code here.
- */
- memset(os_page_status, 0xff, sizeof(int) * os_page_count);
-
- /* Query NUMA status for all the pointers */
- if (pg_numa_query_pages(0, os_page_count, os_page_ptrs, os_page_status) == -1)
- elog(ERROR, "failed NUMA pages inquiry: %m");
- }
-
- /* Initialize the multi-call context, load entries about buffers */
-
- funcctx = SRF_FIRSTCALL_INIT();
-
- /* Switch context when allocating stuff to be used in later calls */
- oldcontext = MemoryContextSwitchTo(funcctx->multi_call_memory_ctx);
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ Size os_page_size;
+ void *os_page_ptrs[MAX_PAGES_PER_BUFFER];
+ int os_page_status[MAX_PAGES_PER_BUFFER];
+ Datum values[NUM_BUFFERCACHE_OS_PAGES_ELEM];
+ bool nulls[NUM_BUFFERCACHE_OS_PAGES_ELEM];
+ char *startptr;
+ int i;
- /* Create a user function context for cross-call persistence */
- fctx = palloc_object(BufferCacheOsPagesContext);
+ InitMaterializedSRF(fcinfo, 0);
- if (get_call_result_type(fcinfo, NULL, &expected_tupledesc) != TYPEFUNC_COMPOSITE)
- elog(ERROR, "return type must be a row type");
+ if (include_numa && pg_numa_init() == -1)
+ elog(ERROR, "libnuma initialization failed or NUMA is not supported on this platform");
- if (expected_tupledesc->natts != NUM_BUFFERCACHE_OS_PAGES_ELEM)
- elog(ERROR, "incorrect number of output arguments");
+ os_page_size = pg_get_shmem_pagesize();
- /* Construct a tuple descriptor for the result rows. */
- tupledesc = CreateTemplateTupleDesc(expected_tupledesc->natts);
- TupleDescInitEntry(tupledesc, (AttrNumber) 1, "bufferid",
- INT4OID, -1, 0);
- TupleDescInitEntry(tupledesc, (AttrNumber) 2, "os_page_num",
- INT8OID, -1, 0);
- TupleDescInitEntry(tupledesc, (AttrNumber) 3, "numa_node",
- INT4OID, -1, 0);
+ Assert((os_page_size % BLCKSZ == 0) || (BLCKSZ % os_page_size == 0));
+ Assert(Max(1, BLCKSZ / os_page_size) + 1 <= MAX_PAGES_PER_BUFFER);
- TupleDescFinalize(tupledesc);
- fctx->tupdesc = BlessTupleDesc(tupledesc);
- fctx->include_numa = include_numa;
+ if (include_numa && firstNumaTouch)
+ elog(DEBUG1, "NUMA: page-faulting the buffercache for proper NUMA readouts");
- /*
- * Each buffer needs at least one entry, but it might be offset in
- * some way, and use one extra entry. So we allocate space for the
- * maximum number of entries we might need, and then count the exact
- * number as we're walking buffers. That way we can do it in one pass,
- * without reallocating memory.
- */
- pages_per_buffer = Max(1, BLCKSZ / os_page_size) + 1;
- max_entries = NBuffers * pages_per_buffer;
+ startptr = (char *) TYPEALIGN_DOWN(os_page_size,
+ (char *) BufferGetBlock(1));
- /* Allocate entries for BufferCacheOsPagesRec records. */
- fctx->record = (BufferCacheOsPagesRec *)
- MemoryContextAllocHuge(CurrentMemoryContext,
- sizeof(BufferCacheOsPagesRec) * max_entries);
+ /*
+ * We don't hold the partition locks, so we don't get a consistent
+ * snapshot across all buffers, but we do grab the buffer header locks,
+ * so the information of each buffer is self-consistent.
+ */
+ for (i = 0; i < NBuffers; i++)
+ {
+ char *buffptr = (char *) BufferGetBlock(i + 1);
+ char *startptr_buff = (char *) TYPEALIGN_DOWN(os_page_size,
+ buffptr);
+ char *endptr_buff = buffptr + BLCKSZ;
+ BufferDesc *bufHdr;
+ uint32 bufferid;
+ int32 page_num;
+ int num_pages_this_buffer = 0;
+ int j;
+ char *ptr;
- /* Return to original context when allocating transient memory */
- MemoryContextSwitchTo(oldcontext);
+ bufHdr = GetBufferDescriptor(i);
+ LockBufHdr(bufHdr);
+ bufferid = BufferDescriptorGetBuffer(bufHdr);
+ UnlockBufHdr(bufHdr);
- if (include_numa && firstNumaTouch)
- elog(DEBUG1, "NUMA: page-faulting the buffercache for proper NUMA readouts");
+ page_num = (startptr_buff - startptr) / os_page_size;
- /*
- * Scan through all the buffers, saving the relevant fields in the
- * fctx->record structure.
- *
- * We don't hold the partition locks, so we don't get a consistent
- * snapshot across all buffers, but we do grab the buffer header
- * locks, so the information of each buffer is self-consistent.
- */
- startptr = (char *) TYPEALIGN_DOWN(os_page_size, (char *) BufferGetBlock(1));
- idx = 0;
- for (i = 0; i < NBuffers; i++)
+ for (ptr = startptr_buff; ptr < endptr_buff; ptr += os_page_size)
{
- char *buffptr = (char *) BufferGetBlock(i + 1);
- BufferDesc *bufHdr;
- uint32 bufferid;
- int32 page_num;
- char *startptr_buff,
- *endptr_buff;
-
- bufHdr = GetBufferDescriptor(i);
-
- /* Lock each buffer header before inspecting. */
- LockBufHdr(bufHdr);
- bufferid = BufferDescriptorGetBuffer(bufHdr);
- UnlockBufHdr(bufHdr);
+ os_page_ptrs[num_pages_this_buffer++] = ptr;
- /* start of the first page of this buffer */
- startptr_buff = (char *) TYPEALIGN_DOWN(os_page_size, buffptr);
-
- /* end of the buffer (no need to align to memory page) */
- endptr_buff = buffptr + BLCKSZ;
-
- Assert(startptr_buff < endptr_buff);
-
- /* calculate ID of the first page for this buffer */
- page_num = (startptr_buff - startptr) / os_page_size;
-
- /* Add an entry for each OS page overlapping with this buffer. */
- for (char *ptr = startptr_buff; ptr < endptr_buff; ptr += os_page_size)
- {
- fctx->record[idx].bufferid = bufferid;
- fctx->record[idx].page_num = page_num;
- fctx->record[idx].numa_node = include_numa ? os_page_status[page_num] : -1;
-
- /* advance to the next entry/page */
- ++idx;
- ++page_num;
- }
-
- /*
- * Check for interrupts here, at the end of the loop, so that the
- * buffer index i remains valid till the next iteration.
- */
- CHECK_FOR_INTERRUPTS();
+ /* Only need to touch memory once per backend process lifetime */
+ if (include_numa && firstNumaTouch)
+ pg_numa_touch_mem_if_required(ptr);
}
- Assert(idx <= max_entries);
-
- if (include_numa)
- Assert(idx >= os_page_count);
-
- /* Set max calls and remember the user function context. */
- funcctx->max_calls = idx;
- funcctx->user_fctx = fctx;
-
- /* Remember this backend touched the pages (only relevant for NUMA) */
if (include_numa)
- firstNumaTouch = false;
- }
-
- funcctx = SRF_PERCALL_SETUP();
-
- /* Get the saved state */
- fctx = funcctx->user_fctx;
-
- if (funcctx->call_cntr < funcctx->max_calls)
- {
- uint32 i = funcctx->call_cntr;
- Datum values[NUM_BUFFERCACHE_OS_PAGES_ELEM];
- bool nulls[NUM_BUFFERCACHE_OS_PAGES_ELEM];
+ {
+ memset(os_page_status, 0xff, sizeof(int) * num_pages_this_buffer);
+ if (pg_numa_query_pages(0, num_pages_this_buffer,
+ os_page_ptrs, os_page_status) == -1)
+ elog(ERROR, "failed NUMA pages inquiry: %m");
+ }
- values[0] = Int32GetDatum(fctx->record[i].bufferid);
+ values[0] = Int32GetDatum(bufferid);
nulls[0] = false;
-
- values[1] = Int64GetDatum(fctx->record[i].page_num);
nulls[1] = false;
- if (fctx->include_numa)
+ for (j = 0; j < num_pages_this_buffer; j++)
{
- /* status is valid node number */
- if (fctx->record[i].numa_node >= 0)
+ values[1] = Int64GetDatum(page_num++);
+
+ if (include_numa && os_page_status[j] >= 0)
{
- values[2] = Int32GetDatum(fctx->record[i].numa_node);
+ values[2] = Int32GetDatum(os_page_status[j]);
nulls[2] = false;
}
else
{
- /* some kind of error (e.g. pages moved to swap) */
values[2] = (Datum) 0;
nulls[2] = true;
}
- }
- else
- {
- values[2] = (Datum) 0;
- nulls[2] = true;
- }
- /* Build and return the tuple. */
- tuple = heap_form_tuple(fctx->tupdesc, values, nulls);
- result = HeapTupleGetDatum(tuple);
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
- SRF_RETURN_NEXT(funcctx, result);
+ /*
+ * Check for interrupts here, at the end of the loop, so that the
+ * buffer index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
- else
- SRF_RETURN_DONE(funcctx);
+
+ if (include_numa)
+ firstNumaTouch = false;
+
+ return (Datum) 0;
}
/*
--
2.43.0
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 14:26 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
@ 2026-08-20 09:11 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Palak Chaturvedi @ 2026-08-20 09:11 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hey Ashutosh,
I tracked down a flake in 003_resize_failures (fails
roughly 1 in 5 runs). Two assertions fail at lines 157 and 163:
got: 't' expected: 'f'
log message "failed to expand buffer pool structures" absent
It's a SIGHUP delivery race. The test calls pg_reload_conf() from one
backend, then calls pg_resize_shared_buffers() from a different one
($resizer). When the signal hasn't reached $resizer yet, NBuffersGUC
is still the old shrunk value. resize_shared_buffers_internal() sees
currentNBuffers == targetNBuffers, prints "no need to resize," returns
true, and the injection point never fires.
Server log from a failing run:
06:07:17.885 [514631] SELECT pg_reload_conf()
06:07:17.885 postmaster: received SIGHUP, reloading configuration files
06:07:17.887 [514627] shared buffers are already at 18, no need to resize
As per discussion, attached patch (0009) moves the ALTER SYSTEM +
pg_reload_conf()
before creating the $resizer backend. The new backend picks up the
updated NBuffersGUC on connect. Injection point attachment and all
assertions stay the same. 30/30 clean after the fix.
Thanks,
Palak
Attachments:
[application/octet-stream] v20260817-0009-buffermgr-fix-SIGHUP-race-in-003_resize_failures.patch (2.2K, ../../CALfch19eG7JX3j0jvXLw7TdH+HOp1ZMacBwMMu07AiPPeNtHhw@mail.gmail.com/2-v20260817-0009-buffermgr-fix-SIGHUP-race-in-003_resize_failures.patch)
download | inline diff:
From cd76508701f9682db7aed8ba58a02643c4ae1fb0 Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Date: Wed, 19 Aug 2026 17:25:52 +0000
Subject: [PATCH v20260817] buffermgr: fix SIGHUP race in 003_resize_failures
The expansion-failure section created the resizer backend before
calling pg_reload_conf(), so the backend could start
pg_resize_shared_buffers() before the SIGHUP updated NBuffersGUC.
When that happened, resize_shared_buffers_internal() saw
currentNBuffers == targetNBuffers and returned true without reaching
the injection point, causing a flaky test failure (~20% rate).
Fix by creating the resizer backend after pg_reload_conf(). A new
backend inherits the current GUC values, so NBuffersGUC is guaranteed
to reflect the ALTER SYSTEM change.
Suggested-by: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
src/test/buffermgr/t/003_resize_failures.pl | 11 ++++++-----
1 file changed, 6 insertions(+), 5 deletions(-)
diff --git a/src/test/buffermgr/t/003_resize_failures.pl b/src/test/buffermgr/t/003_resize_failures.pl
index 9b92df66a59..9d34da75ca9 100644
--- a/src/test/buffermgr/t/003_resize_failures.pl
+++ b/src/test/buffermgr/t/003_resize_failures.pl
@@ -141,17 +141,18 @@ SKIP:
"SELECT name, size FROM pg_shmem_allocations WHERE name IN $resizable_structs ORDER BY name";
my $sizes_before = $node->safe_psql('postgres', $sizes_query);
+ my $expand_target = $max_nbuffers;
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$expand_target'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ # Start the resizer after reload so it inherits the updated NBuffersGUC.
my $resizer = $node->background_psql('postgres');
$resizer->query_safe("SELECT injection_points_set_local()", verbose => 0);
$resizer->query_safe(
"SELECT injection_points_attach('buffer-mgr-resize-struct-fail', 'notice')",
verbose => 0);
- my $expand_target = $max_nbuffers;
- $node->safe_psql('postgres',
- "ALTER SYSTEM SET shared_buffers = '$expand_target'");
- $node->safe_psql('postgres', "SELECT pg_reload_conf()");
-
my $expand_log_offset = -s $node->logfile;
is($resizer->query("SELECT pg_resize_shared_buffers()"),
--
2.43.0
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2026-08-20 11:13 ` Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Yuhang Qiu @ 2026-08-20 11:13 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; chaturvedipalak1911@gmail.com, Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Ashutosh,
I reviewed the v20260817 patches, and here is what I found:
0003 / 0004:
BufTableGetContents() holds all mapping partition locks for the whole scan,
with no CHECK_FOR_INTERRUPTS. pg_buffercache_lookup_table is a view whose name
ends in "table", which is ambiguous. What about pg_buffercache_mappings?
With a resize pending, pg_settings.setting reads "16384 (pending: 32768)",
could break pg_size_bytes(current_setting('shared_buffers')). Maybe we need a
new GUC rather than a composite one.
0006:
configure hasn't been regenerated. The feature might not be compiled by default.
Shrinking PANICs, whenever max_shared_buffers > shared_buffers.
The range handed to madvise(MADV_REMOVE) in ShmemResizeStruct() ends at
maximum_size rather than at the current size, so it covers the PROT_NONE tail
and fails with EACCES.
There are now six read-only shared_memory_* values, and the names are getting
long. A function might fit better than that many GUCs.
shared_memory_size_in_huge_pages is gone, it might break compatibility.
MADV_POPULATE_WRITE requires a new OS kernel version. #ifdef is needed in
PGSharedMemoryEnsureAllocated.
"could not protect shared memory" is emitted from two places, so it's not
possible to tell which one failed.
0007:
EvictExtraBuffers() checks BM_VALID, but BM_TAG_VALID is the flag that means
there's a mapping table entry, so buffers with IO in flight are skipped by the
precheck.
"shared buffer resizing to %d buffers failed" doesn't say why it failed.
Some values like MaxProportionalPins aren't recomputed on resize.
001_resize_fault_tolerance.pl never reaches madvise(): with
max_shared_buffers = 32 buffers and huge pages still on for the TAP cluster,
the range rounds away at 2MB granularity, so everything passes with nothing
freed. That is also why the shrink problem above goes unnoticed.
Best Regards,
Yuhang Qiu.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
@ 2026-08-26 16:36 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Palak Chaturvedi @ 2026-08-26 16:36 UTC (permalink / raw)
To: Yuhang Qiu <iamqyh@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Yuhang,
Thanks for the detailed review of v20260817. I went through
each finding against the code. Responses below, and one fix
patch (0010) attached.
On Thu, 20 Aug 2026 at 16:43, Yuhang Qiu <iamqyh@gmail.com> wrote:
>
> Hi Ashutosh,
>
> I reviewed the v20260817 patches, and here is what I found:
>
> 0003 / 0004:
>
> BufTableGetContents() holds all mapping partition locks for the whole scan,
> with no CHECK_FOR_INTERRUPTS. pg_buffercache_lookup_table is a view whose name
> ends in "table", which is ambiguous. What about pg_buffercache_mappings?
Good catch on the missing CHECK_FOR_INTERRUPTS. The locks themselves
are there to give a consistent view of the mapping table during the
scan, so I don't think we can simply drop them, but you're right that
a long scan with no interrupt check isn't great. Would add it as a TODO.
On the view name, I'd like to note it down and come back to it a bit
later. Right now the discussion on this patch is still fairly
high-level, so I think naming is better revisited once the shape
settles.
>
> With a resize pending, pg_settings.setting reads "16384 (pending: 32768)",
> could break pg_size_bytes(current_setting('shared_buffers')). Maybe we need a
> new GUC rather than a composite one.
>
You're right, and it does break rather than just might. I tried it:
ERROR: invalid size: "128MB (pending: 256MB)"
DETAIL: Invalid size unit: "MB (pending: 256MB)".
That happens for the whole window between the reload and the resize
finishing, so anything parsing that value would be affected.
Whether we go with a separate GUC or keep the composite string feels
like a design call that's better made when the patch is closer to
committable, so I'd suggest holding it for now. I just wanted it on
record that the breakage is real.
> 0006:
>
> configure hasn't been regenerated. The feature might not be compiled by default.
Need little more information on this.
>
> Shrinking PANICs, whenever max_shared_buffers > shared_buffers.
> The range handed to madvise(MADV_REMOVE) in ShmemResizeStruct() ends at
> maximum_size rather than at the current size, so it covers the PROT_NONE tail
> and fails with EACCES.
>
I agree with you on the range. It should stop at result->size so we
only free the pages actually being shrunk away, rather than running on
into the PROT_NONE tail.
One small thing so we're sure we're looking at the same problem: it
seems whether this actually PANICs depends on the kernel. On my box
(Linux 6.17, huge_pages=off) those over-wide MADV_REMOVE calls all
return 0, so I don't hit the PANIC locally. The range is still wrong
either way, so this doesn't change the conclusion.
> There are now six read-only shared_memory_* values, and the names are getting
> long. A function might fit better than that many GUCs.
> shared_memory_size_in_huge_pages is gone, it might break compatibility.
>
Both of these seem reasonable to me. Same as the naming point earlier,
I'd like to hold them until we're closer to commit.
> MADV_POPULATE_WRITE requires a new OS kernel version. #ifdef is needed in
> PGSharedMemoryEnsureAllocated.
>
I think the compile-time guard is actually already there, just
indirectly. The madvise call sits inside that function's existing
#ifndef HAVE_RESIZABLE_SHMEM block, and HAVE_RESIZABLE_SHMEM itself
requires HAVE_DECL_MADV_POPULATE_WRITE, so an extra #ifdef would end
up redundant.
> "could not protect shared memory" is emitted from two places, so it's not
> possible to tell which one failed.
>
Partly covered already, I think. The caller does report which
structure failed, so that part reaches the log. What's missing is just
which of the two mprotect calls it was, and those two fail for fairly
different reasons.
I've attached 0011 which gives each call site its own message, "could
not make shared memory read-write" and "could not make reserved shared
memory inaccessible".
> 0007:
>
> EvictExtraBuffers() checks BM_VALID, but BM_TAG_VALID is the flag that means
> there's a mapping table entry, so buffers with IO in flight are skipped by the
> precheck.
>
Agreed, and thanks for spotting this one. BM_TAG_VALID goes on when
BufferAlloc publishes the tag into the mapping table, while BM_VALID
only goes on after the read finishes. So a buffer sitting between the
two has a live hash entry that the current precheck walks past, and if
it happens to be above the new target when the shrink runs, that entry
gets orphaned.
I did try to reproduce it, including widening the window artificially,
but couldn't catch it. It seems to need AIO in the mix, so it would be
good to discuss this one a bit more.
> "shared buffer resizing to %d buffers failed" doesn't say why it failed.
>
Here I'd gently push back. I think that line is only meant as the
final summary, and the actual reason gets logged by whichever phase
failed, right before it. For example:
WARNING: failed to expand buffer pool structures
WARNING: shared buffer resizing to 25600 buffers failed
So the information is there, just on the previous line.
> Some values like MaxProportionalPins aren't recomputed on resize.
>
Confirmed, and this is what the attached 0010 fixes.
InitBufferManagerAccess() computes it once at startup and nothing
revisits it afterwards, so after a shrink every backend keeps using a
cap sized for the old pool. The patch moves the computation into
RecomputeMaxProportionalPins() and calls it from
ProcessBarrierBufferPoolSize(), so each backend refreshes as it takes
in the new pool size.
> 001_resize_fault_tolerance.pl never reaches madvise(): with
> max_shared_buffers = 32 buffers and huge pages still on for the TAP cluster,
> the range rounds away at 2MB granularity, so everything passes with nothing
> freed. That is also why the shrink problem above goes unnoticed.
>
Confirmed, and there's one more detail that makes it worse. The
buffermgr_test.conf that sets huge_pages=off is wired in as
TEMP_CONFIG for the regress target only, so the TAP clusters never
pick it up, and 001 doesn't set it itself either.
Could you share the error you saw, or how to reproduce it?
HugePages_Total is 0 on my machine, so the shrink does reach madvise
here and 001 still passes for me, and I'd like to be sure I'm seeing
the same thing you are.
Thanks again for going through the patch so carefully.
Thanks,
Palak
Attachments:
[application/octet-stream] v20260817-0010-buffermgr-recompute-MaxProportionalPins-after-buffer.patch (3.1K, ../../CALfch1_s98q32Ch3mo9J+kN0ycuGvLD-24fjVoQOAX1QyKGRNw@mail.gmail.com/2-v20260817-0010-buffermgr-recompute-MaxProportionalPins-after-buffer.patch)
download | inline diff:
From 1bb746905affe8955e9475ec68e5986aaf001fd6 Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Date: Fri, 21 Aug 2026 06:28:25 +0000
Subject: [PATCH v20260817 10/11] buffermgr: recompute MaxProportionalPins
after buffer pool resize
MaxProportionalPins is computed once at startup as
NBuffers / (MaxBackends + NUM_AUXILIARY_PROCS). After a buffer pool
resize, NBuffers changes but MaxProportionalPins stays stale. If the
pool shrinks, backends can hold too many pins relative to the new pool
size, defeating the per-backend pin limit.
Extract RecomputeMaxProportionalPins() from InitBufferManagerAccess()
and call it from ProcessBarrierBufferPoolSize() so every backend
recomputes the limit when it absorbs the new pool size.
Reported-by: Yuhang Qiu <iamqyh@gmail.com>
---
src/backend/storage/buffer/buf_resize.c | 1 +
src/backend/storage/buffer/bufmgr.c | 11 ++++++++++-
src/include/storage/bufmgr.h | 1 +
3 files changed, 12 insertions(+), 1 deletion(-)
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
index 90f5fb1c71d..d8b5a0ad845 100644
--- a/src/backend/storage/buffer/buf_resize.c
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -503,6 +503,7 @@ ProcessBarrierBufferPoolSize(void)
Assert(activeNBuffers == pg_atomic_read_u32(&BufferControl->activeNBuffers));
NBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+ RecomputeMaxProportionalPins();
return true;
}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 2aac77ccc35..787419cc7f2 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -275,6 +275,15 @@ static int PrivateRefCountEntryLast = -1;
static uint32 MaxProportionalPins;
+/*
+ * Recompute the per-backend pin limit after a buffer pool resize.
+ */
+void
+RecomputeMaxProportionalPins(void)
+{
+ MaxProportionalPins = NBuffers / (MaxBackends + NUM_AUXILIARY_PROCS);
+}
+
static void ReservePrivateRefCountEntry(void);
static PrivateRefCountEntry *NewPrivateRefCountEntry(Buffer buffer);
static PrivateRefCountEntry *GetPrivateRefCountEntry(Buffer buffer, bool do_move);
@@ -4328,7 +4337,7 @@ InitBufferManagerAccess(void)
* allow plenty of pins. LimitAdditionalPins() and
* GetAdditionalPinLimit() can be used to check the remaining balance.
*/
- MaxProportionalPins = NBuffers / (MaxBackends + NUM_AUXILIARY_PROCS);
+ RecomputeMaxProportionalPins();
memset(&PrivateRefCountArray, 0, sizeof(PrivateRefCountArray));
memset(&PrivateRefCountArrayKeys, 0, sizeof(PrivateRefCountArrayKeys));
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 187e9b29a9e..650989fcdcd 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -287,6 +287,7 @@ extern Buffer ExtendBufferedRelTo(BufferManagerRelation bmr,
ReadBufferMode mode);
extern void InitBufferManagerAccess(void);
+extern void RecomputeMaxProportionalPins(void);
extern void AtEOXact_Buffers(bool isCommit);
#ifdef USE_ASSERT_CHECKING
extern void AssertBufferLocksPermitCatalogRead(void);
--
2.43.0
[application/octet-stream] v20260817-0011-shmem-distinguish-the-two-mprotect-failure-messages.patch (1.8K, ../../CALfch1_s98q32Ch3mo9J+kN0ycuGvLD-24fjVoQOAX1QyKGRNw@mail.gmail.com/3-v20260817-0011-shmem-distinguish-the-two-mprotect-failure-messages.patch)
download | inline diff:
From 1535129164892679dd7a94ee5007e7ba0e128b5d Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Date: Wed, 26 Aug 2026 16:15:47 +0000
Subject: [PATCH v20260817 11/11] shmem: distinguish the two mprotect failure
messages
PGSharedMemoryProtect() calls mprotect() twice, once to make the active
region read-write and once to make the reserved tail inaccessible, and
both failure paths emitted the identical "could not protect shared
memory" message. The caller reports which structure failed, but not
which of the two calls, even though they fail for quite different
reasons.
Give each call site its own message.
Reported-by: Yuhang Qiu <iamqyh@gmail.com>
Discussion: https://postgr.es/m/B6CC6AF0-F7B3-4389-9740-DD1288EC15DB@gmail.com
---
src/backend/port/sysv_shmem.c | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index c052776e94c..179ceab57ff 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1200,7 +1200,8 @@ PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
if (mprotect(rw_start, (char *) rw_end - (char *) rw_start,
PROT_READ | PROT_WRITE) != 0)
{
- ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ ereport(WARNING,
+ errmsg("could not make shared memory read-write: %m"));
return false;
}
}
@@ -1210,7 +1211,8 @@ PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
if (mprotect(rw_end, (char *) prot_end - (char *) rw_end,
PROT_NONE) != 0)
{
- ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ ereport(WARNING,
+ errmsg("could not make reserved shared memory inaccessible: %m"));
return false;
}
}
--
2.43.0
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
@ 2026-08-27 09:01 ` Yuhang Qiu <iamqyh@gmail.com>
2026-09-07 13:00 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Yuhang Qiu @ 2026-08-27 09:01 UTC (permalink / raw)
To: Palak Chaturvedi <chaturvedipalak1911@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Palak,
0010 and 0011 both look good to me.
I missed the existing `HAVE_RESIZABLE_SHMEM` guard. Ignore that comment.
> It seems whether this actually PANICs depends on the kernel. On my box
> (Linux 6.17, huge_pages=off) those over-wide MADV_REMOVE calls all
> return 0, so I don't hit the PANIC locally.
I reproduced it on Linux 5.10 with `huge_pages=off`. The code calls
`MADV_REMOVE` from the new end to `maximum_size`, crossing from the RW
area into the existing `PROT_NONE` tail. The call returns EACCES and the
resize PANICs.
I tested this fix:
```c
char *current_end = (char *) TYPEALIGN(page_size,
(char *) result->location + result->size);
char *reserved_end = (char *) TYPEALIGN_DOWN(page_size,
(char *) result->location + result->maximum_size);
char *max_end = Min(current_end, reserved_end);
```
The first bound stops at the current allocation; the second preserves the
last page when it is shared with the next structure. Linux 6.7 changed
the check from `VM_WRITE` to `VM_MAYWRITE` [1], which explains why the
over-wide call succeeds on 6.17.
> I did try to reproduce it, including widening the window artificially,
> but couldn't catch it. It seems to need AIO in the mix.
The race is:
1. A backend takes a buffer above the shrink target before processing the
new-allocation barrier.
2. It publishes the mapping entry (`BM_TAG_VALID` is set), while `BM_VALID`
is still clear.
3. It processes the barrier before the read completes. This can also
happen on the synchronous path between `StartReadBuffers()` and
`WaitReadBuffers()`.
4. `EvictExtraBuffers()` sees `BM_VALID` clear and skips the buffer, leaving
the mapping entry behind after the shrink.
To force it, hold the read completion after step 2, let the backend process
the barrier, and then start the shrink.
Both the precheck and the locked assertion should use `BM_TAG_VALID`, and
the latter should check the state returned by `LockBufHdr()`. The shrink
will then roll back instead of leaving an orphaned entry.
Running `autoconf` will update `configure` with new `configure.ac`.
> Could you share the error you saw, or how to reproduce it?
I tested it with `huge_pages = off` on Linux 5.10. The error is EACCES which
I explained above.
Best regards,
Yuhang Qiu
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=e8e17ee90eaf650c855adb...
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
@ 2026-09-07 13:00 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-08 03:09 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Palak Chaturvedi @ 2026-09-07 13:00 UTC (permalink / raw)
To: Yuhang Qiu <iamqyh@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Yuhang,
On Thu, 27 Aug 2026 at 14:31, Yuhang Qiu <iamqyh@gmail.com> wrote:
>
> Hi Palak,
>
> 0010 and 0011 both look good to me.
>
Thanks
> I missed the existing `HAVE_RESIZABLE_SHMEM` guard. Ignore that comment.
>
Got it.
> > It seems whether this actually PANICs depends on the kernel. On my box
> > (Linux 6.17, huge_pages=off) those over-wide MADV_REMOVE calls all
> > return 0, so I don't hit the PANIC locally.
>
> I reproduced it on Linux 5.10 with `huge_pages=off`. The code calls
> `MADV_REMOVE` from the new end to `maximum_size`, crossing from the RW
> area into the existing `PROT_NONE` tail. The call returns EACCES and the
> resize PANICs.
>
> I tested this fix:
> ```c
> char *current_end = (char *) TYPEALIGN(page_size,
> (char *) result->location + result->size);
> char *reserved_end = (char *) TYPEALIGN_DOWN(page_size,
> (char *) result->location + result->maximum_size);
> char *max_end = Min(current_end, reserved_end);
> ```
>
Tried this. This works fine. Attached a patch with this. Also did the testing
on my box with huge_pages=off and confirmed the EACCES/PANIC before the
fix and a clean resize after.
> The first bound stops at the current allocation; the second preserves the
> last page when it is shared with the next structure. Linux 6.7 changed
> the check from `VM_WRITE` to `VM_MAYWRITE` [1], which explains why the
> over-wide call succeeds on 6.17.
>
> > I did try to reproduce it, including widening the window artificially,
> > but couldn't catch it. It seems to need AIO in the mix.
>
> The race is:
> 1. A backend takes a buffer above the shrink target before processing the
> new-allocation barrier.
> 2. It publishes the mapping entry (`BM_TAG_VALID` is set), while `BM_VALID`
> is still clear.
> 3. It processes the barrier before the read completes. This can also
> happen on the synchronous path between `StartReadBuffers()` and
> `WaitReadBuffers()`.
> 4. `EvictExtraBuffers()` sees `BM_VALID` clear and skips the buffer, leaving
> the mapping entry behind after the shrink.
>
> To force it, hold the read completion after step 2, let the backend process
> the barrier, and then start the shrink.
>
> Both the precheck and the locked assertion should use `BM_TAG_VALID`, and
> the latter should check the state returned by `LockBufHdr()`. The shrink
> will then roll back instead of leaving an orphaned entry.
>
Thanks for the steps. Reproduced this using the steps via creating a test
file and adding an injection point right after BM_TAG_VALID is set, before
the read completes. Was able to fail it at the Assert (signal 6) with the
code as posted. Not attaching the test file or injection point as part of
the patch, happy to share the test
file separately if useful, just let me know.
Here is the patch for EvictExtraBuffers. Changed both the precheck and the
locked assertion to BM_TAG_VALID, and fixed the assertion to check the
state returned by LockBufHdr() instead of the stale unlocked read. With
this, the shrink rolls back and logs the pinned buffer instead of leaving
an orphaned mapping. Also went through every other function in bufmgr.c
that walks buffers and checks BM_VALID (EvictAllUnpinnedBuffers,
EvictRelUnpinnedBuffers, SyncOneBuffer, FlushRelationBuffers,
FlushDatabaseBuffers, the MarkDirty helpers) — none of those are affected,
a mid-IO buffer can't be dirty so skipping it there is correct.
> Running `autoconf` will update `configure` with new `configure.ac`.
>
Right, configure.ac has AC_CHECK_DECLS for MADV_POPULATE_WRITE and
MADV_REMOVE but configure was never regenerated, so autotools builds
never get HAVE_RESIZABLE_SHMEM defined. Ran autoconf 2.69 and attached
the regenerated configure.
> > Could you share the error you saw, or how to reproduce it?
> I tested it with `huge_pages = off` on Linux 5.10. The error is EACCES which
> I explained above.
>
> Best regards,
> Yuhang Qiu
>
> [1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=e8e17ee90eaf650c855adb...
>
Thanks,
Palak
Attachments:
[application/octet-stream] v20260907-0013-buffermgr-fix-EvictExtraBuffers-BM_TAG_VALID.patch (1.3K, ../../CALfch1_K04G39nNrZCBw+pbqG4aXMXsAcv04A_MRhEtjJi2T8A@mail.gmail.com/2-v20260907-0013-buffermgr-fix-EvictExtraBuffers-BM_TAG_VALID.patch)
download | inline diff:
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index bb436734585..06437161135 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -9143,22 +9143,24 @@ EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
buf_state = pg_atomic_read_u64(&desc->state);
/*
- * Nobody is expected to allocate new buffers while resizing is going
- * on hence unlocked precheck should be safe and saves some cycles.
+ * A buffer whose tag is not published carries no hash-table mapping
+ * that shrinking has to clean up. BM_VALID is not sufficient here:
+ * a buffer mid-IO has BM_TAG_VALID but not yet BM_VALID, and skipping
+ * it would leave an orphaned mapping in the shrunk pool.
*/
- if (!(buf_state & BM_VALID))
+ if (!(buf_state & BM_TAG_VALID))
continue;
ResourceOwnerEnlarge(CurrentResourceOwner);
ReservePrivateRefCountEntry();
- LockBufHdr(desc);
+ buf_state = LockBufHdr(desc);
/*
* Now that we have locked buffer descriptor, make sure that the
- * buffer without valid data has been skipped above.
+ * buffer without a tag has been skipped above.
*/
- Assert(buf_state & BM_VALID);
+ Assert(buf_state & BM_TAG_VALID);
if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
{
[application/octet-stream] v20260907-0012-shmem-fix-madv-remove-range.patch (979B, ../../CALfch1_K04G39nNrZCBw+pbqG4aXMXsAcv04A_MRhEtjJi2T8A@mail.gmail.com/3-v20260907-0012-shmem-fix-madv-remove-range.patch)
download | inline diff:
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index fbc05f0dc13..c9a9b7ec66c 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -925,7 +925,14 @@ ShmemResizeStruct(const char *name, Size new_size)
new_end = (char *) TYPEALIGN(page_size, (char *) result->location + new_size);
if (new_size < result->size)
{
- char *max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+ /*
+ * Only pages currently RW should be freed; anything past result->size
+ * is PROT_NONE reserved tail. Also stop before the last page shared
+ * with the next structure.
+ */
+ char *current_end = (char *) TYPEALIGN(page_size, (char *) result->location + result->size);
+ char *reserved_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+ char *max_end = Min(current_end, reserved_end);
if (max_end > new_end)
{
[application/octet-stream] v20260907-0014-configure-regenerate-for-MADV-checks.patch (1.1K, ../../CALfch1_K04G39nNrZCBw+pbqG4aXMXsAcv04A_MRhEtjJi2T8A@mail.gmail.com/4-v20260907-0014-configure-regenerate-for-MADV-checks.patch)
download | inline diff:
diff --git a/configure b/configure
index d42a7a794ff..6e61394a81d 100755
--- a/configure
+++ b/configure
@@ -16414,6 +16414,32 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux-specific madvise constants needed for resizable shared memory. See similar checks in meson.build for explanation of why these checks are here.
+ac_fn_c_check_decl "$LINENO" "MADV_POPULATE_WRITE" "ac_cv_have_decl_MADV_POPULATE_WRITE" "#include <sys/mman.h>
+"
+if test "x$ac_cv_have_decl_MADV_POPULATE_WRITE" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_MADV_POPULATE_WRITE $ac_have_decl
+_ACEOF
+
+ac_fn_c_check_decl "$LINENO" "MADV_REMOVE" "ac_cv_have_decl_MADV_REMOVE" "#include <sys/mman.h>
+"
+if test "x$ac_cv_have_decl_MADV_REMOVE" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_MADV_REMOVE $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-07 13:00 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
@ 2026-09-08 03:09 ` Yuhang Qiu <iamqyh@gmail.com>
2026-09-08 11:24 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Yuhang Qiu @ 2026-09-08 03:09 UTC (permalink / raw)
To: Palak Chaturvedi <chaturvedipalak1911@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Palak,
> Changed both the precheck and the locked assertion to BM_TAG_VALID
I missed a race in my earlier suggestion. Restricting new allocations
does not prevent another backend from invalidating an existing buffer.
For example, DROP TABLE can clear the tag between the unlocked precheck
and LockBufHdr(), causing the assertion to fail.
Could we recheck BM_TAG_VALID under the lock, and unlock and continue
if it is already clear?
Also, the earlier comment in 0012 still says "We do not consider the
current end of the structure". That is no longer true for shrinking.
Best regards,
Yuhang Qiu
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-07 13:00 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-08 03:09 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
@ 2026-09-08 11:24 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-11 10:56 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Palak Chaturvedi @ 2026-09-08 11:24 UTC (permalink / raw)
To: Yuhang Qiu <iamqyh@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hey Yuhang,
On Mon, 7 Sept 2026 at 20:09, Yuhang Qiu <iamqyh@gmail.com> wrote:
>
> Hi Palak,
>
> > Changed both the precheck and the locked assertion to BM_TAG_VALID
>
> I missed a race in my earlier suggestion. Restricting new allocations
> does not prevent another backend from invalidating an existing buffer.
> For example, DROP TABLE can clear the tag between the unlocked precheck
> and LockBufHdr(), causing the assertion to fail.
>
Reproduced this. I added a test-only injection point right after DROP
TABLE clears BM_TAG_VALID but before EvictExtraBuffers() takes
LockBufHdr(), and the assertion trips every time. Also found it
independently while running the stress tests, 6 hits, all TRAP on the
BM_TAG_VALID assertion.
> Could we recheck BM_TAG_VALID under the lock, and unlock and continue
> if it is already clear?
>
Yes. v20260908-0016 rechecks BM_TAG_VALID under LockBufHdr() and does
UnlockBufHdr() + continue instead of asserting when the tag was
cleared concurrently.
> Also, the earlier comment in 0012 still says "We do not consider the
> current end of the structure". That is no longer true for shrinking.
>
Right, that line predates the MADV_REMOVE range fix. Removed it in
v20260908-0015.
Thanks,
Palak
Attachments:
[application/octet-stream] v20260908-0016-buffermgr-fix-EvictExtraBuffers-BM_TAG_VALID.patch (1.6K, ../../CALfch18bfYdXz1E7OZO45ECKN9Sn=L-tYXg=mhWa0b2k3c7Qnw@mail.gmail.com/2-v20260908-0016-buffermgr-fix-EvictExtraBuffers-BM_TAG_VALID.patch)
download | inline diff:
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index bb436734585..db1cb8648ab 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -9143,22 +9162,31 @@ EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
buf_state = pg_atomic_read_u64(&desc->state);
/*
- * Nobody is expected to allocate new buffers while resizing is going
- * on hence unlocked precheck should be safe and saves some cycles.
+ * A buffer whose tag is not published carries no hash-table mapping
+ * that shrinking has to clean up. BM_VALID is not sufficient here:
+ * a buffer mid-IO has BM_TAG_VALID but not yet BM_VALID, and skipping
+ * it would leave an orphaned mapping in the shrunk pool.
*/
- if (!(buf_state & BM_VALID))
+ if (!(buf_state & BM_TAG_VALID))
continue;
ResourceOwnerEnlarge(CurrentResourceOwner);
ReservePrivateRefCountEntry();
- LockBufHdr(desc);
+ buf_state = LockBufHdr(desc);
/*
- * Now that we have locked buffer descriptor, make sure that the
- * buffer without valid data has been skipped above.
+ * The unlocked precheck above is only a hint: a concurrent
+ * InvalidateBuffer() (e.g. DROP TABLE, TRUNCATE, DROP DATABASE) can
+ * clear BM_TAG_VALID between the precheck and this lock without
+ * needing our pin. Recheck under the lock and skip if so; there is
+ * nothing left for us to evict.
*/
- Assert(buf_state & BM_VALID);
+ if (!(buf_state & BM_TAG_VALID))
+ {
+ UnlockBufHdr(desc);
+ continue;
+ }
if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
{
[application/octet-stream] v20260908-0015-shmem-fix-madv-remove-range.patch (1.3K, ../../CALfch18bfYdXz1E7OZO45ECKN9Sn=L-tYXg=mhWa0b2k3c7Qnw@mail.gmail.com/3-v20260908-0015-shmem-fix-madv-remove-range.patch)
download | inline diff:
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index fbc05f0dc13..cb1f3586d4c 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -917,15 +917,18 @@ ShmemResizeStruct(const char *name, Size new_size)
* When shrinking, release memory pages beyond the new end, but not the
* page containing maximal end of the structure, as it may be used by the
* next structure.
- *
- * We do not consider the current end of the structure as it simplifies
- * the calculations. Instead we rely on the underlying APIs not to touch
- * the memory pages that will not be affected by the change in size.
*/
new_end = (char *) TYPEALIGN(page_size, (char *) result->location + new_size);
if (new_size < result->size)
{
- char *max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+ /*
+ * Only pages currently RW should be freed; anything past result->size
+ * is PROT_NONE reserved tail. Also stop before the last page shared
+ * with the next structure.
+ */
+ char *current_end = (char *) TYPEALIGN(page_size, (char *) result->location + result->size);
+ char *reserved_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+ char *max_end = Min(current_end, reserved_end);
if (max_end > new_end)
{
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-07 13:00 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-08 03:09 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-08 11:24 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
@ 2026-09-11 10:56 ` Yuhang Qiu <iamqyh@gmail.com>
2026-09-17 16:02 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Yuhang Qiu @ 2026-09-11 10:56 UTC (permalink / raw)
To: Palak Chaturvedi <chaturvedipalak1911@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Palak,
> Yes. v20260908-0016 rechecks BM_TAG_VALID under LockBufHdr() and does
> UnlockBufHdr() + continue instead of asserting when the tag was
> cleared concurrently.
>
> Right, that line predates the MADV_REMOVE range fix. Removed it in
> v20260908-0015.
Thanks. 0015 and 0016 LGTM.
I also tested the patches and did some more review.
With shared_buffers=32MB and max_shared_buffers left at its default, parallel
workers fail to start, even without any resize:
FATAL: failed to initialize shared_buffers to 16384
CONTEXT: parallel worker
RestoreGUCState() resets the GUC to its 128MB boot value, which exceeds
MaxNBuffers and fails the check hook.
There is also a performance issue in the grow path:
PGSharedMemoryEnsureAllocated() gets the whole structure's range, rather than
just the added range. Even a small increase makes MADV_POPULATE_WRITE walk the
existing range again, adding overhead for large buffer pools. Could we limit
this to the page-aligned [current_end, new_end), as in the shrink path?
Two minor points:
* check_shared_buffers() allows equality with MaxNBuffers, so "must be less than"
should be "must not exceed".
* In proc.c, "specifid there" should be "specified there".
Some issues raised in earlier reviews remain unaddressed. Could you include a
status list with the next update, separating resolved and unresolved items?
That would help reviewers who are new to the thread or haven't followed it for
a while.
Best regards,
Yuhang Qiu
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-07 13:00 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-08 03:09 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-08 11:24 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-11 10:56 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
@ 2026-09-17 16:02 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-22 03:45 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-22 07:02 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 2 replies; 167+ messages in thread
From: Palak Chaturvedi @ 2026-09-17 16:02 UTC (permalink / raw)
To: Yuhang Qiu <iamqyh@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Yuhang,
Sorry for the late reply. Was working through the patches in order.
On Fri, 11 Sept 2026 at 03:56, Yuhang Qiu <iamqyh@gmail.com> wrote:
>
> Hi Palak,
>
> > Yes. v20260908-0016 rechecks BM_TAG_VALID under LockBufHdr() and does
> > UnlockBufHdr() + continue instead of asserting when the tag was
> > cleared concurrently.
> >
> > Right, that line predates the MADV_REMOVE range fix. Removed it in
> > v20260908-0015.
>
> Thanks. 0015 and 0016 LGTM.
Thanks.
> I also tested the patches and did some more review.
>
> With shared_buffers=32MB and max_shared_buffers left at its default,
parallel
> workers fail to start, even without any resize:
> FATAL: failed to initialize shared_buffers to 16384
> CONTEXT: parallel worker
>
> RestoreGUCState() resets the GUC to its 128MB boot value, which exceeds
> MaxNBuffers and fails the check hook.
Reproduced this with shared_buffers=32MB and max_shared_buffers left at
its default. Fixed in 0015 by allowing the temporary PGC_S_DEFAULT
assignment only while InitializingParallelWorker is set.
Also checked removing the last configured value. Allowing PGC_S_DEFAULT
unconditionally accepts the 128MB default above the reserved 32MB
maximum, followed by a PANIC on resize. The revised check rejects this
reset.
Added test 008 for both cases. All nine assertions pass, including
launching two workers. Without the fix, it fails with the reported
initialization error. It also fails with the unconditional
default-source exception.
> There is also a performance issue in the grow path:
> PGSharedMemoryEnsureAllocated() gets the whole structure's range, rather
than
> just the added range. Even a small increase makes MADV_POPULATE_WRITE
walk the
> existing range again, adding overhead for large buffer pools. Could we
limit
> this to the page-aligned [current_end, new_end), as in the shrink path?
Changed this in 0013 to populate only the newly added pages. Also did
the testing. For a 512MB to 528MB grow, seven untraced runs gave a median
resize time of 332.7ms with the old range and 11.4ms with the new range.
Repeated the optimized run after the control and got 11.4ms again.
This was on Linux 6.17 with huge_pages=off, no swap and no concurrent
workload. These are total resize times for this case. Existing pages
that have been swapped out will be brought back when accessed. I have
not measured the effect under memory pressure with swap enabled.
> Two minor points:
> * check_shared_buffers() allows equality with MaxNBuffers, so "must be
less than"
> should be "must not exceed".
> * In proc.c, "specifid there" should be "specified there".
Changed "must be less than" to "must not exceed" in 0015 and corrected
"specifid" to "specified" in 0016.
> Some issues raised in earlier reviews remain unaddressed. Could you
include a
> status list with the next update, separating resolved and unresolved
items?
> That would help reviewers who are new to the thread or haven't followed
it for
> a while.
Included the complete series and renumbered it as v20260917-0001
through 0019. The numbers below refer to this new set. Patches 0001-0010
are Ashutosh's base commits, including his fixups, with authorship
preserved. Patches 0011-0019 are my earlier fixes and the current
follow-ups. No earlier attachments need to be applied separately.
Also went through the TODOs. Below is what is done and what remains.
File and line references are after applying all 19 patches.
Done:
- 0011: Included the earlier pg_buffercache fix to process one buffer
at a time in pg_buffercache_os_pages.
- 0012: Recompute MaxProportionalPins when a backend processes a resize
barrier, so the pin limit follows the new pool size.
- 0013: Retained the corrected MADV_REMOVE shrink range and included
the grow-range change discussed above. Also included the configure
declaration checks for the madvise constants.
- 0014: Check BM_TAG_VALID during eviction. A buffer can already have
a lookup-table entry while its read is still in progress, so checking
only whether its contents are valid can miss it. Recheck under the
header lock and roll back if eviction fails. Test 006 checks rollback,
reader completion and successful retry.
003 already covers the fully-valid pinned-buffer case.
- 0015: Fixed the parallel-worker initialization and configuration-reset
cases discussed above.
- 0016: Covered the TODO to document unsupported platforms. The
documentation now says that the resize function reports an error
when resizable shared memory is unavailable. Also repaired the
garbled paragraph and the typo.
See doc/src/sgml/func/func-admin.sgml:163.
- 0016: Covered the numeric memory-overhead example TODO. Verified with
max_shared_buffers=1GB and 8kB blocks on the 64-bit build. The lookup
table and checkpoint sort array together allocate 9,976,780 bytes,
about 9.5MB. Got the same result with shared_buffers=32MB and 128MB.
See doc/src/sgml/config.sgml:1875.
- 0017: Covered the TODO to validate hints inside ReadRecentBuffer().
A buffer number saved before shrink may be outside the new pool.
The function now rejects it before accessing the descriptor,
allowing the caller to do a normal lookup.
See src/backend/storage/buffer/bufmgr.c:843.
- 0017: Covered the TODO to check pg_shmem_allocations after rollback.
Test 001 now compares allocation metadata for all five buffer-manager
structures instead of checking only buffer counts. Also disabled
huge pages so rounding does not hide the small allocation changes.
All 414 assertions pass. This checks allocation metadata, not whether
physical memory was released.
See src/test/buffermgr/t/001_resize_fault_tolerance.pl:290.
The snapshot helper is at line 48; huge_pages=off is at line 29.
- 0018: Included the earlier SIGHUP test-ordering fix from
v20260817-0009. Start the resizer after configuration reload so the
expansion-failure test does not attempt resize with the old target.
- 0019: Included the earlier mprotect diagnostic fix from
v20260908-0011. The messages distinguish failure to make the active
range writable from failure to protect the reserved tail.
Not done:
- The variable-naming changes raised earlier are still pending.
- Direct stale-hint coverage for 0017. The code check is included, but
there is no test that keeps a buffer-number hint across a shrink
which removes that buffer slot. Such a test should check that the
old hint is rejected and the caller finds the page by normal lookup.
The rollback tests do not exercise this case.
- SHOW shared_buffers still includes the pending target, so its output
cannot be passed to pg_size_bytes(). Should SHOW return a parseable
size and leave the pending target to pg_get_buffer_resize_status()?
- A failed grow may leave some pages allocated in the unused range.
MADV_POPULATE_WRITE can fail after populating part of the range, and
making it inaccessible again does not free those pages. I have not
tested this failure case or added cleanup for it.
Thanks,
Palak
Attachments:
[application/octet-stream] v20260917-0004-Pass-use_units-parameter-to-GucShowHook-functions.patch (15.6K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/3-v20260917-0004-Pass-use_units-parameter-to-GucShowHook-functions.patch)
download | inline diff:
From 8f77ca03a58e401d4b3e7291b3146ad82cca1169 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 25 Feb 2026 15:48:14 +0530
Subject: [PATCH] Pass use_units parameter to GucShowHook functions
ShowGUCOption() prints the value of a GUC variable with units if the variable
has a unit and the caller passes use_units as true. However when a GUC variable
has a show_hook associated with it, the show_hook does not receive use_units
parameter. Hence a show_hook cannot determine whether to show the value with
units or not. This was not a problem until now because all the GUC variables
with show_hook were either didn't have any units associated with them or their
show_hook used a fixed unit to print the value of the variable.
shared_buffers is a GUC variable whose value is shows in different units based
on the GUC variable's value. With shared buffer pool resizing, we want to show
the pending size of the buffer pool, if any with appropriate units. We do this
using a show_hook, which needs use_units input. This commit adds use_units
parameter to GucShowHook and passes it from ShowGUCOption(). The actual commit
which uses this new parameter will be in a followup commit.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Reported by: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
---
src/backend/access/transam/xlog.c | 8 ++++----
src/backend/commands/variable.c | 14 +++++++-------
src/backend/executor/instrument.c | 2 +-
src/backend/libpq/pqcomm.c | 8 ++++----
src/backend/utils/misc/guc.c | 10 +++++-----
src/include/access/xlog.h | 2 +-
src/include/utils/guc.h | 2 +-
src/include/utils/guc_hooks.h | 30 +++++++++++++++---------------
8 files changed, 38 insertions(+), 38 deletions(-)
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index bd5fa6d896b..d38d7b94664 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4975,7 +4975,7 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
/* guc hook */
const char *
-show_data_checksums(void)
+show_data_checksums(bool use_units)
{
return get_checksum_state_string(LocalDataChecksumState);
}
@@ -5211,7 +5211,7 @@ InitializeWalConsistencyChecking(void)
* GUC show_hook for archive_command
*/
const char *
-show_archive_command(void)
+show_archive_command(bool use_units)
{
if (XLogArchivingActive())
return XLogArchiveCommand;
@@ -5223,7 +5223,7 @@ show_archive_command(void)
* GUC show_hook for in_hot_standby
*/
const char *
-show_in_hot_standby(void)
+show_in_hot_standby(bool use_units)
{
/*
* We display the actual state based on shared memory, so that this GUC
@@ -5238,7 +5238,7 @@ show_in_hot_standby(void)
* GUC show_hook for effective_wal_level
*/
const char *
-show_effective_wal_level(void)
+show_effective_wal_level(bool use_units)
{
if (wal_level == WAL_LEVEL_MINIMAL)
return "minimal";
diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c
index 8afd252fc8c..491c8aa392d 100644
--- a/src/backend/commands/variable.c
+++ b/src/backend/commands/variable.c
@@ -389,7 +389,7 @@ assign_timezone(const char *newval, void *extra)
* show_timezone: GUC show_hook for timezone
*/
const char *
-show_timezone(void)
+show_timezone(bool use_units)
{
const char *tzn;
@@ -462,7 +462,7 @@ assign_log_timezone(const char *newval, void *extra)
* show_log_timezone: GUC show_hook for log_timezone
*/
const char *
-show_log_timezone(void)
+show_log_timezone(bool use_units)
{
const char *tzn;
@@ -674,7 +674,7 @@ assign_random_seed(double newval, void *extra)
}
const char *
-show_random_seed(void)
+show_random_seed(bool use_units)
{
return "unavailable";
}
@@ -1031,7 +1031,7 @@ assign_role(const char *newval, void *extra)
}
const char *
-show_role(void)
+show_role(bool use_units)
{
/*
* Check whether SET ROLE is active; if not return "none". This is a
@@ -1179,7 +1179,7 @@ assign_io_combine_limit(int newval, void *extra)
* GUC show_hook for data_directory_mode
*/
const char *
-show_data_directory_mode(void)
+show_data_directory_mode(bool use_units)
{
static char buf[12];
@@ -1191,7 +1191,7 @@ show_data_directory_mode(void)
* GUC show_hook for log_file_mode
*/
const char *
-show_log_file_mode(void)
+show_log_file_mode(bool use_units)
{
static char buf[12];
@@ -1203,7 +1203,7 @@ show_log_file_mode(void)
* GUC show_hook for unix_socket_permissions
*/
const char *
-show_unix_socket_permissions(void)
+show_unix_socket_permissions(bool use_units)
{
static char buf[12];
diff --git a/src/backend/executor/instrument.c b/src/backend/executor/instrument.c
index ffbcd572133..d4c8373336a 100644
--- a/src/backend/executor/instrument.c
+++ b/src/backend/executor/instrument.c
@@ -423,7 +423,7 @@ assign_timing_clock_source(int newval, void *extra)
}
const char *
-show_timing_clock_source(void)
+show_timing_clock_source(bool use_units)
{
switch (timing_clock_source)
{
diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c
index aaae7214f13..b69bff03f09 100644
--- a/src/backend/libpq/pqcomm.c
+++ b/src/backend/libpq/pqcomm.c
@@ -1972,7 +1972,7 @@ assign_tcp_keepalives_idle(int newval, void *extra)
* GUC show_hook for tcp_keepalives_idle
*/
const char *
-show_tcp_keepalives_idle(void)
+show_tcp_keepalives_idle(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -1995,7 +1995,7 @@ assign_tcp_keepalives_interval(int newval, void *extra)
* GUC show_hook for tcp_keepalives_interval
*/
const char *
-show_tcp_keepalives_interval(void)
+show_tcp_keepalives_interval(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -2018,7 +2018,7 @@ assign_tcp_keepalives_count(int newval, void *extra)
* GUC show_hook for tcp_keepalives_count
*/
const char *
-show_tcp_keepalives_count(void)
+show_tcp_keepalives_count(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
@@ -2041,7 +2041,7 @@ assign_tcp_user_timeout(int newval, void *extra)
* GUC show_hook for tcp_user_timeout
*/
const char *
-show_tcp_user_timeout(void)
+show_tcp_user_timeout(bool use_units)
{
/* See comments in assign_tcp_keepalives_idle */
static char nbuf[16];
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 774bbc9be5f..1a5a168bc1a 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -5381,7 +5381,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_bool *conf = &record->_bool;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
val = *conf->variable ? "on" : "off";
}
@@ -5392,7 +5392,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_int *conf = &record->_int;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
{
/*
@@ -5421,7 +5421,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_real *conf = &record->_real;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
{
double result = *conf->variable;
@@ -5446,7 +5446,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_string *conf = &record->_string;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else if (*conf->variable && **conf->variable)
val = *conf->variable;
else
@@ -5459,7 +5459,7 @@ ShowGUCOption(const struct config_generic *record, bool use_units)
const struct config_enum *conf = &record->_enum;
if (conf->show_hook)
- val = conf->show_hook();
+ val = conf->show_hook(use_units);
else
val = config_enum_lookup_by_value(record, *conf->variable);
}
diff --git a/src/include/access/xlog.h b/src/include/access/xlog.h
index 4dd98624204..ee15b4cbb16 100644
--- a/src/include/access/xlog.h
+++ b/src/include/access/xlog.h
@@ -255,7 +255,7 @@ extern bool DataChecksumsInProgressOn(void);
extern void SetDataChecksumsOnInProgress(void);
extern void SetDataChecksumsOn(void);
extern void SetDataChecksumsOff(void);
-extern const char *show_data_checksums(void);
+extern const char *show_data_checksums(bool use_units);
extern const char *get_checksum_state_string(uint32 state);
extern void InitLocalDataChecksumState(void);
extern void SetLocalDataChecksumState(uint32 data_checksum_version);
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 8057d7870ad..2a6e2ed18b3 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -192,7 +192,7 @@ typedef void (*GucRealAssignHook) (double newval, void *extra);
typedef void (*GucStringAssignHook) (const char *newval, void *extra);
typedef void (*GucEnumAssignHook) (int newval, void *extra);
-typedef const char *(*GucShowHook) (void);
+typedef const char *(*GucShowHook) (bool use_units);
/*
* Miscellaneous
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 6a76f8d5ed6..df048517a0f 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -28,7 +28,7 @@
extern bool check_application_name(char **newval, void **extra,
GucSource source);
extern void assign_application_name(const char *newval, void *extra);
-extern const char *show_archive_command(void);
+extern const char *show_archive_command(bool use_units);
extern bool check_autovacuum_work_mem(int *newval, void **extra,
GucSource source);
extern bool check_vacuum_buffer_usage_limit(int *newval, void **extra,
@@ -46,7 +46,7 @@ extern void assign_client_encoding(const char *newval, void *extra);
extern bool check_cluster_name(char **newval, void **extra, GucSource source);
extern bool check_commit_ts_buffers(int *newval, void **extra,
GucSource source);
-extern const char *show_data_directory_mode(void);
+extern const char *show_data_directory_mode(bool use_units);
extern bool check_datestyle(char **newval, void **extra, GucSource source);
extern void assign_datestyle(const char *newval, void *extra);
extern bool check_debug_io_direct(char **newval, void **extra, GucSource source);
@@ -61,11 +61,11 @@ extern bool check_default_text_search_config(char **newval, void **extra, GucSou
extern void assign_default_text_search_config(const char *newval, void *extra);
extern bool check_default_with_oids(bool *newval, void **extra,
GucSource source);
-extern const char *show_effective_wal_level(void);
+extern const char *show_effective_wal_level(bool use_units);
extern bool check_huge_page_size(int *newval, void **extra, GucSource source);
extern void assign_io_method(int newval, void *extra);
extern bool check_io_max_concurrency(int *newval, void **extra, GucSource source);
-extern const char *show_in_hot_standby(void);
+extern const char *show_in_hot_standby(bool use_units);
extern bool check_locale_messages(char **newval, void **extra, GucSource source);
extern void assign_locale_messages(const char *newval, void *extra);
extern bool check_locale_monetary(char **newval, void **extra, GucSource source);
@@ -77,11 +77,11 @@ extern void assign_locale_time(const char *newval, void *extra);
extern bool check_log_destination(char **newval, void **extra,
GucSource source);
extern void assign_log_destination(const char *newval, void *extra);
-extern const char *show_log_file_mode(void);
+extern const char *show_log_file_mode(bool use_units);
extern bool check_log_stats(bool *newval, void **extra, GucSource source);
extern bool check_log_timezone(char **newval, void **extra, GucSource source);
extern void assign_log_timezone(const char *newval, void *extra);
-extern const char *show_log_timezone(void);
+extern const char *show_log_timezone(bool use_units);
extern void assign_maintenance_io_concurrency(int newval, void *extra);
extern void assign_io_max_combine_limit(int newval, void *extra);
extern void assign_io_combine_limit(int newval, void *extra);
@@ -97,7 +97,7 @@ extern bool check_primary_slot_name(char **newval, void **extra,
GucSource source);
extern bool check_random_seed(double *newval, void **extra, GucSource source);
extern void assign_random_seed(double newval, void *extra);
-extern const char *show_random_seed(void);
+extern const char *show_random_seed(bool use_units);
extern bool check_recovery_prefetch(int *new_value, void **extra,
GucSource source);
extern void assign_recovery_prefetch(int new_value, void *extra);
@@ -118,7 +118,7 @@ extern bool check_recovery_target_xid(char **newval, void **extra,
extern void assign_recovery_target_xid(const char *newval, void *extra);
extern bool check_role(char **newval, void **extra, GucSource source);
extern void assign_role(const char *newval, void *extra);
-extern const char *show_role(void);
+extern const char *show_role(bool use_units);
extern bool check_restrict_nonsystem_relation_kind(char **newval, void **extra,
GucSource source);
extern void assign_restrict_nonsystem_relation_kind(const char *newval, void *extra);
@@ -143,32 +143,32 @@ extern void assign_synchronous_commit(int newval, void *extra);
extern void assign_syslog_facility(int newval, void *extra);
extern void assign_syslog_ident(const char *newval, void *extra);
extern void assign_tcp_keepalives_count(int newval, void *extra);
-extern const char *show_tcp_keepalives_count(void);
+extern const char *show_tcp_keepalives_count(bool use_units);
extern void assign_tcp_keepalives_idle(int newval, void *extra);
-extern const char *show_tcp_keepalives_idle(void);
+extern const char *show_tcp_keepalives_idle(bool use_units);
extern void assign_tcp_keepalives_interval(int newval, void *extra);
-extern const char *show_tcp_keepalives_interval(void);
+extern const char *show_tcp_keepalives_interval(bool use_units);
extern void assign_tcp_user_timeout(int newval, void *extra);
-extern const char *show_tcp_user_timeout(void);
+extern const char *show_tcp_user_timeout(bool use_units);
extern bool check_temp_buffers(int *newval, void **extra, GucSource source);
extern bool check_temp_tablespaces(char **newval, void **extra,
GucSource source);
extern void assign_temp_tablespaces(const char *newval, void *extra);
extern bool check_timezone(char **newval, void **extra, GucSource source);
extern void assign_timezone(const char *newval, void *extra);
-extern const char *show_timezone(void);
+extern const char *show_timezone(bool use_units);
extern bool check_timezone_abbreviations(char **newval, void **extra,
GucSource source);
extern void assign_timezone_abbreviations(const char *newval, void *extra);
extern void assign_timing_clock_source(int newval, void *extra);
extern bool check_timing_clock_source(int *newval, void **extra, GucSource source);
-extern const char *show_timing_clock_source(void);
+extern const char *show_timing_clock_source(bool use_units);
extern bool check_transaction_buffers(int *newval, void **extra, GucSource source);
extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source);
extern bool check_transaction_isolation(int *newval, void **extra, GucSource source);
extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source);
extern void assign_transaction_timeout(int newval, void *extra);
-extern const char *show_unix_socket_permissions(void);
+extern const char *show_unix_socket_permissions(bool use_units);
extern bool check_wal_buffers(int *newval, void **extra, GucSource source);
extern bool check_wal_consistency_checking(char **newval, void **extra,
GucSource source);
[application/octet-stream] v20260917-0003-Add-a-view-to-read-contents-of-shared-buffer-lookup-table.patch (14.6K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/4-v20260917-0003-Add-a-view-to-read-contents-of-shared-buffer-lookup-table.patch)
download | inline diff:
From af0751483517539e035cbe6b5a81ba711118a2b4 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 25 Aug 2025 19:23:50 +0530
Subject: [PATCH] Add a view to read contents of shared buffer lookup table
The view exposes the contents of the shared buffer lookup table for
debugging, testing and investigation.
This helped me in debugging issues where the buffer descriptor array and
buffer lookup table were out of sync; either the buffer lookup table had
a mapping page->buffer which wasn't present in the buffer descriptor
array or a page in the buffer descriptor array didn't have corresponding
entry in the buffer lookup table. pg_buffercache doesn't help with those
kind of issues. Also doing that under the debugger in very painful.
I intend to keep this patch while the rest of the code matures. If it is
found useful as a debugging tool, we may consider make it committable
and commit it.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../expected/pg_buffercache.out | 39 ++++++++
.../pg_buffercache--1.5--1.6.sql | 24 +++++
contrib/pg_buffercache/pg_buffercache_pages.c | 17 ++++
contrib/pg_buffercache/sql/pg_buffercache.sql | 20 +++++
doc/src/sgml/system-views.sgml | 89 +++++++++++++++++++
src/backend/storage/buffer/buf_table.c | 58 ++++++++++++
src/include/storage/buf_internals.h | 2 +
7 files changed, 249 insertions(+)
diff --git a/contrib/pg_buffercache/expected/pg_buffercache.out b/contrib/pg_buffercache/expected/pg_buffercache.out
index c52a8491ff9..1ce5a970334 100644
--- a/contrib/pg_buffercache/expected/pg_buffercache.out
+++ b/contrib/pg_buffercache/expected/pg_buffercache.out
@@ -33,6 +33,26 @@ SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
t
(1 row)
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -46,6 +66,10 @@ SELECT * FROM pg_buffercache_summary();
ERROR: permission denied for function pg_buffercache_summary
SELECT * FROM pg_buffercache_usage_counts();
ERROR: permission denied for function pg_buffercache_usage_counts
+SELECT * FROM pg_buffercache_lookup_table_entries();
+ERROR: permission denied for function pg_buffercache_lookup_table_entries
+SELECT * FROM pg_buffercache_lookup_table;
+ERROR: permission denied for view pg_buffercache_lookup_table
RESET role;
-- Check that pg_monitor is allowed to query view / function
SET ROLE pg_monitor;
@@ -81,6 +105,21 @@ FROM pg_buffercache_pages() AS p
LIMIT 1;
ERROR: function return row and query-specified return row do not match
DETAIL: Returned type boolean at ordinal position 7, but query expects text.
+RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+ ?column?
+----------
+ t
+(1 row)
+
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+ ?column?
+----------
+ t
+(1 row)
+
RESET role;
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
index 458f054a691..9bf58567878 100644
--- a/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
+++ b/contrib/pg_buffercache/pg_buffercache--1.5--1.6.sql
@@ -44,3 +44,27 @@ CREATE FUNCTION pg_buffercache_evict_all(
OUT buffers_skipped int4)
AS 'MODULE_PATHNAME', 'pg_buffercache_evict_all'
LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Add the buffer lookup table function
+CREATE FUNCTION pg_buffercache_lookup_table_entries(
+ OUT tablespace oid,
+ OUT database oid,
+ OUT relfilenode oid,
+ OUT forknum int2,
+ OUT blocknum int8,
+ OUT bufferid int4)
+RETURNS SETOF record
+AS 'MODULE_PATHNAME', 'pg_buffercache_lookup_table_entries'
+LANGUAGE C PARALLEL SAFE VOLATILE;
+
+-- Create a view for convenient access.
+CREATE VIEW pg_buffercache_lookup_table AS
+ SELECT * FROM pg_buffercache_lookup_table_entries();
+
+-- Don't want these to be available to public.
+REVOKE ALL ON FUNCTION pg_buffercache_lookup_table_entries() FROM PUBLIC;
+REVOKE ALL ON pg_buffercache_lookup_table FROM PUBLIC;
+
+-- Grant access to monitoring role.
+GRANT EXECUTE ON FUNCTION pg_buffercache_lookup_table_entries() TO pg_read_all_stats;
+GRANT SELECT ON pg_buffercache_lookup_table TO pg_read_all_stats;
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 510455998aa..9512f1efa2f 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -77,6 +77,7 @@ PG_FUNCTION_INFO_V1(pg_buffercache_evict_all);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_relation);
PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_all);
+PG_FUNCTION_INFO_V1(pg_buffercache_lookup_table_entries);
/* Only need to touch memory once per backend process lifetime */
@@ -922,3 +923,19 @@ pg_buffercache_mark_dirty_all(PG_FUNCTION_ARGS)
PG_RETURN_DATUM(result);
}
+
+/*
+ * Return lookup table content as a set of records.
+ */
+Datum
+pg_buffercache_lookup_table_entries(PG_FUNCTION_ARGS)
+{
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+ /* Fill the tuplestore */
+ BufTableGetContents(rsinfo->setResult, rsinfo->setDesc);
+
+ return (Datum) 0;
+}
diff --git a/contrib/pg_buffercache/sql/pg_buffercache.sql b/contrib/pg_buffercache/sql/pg_buffercache.sql
index be89b5f5a3a..d0726b2b2f9 100644
--- a/contrib/pg_buffercache/sql/pg_buffercache.sql
+++ b/contrib/pg_buffercache/sql/pg_buffercache.sql
@@ -18,6 +18,18 @@ from pg_buffercache_summary();
SELECT count(*) > 0 FROM pg_buffercache_usage_counts() WHERE buffers >= 0;
+-- Test the buffer lookup table function and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table_entries();
+
+-- Check that pg_buffercache_lookup_table view works and count is <= shared_buffers
+select count(*) <= (select setting::bigint
+ from pg_settings
+ where name = 'shared_buffers')
+from pg_buffercache_lookup_table;
+
-- Check that the functions / views can't be accessed by default. To avoid
-- having to create a dedicated user, use the pg_database_owner pseudo-role.
SET ROLE pg_database_owner;
@@ -26,6 +38,8 @@ SELECT * FROM pg_buffercache_os_pages;
SELECT * FROM pg_buffercache_pages() AS p (wrong int);
SELECT * FROM pg_buffercache_summary();
SELECT * FROM pg_buffercache_usage_counts();
+SELECT * FROM pg_buffercache_lookup_table_entries();
+SELECT * FROM pg_buffercache_lookup_table;
RESET role;
-- Check that pg_monitor is allowed to query view / function
@@ -42,6 +56,12 @@ FROM pg_buffercache_pages() AS p
LIMIT 1;
RESET role;
+-- Check that pg_read_all_stats is allowed to query buffer lookup table
+SET ROLE pg_read_all_stats;
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table_entries();
+SELECT count(*) >= 0 FROM pg_buffercache_lookup_table;
+RESET role;
+
------
---- Test pg_buffercache_evict* and pg_buffercache_mark_dirty* functions
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 5ea19d68622..6b905498337 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -71,6 +71,11 @@
<entry>backend memory contexts</entry>
</row>
+ <row>
+ <entry><link linkend="view-pg-buffer-lookup-table"><structname>pg_buffer_lookup_table</structname></link></entry>
+ <entry>shared buffer lookup table</entry>
+ </row>
+
<row>
<entry><link linkend="view-pg-config"><structname>pg_config</structname></link></entry>
<entry>compile-time configuration parameters</entry>
@@ -929,6 +934,90 @@ AND c1.path[c2.level] = c2.path[c2.level];
</para>
</sect1>
+ <sect1 id="view-pg-buffer-lookup-table">
+ <title><structname>pg_buffer_lookup_table</structname></title>
+ <indexterm>
+ <primary>pg_buffer_lookup_table</primary>
+ </indexterm>
+ <para>
+ The <structname>pg_buffer_lookup_table</structname> view exposes the current
+ contents of the shared buffer lookup table. Each row represents an entry in
+ the lookup table mapping a relation page to the ID of buffer in which it is
+ cached. The shared buffer lookup table is locked for a short duration while
+ reading so as to ensure consistency. This may affect performance if this view
+ is queried very frequently.
+ </para>
+ <table id="pg-buffer-lookup-table-view" xreflabel="pg_buffer_lookup_table">
+ <title><structname>pg_buffer_lookup_table</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>tablespace</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the tablespace containing the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>database</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database containing the relation (zero for shared relations)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relfilenode</structfield> <type>oid</type>
+ </para>
+ <para>
+ relfilenode identifying the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>forknum</structfield> <type>int2</type>
+ </para>
+ <para>
+ Fork number within the relation (see <xref linkend="storage-file-layout"/>)
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>blocknum</structfield> <type>int8</type>
+ </para>
+ <para>
+ Block number within the relation
+ </para></entry>
+ </row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>bufferid</structfield> <type>int4</type>
+ </para>
+ <para>
+ ID of the buffer caching the page
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ <para>
+ Access to this view is restricted to members of the
+ <literal>pg_read_all_stats</literal> role by default.
+ </para>
+ </sect1>
+
<sect1 id="view-pg-config">
<title><structname>pg_config</structname></title>
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 5c8ccbee13f..c82b71deaa6 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -21,7 +21,12 @@
*/
#include "postgres.h"
+#include "fmgr.h"
+#include "funcapi.h"
#include "storage/buf_internals.h"
+#include "utils/rel.h"
+#include "utils/builtins.h"
+#include "storage/lwlock.h"
#include "storage/subsystems.h"
/* entry for buffer lookup hashtable */
@@ -167,3 +172,56 @@ BufTableDelete(BufferTag *tagPtr, uint32 hashcode)
if (!result) /* shouldn't happen */
elog(ERROR, "shared buffer hash table corrupted");
}
+
+/*
+ * BufTableGetContents
+ * Fill the given tuplestore with contents of the shared buffer lookup table
+ *
+ * This function is used by pg_buffercache extension to expose buffer lookup
+ * table contents via SQL. The caller is responsible for setting up the
+ * tuplestore and result set info.
+ */
+void
+BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc)
+{
+/* Expected number of attributes of the buffer lookup table entry. */
+#define BUFTABLE_CONTENTS_COLS 6
+
+ HASH_SEQ_STATUS hstat;
+ BufferLookupEnt *ent;
+ Datum values[BUFTABLE_CONTENTS_COLS];
+ bool nulls[BUFTABLE_CONTENTS_COLS];
+ int i;
+
+ memset(nulls, 0, sizeof(nulls));
+
+ Assert(tupdesc->natts == BUFTABLE_CONTENTS_COLS);
+
+ /*
+ * Lock all buffer mapping partitions to ensure a consistent view of the
+ * hash table during the scan. Must grab LWLocks in partition-number order
+ * to avoid LWLock deadlock.
+ */
+ for (i = 0; i < NUM_BUFFER_PARTITIONS; i++)
+ LWLockAcquire(BufMappingPartitionLockByIndex(i), LW_SHARED);
+
+ hash_seq_init(&hstat, SharedBufHash);
+ while ((ent = (BufferLookupEnt *) hash_seq_search(&hstat)) != NULL)
+ {
+ values[0] = ObjectIdGetDatum(ent->key.spcOid);
+ values[1] = ObjectIdGetDatum(ent->key.dbOid);
+ values[2] = ObjectIdGetDatum(ent->key.relNumber);
+ values[3] = ObjectIdGetDatum(ent->key.forkNum);
+ values[4] = Int64GetDatum(ent->key.blockNum);
+ values[5] = Int32GetDatum(ent->id);
+
+ tuplestore_putvalues(tupstore, tupdesc, values, nulls);
+ }
+
+ /*
+ * Release all buffer mapping partition locks in the reverse order so as
+ * to avoid LWLock deadlock.
+ */
+ for (i = NUM_BUFFER_PARTITIONS - 1; i >= 0; i--)
+ LWLockRelease(BufMappingPartitionLockByIndex(i));
+}
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index a7606c6e92b..678065b3de2 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -29,6 +29,7 @@
#include "storage/spin.h"
#include "utils/relcache.h"
#include "utils/resowner.h"
+#include "utils/tuplestore.h"
/*
* Buffer state is a single 64-bit variable where following data is combined.
@@ -599,6 +600,7 @@ extern uint32 BufTableHashCode(BufferTag *tagPtr);
extern int BufTableLookup(BufferTag *tagPtr, uint32 hashcode);
extern int BufTableInsert(BufferTag *tagPtr, uint32 hashcode, int buf_id);
extern void BufTableDelete(BufferTag *tagPtr, uint32 hashcode);
+extern void BufTableGetContents(Tuplestorestate *tupstore, TupleDesc tupdesc);
/* localbuf.c */
extern bool PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount);
[application/octet-stream] v20260917-0001-Add-BgBufferSync-sanity-Asserts.patch (2.1K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/5-v20260917-0001-Add-BgBufferSync-sanity-Asserts.patch)
download | inline diff:
From 682e6c5e86c06e882b13d3802115760e6a8617c7 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 21 Jul 2026 16:16:48 +0530
Subject: [PATCH] Add BgBufferSync sanity Asserts
Assert that BgBufferSync() is only ever invoked from the background
writer process, matching the actual caller in BackgroundWriterMain.
Also move the existing Assert(strategy_delta >= 0) to the end of the
enclosing block so that the surrounding elog(DEBUG2) messages get a
chance to reach the server log before the Assert fires, easing diagnosis
if the Assert ever fails.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
src/backend/storage/buffer/bufmgr.c | 15 +++++++++++++--
1 file changed, 13 insertions(+), 2 deletions(-)
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 169829eb020..536c92d47e4 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3894,6 +3894,8 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ Assert(AmBackgroundWriterProcess());
+
/*
* Find out where the clock-sweep currently is, and how many buffer
* allocations have happened since our last call.
@@ -3929,8 +3931,6 @@ BgBufferSync(WritebackContext *wb_context)
strategy_delta = strategy_buf_id - prev_strategy_buf_id;
strategy_delta += (long) passes_delta * NBuffers;
- Assert(strategy_delta >= 0);
-
if ((int32) (next_passes - strategy_passes) > 0)
{
/* we're one pass ahead of the strategy point */
@@ -3970,6 +3970,17 @@ BgBufferSync(WritebackContext *wb_context)
next_passes = strategy_passes;
bufs_to_lap = NBuffers;
}
+
+ /*
+ * We do not expect the current strategy point to be behind the
+ * previous one.
+ *
+ * If this Assert fails, we would have the debug messages printed in
+ * the server error log at appropriate debug level. Hence Asserting
+ * here provides minor convenience compared to Asserting right after
+ * calculating the difference.
+ */
+ Assert(strategy_delta >= 0);
}
else
{
[application/octet-stream] v20260917-0005-PID-of-the-backend-process-backing-the-BackgroundPsql-session.patch (3.0K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/6-v20260917-0005-PID-of-the-backend-process-backing-the-BackgroundPsql-session.patch)
download | inline diff:
From 8ef74820a2743885e4a4801090e8aabae01bd197 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 23 Jun 2026 09:19:06 +0530
Subject: [PATCH] PID of the backend process backing the BackgroundPsql session
Many tests which use background psql sessions also fetch the pid of the backend
process as a separate step. They end up maintaing a separate variable for the
pid and pass it to the subroutines along with the BackgroundPsql object when
needed. Further the PID can not be fetched if the backend is blocked in a query
or an injection point or if the backend has died. This can be avoided by
proactively fetching the PID of the backend process when the BackgroundPsql is
started and storing it in the BackgroundPsql object. The PID can then be fetched
from the BackgroundPsql object when needed.
The pid is reset when the BackgroundPsql session is finished. Many of the tests
which use BackgroundPsql explicitly call {run}->finish to finish the session
when the backend is gone. They will have stale PID in the BackgroundPsql object.
The commit introduces a finish method in BackgroundPsql which will finish the
session and reset the PID.
Note to the reviewer:
This commit adds code to set the PID proactively and also the finish method.
These will be used in a future buffer resize test. But I have not modified the
existing tests to use the new capabilities. If we find this change useful, we
can modify the existing tests to use the new capabilities before committing this
change.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 29 ++++++++++++++++++-
1 file changed, 28 insertions(+), 1 deletion(-)
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index d7797225451..699334320d9 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -173,6 +173,14 @@ sub wait_connect
$self->{stderr} = '';
die "psql startup timed out" if $self->{timeout}->is_expired;
+
+ # Many tests which use background psql sessions also fetch the pid of the
+ # backend process, so we capture it here. The Callers that need the pid
+ # after a blocking query or after the backend has died can read it from
+ # $self->{backend_pid}.
+ my $pid = $self->query('SELECT pg_backend_pid()', verbose => 0);
+ chomp $pid;
+ $self->{backend_pid} = $pid;
}
=pod
@@ -190,6 +198,25 @@ sub quit
$self->{stdin} .= "\\q\n";
+ return $self->finish;
+}
+
+=pod
+
+=item $session->finish
+
+Reap the underlying IPC::Run handle without sending \q. Intended for
+sessions whose psql process has already exited (e.g. after the server
+terminated the backend or the client connection was killed).
+
+=cut
+
+sub finish
+{
+ my ($self) = @_;
+
+ $self->{backend_pid} = undef;
+
return $self->{run}->finish;
}
@@ -212,7 +239,7 @@ sub reconnect_and_clear
{
$self->{stdin} .= "\\q\n";
}
- $self->{run}->finish;
+ $self->finish;
# restart
$self->{run}->run();
[application/octet-stream] v20260917-0002-Decouple-GUC-shared_buffers-and-size-of-the-buffer-pool.patch (18.6K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/7-v20260917-0002-Decouple-GUC-shared_buffers-and-size-of-the-buffer-pool.patch)
download | inline diff:
From 264428d9d50fae622c4504ec3f5d832d7e305594 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Thu, 4 Jun 2026 15:10:56 +0530
Subject: [PATCH] Decouple GUC shared_buffers and size of the buffer pool
A single variable NBuffers holds the value of the GUC 'shared_buffers' and the
size of the buffer pool in number of buffers. With the introduction of resizable
shared buffer pool feature, the value of 'shared_buffers' GUC and the size of
the buffer pool can be different. This commit prepares for the same by
decoupling the two. A new variable NBuffersGUC holds the value of the GUC
'shared_buffers'. The variable NBuffers continues to hold the size of the
buffer pool. It is set to the value of NBuffersGUC during initialization.
Because of this decoupling, the code which references NBuffers as the value of
the GUC 'shared_buffers' is clearly differentiated from the code which
references NBuffers as the size of the buffer pool.
Also change comments to use term "size of the buffer pool" or "number of
buffers" instead of referencing variable NBuffers.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
---
src/backend/access/heap/heapam.c | 15 ++++++++-------
src/backend/access/transam/slru.c | 2 +-
src/backend/access/transam/xlog.c | 8 ++++----
src/backend/optimizer/path/costsize.c | 8 ++++----
src/backend/postmaster/checkpointer.c | 11 ++++++-----
src/backend/storage/aio/aio_init.c | 2 +-
src/backend/storage/buffer/buf_init.c | 17 +++++++++++++----
src/backend/storage/buffer/buf_table.c | 14 ++++++++------
src/backend/storage/buffer/bufmgr.c | 13 +++++++------
src/backend/storage/buffer/freelist.c | 7 ++++---
src/backend/utils/init/globals.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 2 +-
src/include/miscadmin.h | 2 +-
src/include/storage/buf.h | 2 +-
src/include/storage/buf_internals.h | 4 ++--
src/include/storage/bufmgr.h | 3 ++-
16 files changed, 64 insertions(+), 48 deletions(-)
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 8b488cfd8f6..8eb84bd51d6 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -383,13 +383,14 @@ initscan(HeapScanDesc scan, ScanKey key, bool keep_startblock)
scan->rs_nblocks = RelationGetNumberOfBlocks(scan->rs_base.rs_rd);
/*
- * If the table is large relative to NBuffers, use a bulk-read access
- * strategy and enable synchronized scanning (see syncscan.c). Although
- * the thresholds for these features could be different, we make them the
- * same so that there are only two behaviors to tune rather than four.
- * (However, some callers need to be able to disable one or both of these
- * behaviors, independently of the size of the table; also there is a GUC
- * variable that can disable synchronized scanning.)
+ * If the table is large relative to the size of the buffer pool, use a
+ * bulk-read access strategy and enable synchronized scanning (see
+ * syncscan.c). Although the thresholds for these features could be
+ * different, we make them the same so that there are only two behaviors
+ * to tune rather than four. (However, some callers need to be able to
+ * disable one or both of these behaviors, independently of the size of
+ * the table; also there is a GUC variable that can disable synchronized
+ * scanning.)
*
* Note that table_block_parallelscan_initialize has a very similar test;
* if you change this, consider changing that one, too.
diff --git a/src/backend/access/transam/slru.c b/src/backend/access/transam/slru.c
index 47dd52d6749..b47ace37974 100644
--- a/src/backend/access/transam/slru.c
+++ b/src/backend/access/transam/slru.c
@@ -236,7 +236,7 @@ SimpleLruAutotuneBuffers(int divisor, int max)
{
return Min(max - (max % SLRU_BANK_SIZE),
Max(SLRU_BANK_SIZE,
- NBuffers / divisor - (NBuffers / divisor) % SLRU_BANK_SIZE));
+ NBuffersGUC / divisor - (NBuffersGUC / divisor) % SLRU_BANK_SIZE));
}
/*
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index b23d8bbbdad..bd5fa6d896b 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -5017,14 +5017,14 @@ GetFakeLSNForUnloggedRel(void)
* and a minimum of 8 blocks (which was the default value prior to PostgreSQL
* 9.1, when auto-tuning was added).
*
- * This should not be called until NBuffers has received its final value.
+ * This should not be called until NBuffersGUC has received its final value.
*/
static int
XLOGChooseNumBuffers(void)
{
int xbuffers;
- xbuffers = NBuffers / 32;
+ xbuffers = NBuffersGUC / 32;
if (xbuffers > (wal_segment_size / XLOG_BLCKSZ))
xbuffers = (wal_segment_size / XLOG_BLCKSZ);
if (xbuffers < 8)
@@ -5297,8 +5297,8 @@ XLOGShmemRequest(void *arg)
/*
* If the value of wal_buffers is -1, use the preferred auto-tune value.
* This isn't an amazingly clean place to do this, but we must wait till
- * NBuffers has received its final value, and must do it before using the
- * value of XLOGbuffers to do anything important.
+ * NBuffersGUC has received its final value, and must do it before using
+ * the value of XLOGbuffers to do anything important.
*
* We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
* However, if the DBA explicitly set wal_buffers = -1 in the config file,
diff --git a/src/backend/optimizer/path/costsize.c b/src/backend/optimizer/path/costsize.c
index ac523ecf9a8..8be50631e9c 100644
--- a/src/backend/optimizer/path/costsize.c
+++ b/src/backend/optimizer/path/costsize.c
@@ -19,10 +19,10 @@
* is normally considerably less than random_page_cost. (However, if the
* database is fully cached in RAM, it is reasonable to set them equal.)
*
- * We also use a rough estimate "effective_cache_size" of the number of
- * disk pages in Postgres + OS-level disk cache. (We can't simply use
- * NBuffers for this purpose because that would ignore the effects of
- * the kernel's disk cache.)
+ * We also use a rough estimate "effective_cache_size" of the number of disk
+ * pages in Postgres + OS-level disk cache. (We can't simply use size of the
+ * buffer pool for this purpose because that would ignore the effects of the
+ * kernel's disk cache.)
*
* Obviously, taking constants for these values is an oversimplification,
* but it's tough enough to get any useful estimates even at this level of
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index f6351f1eb10..0db3c9e1a38 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -963,12 +963,13 @@ CheckpointerShmemRequest(void *arg)
Size size;
/*
- * The size of the requests[] array is arbitrarily set equal to NBuffers.
- * But there is a cap of MAX_CHECKPOINT_REQUESTS to prevent accumulating
- * too many checkpoint requests in the ring buffer.
+ * The size of the requests[] array is arbitrarily set equal to the
+ * initial size of buffer pool. But there is a cap of
+ * MAX_CHECKPOINT_REQUESTS to prevent accumulating too many checkpoint
+ * requests in the ring buffer.
*/
size = offsetof(CheckpointerShmemStruct, requests);
- size = add_size(size, mul_size(Min(NBuffers,
+ size = add_size(size, mul_size(Min(NBuffersGUC,
MAX_CHECKPOINT_REQUESTS),
sizeof(CheckpointerRequest)));
ShmemRequestStruct(.name = "Checkpointer Data",
@@ -985,7 +986,7 @@ static void
CheckpointerShmemInit(void *arg)
{
SpinLockInit(&CheckpointerShmem->ckpt_lck);
- CheckpointerShmem->max_requests = Min(NBuffers, MAX_CHECKPOINT_REQUESTS);
+ CheckpointerShmem->max_requests = Min(NBuffersGUC, MAX_CHECKPOINT_REQUESTS);
CheckpointerShmem->head = CheckpointerShmem->tail = 0;
ConditionVariableInit(&CheckpointerShmem->start_cv);
ConditionVariableInit(&CheckpointerShmem->done_cv);
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index de50e6a8a31..81825dd1b73 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -109,7 +109,7 @@ AioChooseMaxConcurrency(void)
/* Similar logic to LimitAdditionalPins() */
max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
- max_proportional_pins = NBuffers / max_backends;
+ max_proportional_pins = NBuffersGUC / max_backends;
max_proportional_pins = Max(max_proportional_pins, 1);
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 1407c930c56..9ddf6551fcd 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -77,21 +77,21 @@ static void
BufferManagerShmemRequest(void *arg)
{
ShmemRequestStruct(.name = "Buffer Descriptors",
- .size = NBuffers * sizeof(BufferDescPadded),
+ .size = NBuffersGUC * sizeof(BufferDescPadded),
/* Align descriptors to a cacheline boundary. */
.alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &BufferDescriptors,
);
ShmemRequestStruct(.name = "Buffer Blocks",
- .size = NBuffers * (Size) BLCKSZ,
+ .size = NBuffersGUC * (Size) BLCKSZ,
/* Align buffer pool on IO page size boundary. */
.alignment = PG_IO_ALIGN_SIZE,
.ptr = (void **) &BufferBlocks,
);
ShmemRequestStruct(.name = "Buffer IO Condition Variables",
- .size = NBuffers * sizeof(ConditionVariableMinimallyPadded),
+ .size = NBuffersGUC * sizeof(ConditionVariableMinimallyPadded),
/* Align descriptors to a cacheline boundary. */
.alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &BufferIOCVArray,
@@ -105,7 +105,7 @@ BufferManagerShmemRequest(void *arg)
* painful.
*/
ShmemRequestStruct(.name = "Checkpoint BufferIds",
- .size = NBuffers * sizeof(CkptSortItem),
+ .size = NBuffersGUC * sizeof(CkptSortItem),
.ptr = (void **) &CkptBufferIds,
);
}
@@ -119,6 +119,12 @@ BufferManagerShmemRequest(void *arg)
static void
BufferManagerShmemInit(void *arg)
{
+ /*
+ * Set the size of the buffer pool, now that it's allocated and ready to
+ * be initialized.
+ */
+ NBuffers = NBuffersGUC;
+
/*
* Initialize all the buffer headers.
*/
@@ -147,6 +153,9 @@ BufferManagerShmemInit(void *arg)
static void
BufferManagerShmemAttach(void *arg)
{
+ /* Update the size of the buffer pool. */
+ NBuffers = NBuffersGUC;
+
/* Initialize per-backend file flush context */
WritebackContextInit(&BackendWritebackContext,
&backend_flush_after);
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index 347bf267d73..5c8ccbee13f 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -42,7 +42,7 @@ const ShmemCallbacks BufTableShmemCallbacks = {
/*
* Register shmem hash table for mapping buffers.
- * size is the desired hash table size (possibly more than NBuffers)
+ * size is the desired hash table size (possibly more than the size of the buffer pool).
*/
void
BufTableShmemRequest(void *arg)
@@ -54,12 +54,14 @@ BufTableShmemRequest(void *arg)
*
* Since we can't tolerate running out of lookup table entries, we must be
* sure to specify an adequate table size here. The maximum steady-state
- * usage is of course NBuffers entries, but BufferAlloc() tries to insert
- * a new entry before deleting the old. In principle this could be
- * happening in each partition concurrently, so we could need as many as
- * NBuffers + NUM_BUFFER_PARTITIONS entries.
+ * usage is of course as many entries as the number of buffers in the
+ * pool, but BufferAlloc() tries to insert a new entry before deleting the
+ * old. In principle this could be happening in each partition
+ * concurrently, so we could need as many as (number of buffers in the
+ * pool) + NUM_BUFFER_PARTITIONS entries. Since we are still requesting
+ * shared memory, use the GUC value instead of the actual size.
*/
- size = NBuffers + NUM_BUFFER_PARTITIONS;
+ size = NBuffersGUC + NUM_BUFFER_PARTITIONS;
ShmemRequestHash(.name = "Shared Buffer Lookup Table",
.nelems = size,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 536c92d47e4..9f86c3320e1 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -223,6 +223,7 @@ int io_max_combine_limit = DEFAULT_IO_COMBINE_LIMIT;
int checkpoint_flush_after = DEFAULT_CHECKPOINT_FLUSH_AFTER;
int bgwriter_flush_after = DEFAULT_BGWRITER_FLUSH_AFTER;
int backend_flush_after = DEFAULT_BACKEND_FLUSH_AFTER;
+int NBuffers = 0; /* number of buffers in the buffer pool */
/* local state for LockBufferForCleanup */
static BufferDesc *PinCountWaitBuf = NULL;
@@ -239,11 +240,11 @@ static BufferDesc *PinCountWaitBuf = NULL;
* and, if so, in what mode.
*
*
- * To avoid - as we used to - requiring an array with NBuffers entries to keep
- * track of local buffers, we use a small sequentially searched array
- * (PrivateRefCountArrayKeys, with the corresponding data stored in
- * PrivateRefCountArray) and an overflow hash table (PrivateRefCountHash) to
- * keep track of backend local pins.
+ * To avoid - as we used to - requiring an array, with as many entries as the
+ * size of buffer pool, to keep track of local buffers, we use a small
+ * sequentially searched array (PrivateRefCountArrayKeys, with the corresponding
+ * data stored in PrivateRefCountArray) and an overflow hash table
+ * (PrivateRefCountHash) to keep track of backend local pins.
*
* Until no more than REFCOUNT_ARRAY_ENTRIES buffers are pinned at once, all
* refcounts are kept track of in the array; after that, new array entries
@@ -3642,7 +3643,7 @@ BufferSync(int flags)
set_bits, 0,
0);
- /* Check for barrier events in case NBuffers is large. */
+ /* Check for barrier events in case the buffer pool is large. */
if (ProcSignalBarrierPending)
ProcessProcSignalBarrier();
}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index fdb5bad7910..4d5ee52ddc0 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -37,7 +37,8 @@ typedef struct
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo NBuffers.
+ * get an actual buffer, it needs to be used modulo size of the buffer
+ * pool.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -522,10 +523,10 @@ GetAccessStrategyWithSize(BufferAccessStrategyType btype, int ring_size_kb)
if (ring_buffers == 0)
return NULL;
- /* Cap to 1/8th of shared_buffers */
+ /* Cap to 1/8th of number of buffers in the buffer pool. */
ring_buffers = Min(NBuffers / 8, ring_buffers);
- /* NBuffers should never be less than 16, so this shouldn't happen */
+ /* Buffer pool should always have more than 16 buffers. */
Assert(ring_buffers > 0);
/* Allocate the object and initialize all elements to zeroes */
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index bbd28d14d99..ccf845e87b9 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -141,7 +141,7 @@ int max_parallel_maintenance_workers = 2;
* MaxBackends is computed by PostmasterMain after modules have had a chance to
* register background workers.
*/
-int NBuffers = 16384;
+int NBuffersGUC = 16384;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index e2b48ea69e9..66ea4afe920 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2716,7 +2716,7 @@
{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
- variable => 'NBuffers',
+ variable => 'NBuffersGUC',
boot_val => '16384',
min => '16',
max => 'INT_MAX / 2',
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 0fc59af02b9..e2df51d275a 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -175,7 +175,7 @@ extern PGDLLIMPORT bool ExitOnAnyError;
extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
-extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersGUC;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
diff --git a/src/include/storage/buf.h b/src/include/storage/buf.h
index b21445522b1..a12d6b9082a 100644
--- a/src/include/storage/buf.h
+++ b/src/include/storage/buf.h
@@ -17,7 +17,7 @@
/*
* Buffer identifiers.
*
- * Zero is invalid, positive is the index of a shared buffer (1..NBuffers),
+ * Zero is invalid, positive is the index of a shared buffer (1..{size of shared buffer pool}),
* negative is the index of a local buffer (-1 .. -NLocBuffer).
*/
typedef int Buffer;
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index e4ff5619b79..a7606c6e92b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -135,8 +135,8 @@ StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2),
/*
* The maximum allowed value of usage_count represents a tradeoff between
- * accuracy and speed of the clock-sweep buffer management algorithm. A
- * large value (comparable to NBuffers) would approximate LRU semantics.
+ * accuracy and speed of the clock-sweep buffer management algorithm. A large
+ * value (comparable to the size of buffer pool) would approximate LRU semantics.
* But it can take as many as BM_MAX_USAGE_COUNT+1 complete cycles of the
* clock-sweep hand to find a free buffer, so in practice we don't want the
* value to be very large.
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 6837b35fc6d..f1f6e601f51 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -159,13 +159,14 @@ typedef struct ReadBuffersOperation ReadBuffersOperation;
typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
-extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int NBuffersGUC;
/* in bufmgr.c */
extern PGDLLIMPORT bool zero_damaged_pages;
extern PGDLLIMPORT int bgwriter_lru_maxpages;
extern PGDLLIMPORT double bgwriter_lru_multiplier;
extern PGDLLIMPORT bool track_io_timing;
+extern PGDLLIMPORT int NBuffers;
#define DEFAULT_EFFECTIVE_IO_CONCURRENCY 16
#define DEFAULT_MAINTENANCE_IO_CONCURRENCY 16
[application/octet-stream] v20260917-0006-Resizable-shared-memory-structures.patch (129.5K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/8-v20260917-0006-Resizable-shared-memory-structures.patch)
download | inline diff:
From 6d37c5bdeb727207b153a21fd56e65d315b5b969 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Tue, 17 Feb 2026 16:51:20 +0530
Subject: [PATCH] Resizable shared memory structures
Resizable shared memory structures can be requested by specifying a new
member ShmemStructOpts::maximum_size. At the startup or when the
structure is created, we reserve address space worth maximum_size in the
shared memory segment. It is expected that the subsystem which creates a
resizable structure would initialize only the memory worth its initial
size given by ShmemStructOpts::maximum_size when creating it. In an
mmap'ed memory, this should allocate memory worth only the initial size.
It should not allocate maximum_size worth of memory initially. As the
structure is resized using ShmemResizeStruct() memory is freed or
allocated in chunks of memory pages when shrinking and expanding the
structure respectively. Optional ShmemStructOpts::minimum_size
specification for resizable shared memory structures allows us to
enforce that a resizable structure cannot be shrunk below a certain
size. If minimum_size is not specified for a resizable structure, it the
minimum size defaults to 0. Two additional columns are added to
pg_shmem_allocations view to report minimum and maximum size
respectively.
The structure for which ShmemStructOpts::maximum_size is not specified
or is set to 0 is considered as a fixed size structure. Existing calls
to ShmemRequestStruct() requsting fixed size structures need no change.
For fixed-size structures, the minimum size and maximum size are set to
the initial size specified in ShmemRequestStructOpts::size. This makes
maximum size and minimum size reported in pg_shmem_allocations view
semantically consistent for both fixed-size and resizable structures. As
a side effect, a request which specifies same minimum_size and
maximum_size will be treated as a fixed size structure.
Minimum, maximum and initial sizes of main shared memory area
=============================================================
With the addition of resizable shared structures, the main shared memory
area has three sizes to be tracked:
- a. the minimum amount of shared memory that will always be needed
- b. the maximum amount of shared memory that may be allocated when all
the resizable structures have grown to their maximum size. It is also
is the size of the address space reserved for the main shared memory
area.
- c. the amount of memory allocated at the server startup i.e. initial
allocation.
mmap needs to use the maximum size to reserve enough address space to
accomodate all the shared structures. But a DBA may provision only the
memory required at the startup initially and increase or descrease the
memory provision as the demand changes.
These three sizes are reported as GUCs shared_memory_minimum_size,
shared_memory_maximum_size and shared_memory_initial_size respectively.
They replace the old GUC shared_memory_size which used to report both
size of the main shared memory area and amount of memory required in it.
Since these two things are not the same anymore, a single GUC is not
sufficient. These GUCs help DBAs to estimate memory to be provisioned
at the startup and during run time as the resizable structures change
their sizes.
Portability
===========
Resizable shared structures feature depends upon existence of function
madvise() and constants MADV_REMOVE and MADV_WRITE_POPULATE. On the
platforms which do not have these or the shared memory types (e.g.
Sys-V) which do not provide ability to free or allocate parts of shared
memory, we disable this feature. A run time GUC have_resizable_shmem
indicates whether a running server supports resizable structures or not.
The commit introduces a compile time flag HAVE_RESIZABLE_SHMEM which is
defined if MADV_REMOVE and MADV_WRITE_POPULATE exist. We don't check
existence of madvise separately, since existence of the constants
implies existence of the function. HAVE_RESIZABLE_SHMEM is not defined
in EXEC_BACKEND builds since that's largely used for Windows where the
APIs to free and allocate memory from and to a given address space are
not known to the author right now. Given that PostgreSQL is used widely
on Linux (with shared memory type mmap), providing this feature on Linux
benefits most of its users. Once we figure out the required Windows
APIs, we will support this feature on Windows as well.
Prohibiting access to the unused portion of a resizable structure
================================================================
We provide ShmemProtectStruct() to add protections on the address space
reserved for a resizable shared memory structure so that the address
space upto its current size is accessible whereas the part beyond that
is inaccessible. Since these protections are backend specific, the
subsystem using the resizable shared memory structures has to make sure
to call ShmemProtectStruct() after every resize before any backend tries
to access the portion of the structure between old and the new size.
This can coordinated using ProcSignalBarrier mechanism in a running
backend. When starting the server, Postmaster adds appropriate
protections when creating these structures. In an EXEC_BACKEND case,
when a backend starts, the protections are added according to the
current size of the backend (through attach_fn call). However, in
non-EXEC_BACKEND case, a new backend inherits stale protections from the
postmaster. It needs to apply the protections according to the current
sizes of these structures.
Following points need more discussion.
Discussion points
=================
adding initial protections in non-EXEC_BACKEND case
----------------------------------------------------------------------
When a backend starts, it needs to add shmem protections it in such a
way that a ProcSignalBarrier conveying the protection change is not
missed. Hence we do it along side InitLocalDataChecksumState(). This has
two problems 1. it delays backend startup process and 2. it loosely ties
resizable shared memory structure to ProcSignalBarrier mechanim. Do we
consider the complexity worth it? Is there any other way to do it
without these two drawbacks?
Further, we call ShmemReprotectResizableStructs() twice in the startup
sequence, once immediately after the place where
AttachSharedMemoryStructs() would be called in EXEC_BACKEND case and
then second time after ProcSignalInit(). The first one is needed so that
the processes can access the resizable structures safely. The second is
needed for the reasons mentioned there. But a change in the protection
during these calls will be missed by the backend and thus lead to
hazardous access. Can we avoid it? If ProcSignalInit() were to be called
nearer to the InitProcess(), it may make calling
ShmemReprotectResizeStructs() easier.
Protection in the resizing backend vs other backends
----------------------------------------------------
The patch provides an API to protect the address spaces not used because
of current size of the resizable structures in every backend. Since
every backend has to protect its own address space, a separate API is
required. But at the same time, in some cases the resize operation
itself needs to change the protection (e.g. when expanding the
structure). So there's some asymmetry between the moment when the
resizing backend adds protection and when other backends add it. That
doesn't affect end result since adding this protection is an idempotent
operation. But still there's some ugliness involved. This part also
needs more discussion.
HAVE_RESIZABLE_SHMEM and have_resizable_shmem
---------------------------------------------
The patch defines a build time macro HAVE_RESIZABLE_SHMEM to avoid
compilation failure on a build platform where the required system calls
or constants are not available. But we also need a run time check to see
whether the shared memory type being used allows resizing. Hence we
provide a GUC have_resizable_shmem which tells whether a running server
supports resizable shared memory or not. Most of the code which deals
with the shared memory is under HAVE_RESIZABLE_SHMEM, but there is some
related to resizable shared memory outside the macro as well. For
example, the members maximum_size, minimum_size are not under this
macro. This is mostly to allow writing readable code in both the modes
and still provide the same user visible views etc. This division of code
within and without macro needs more discussion.
TODOs
=====
There are some TODOs in the code still. We will address them as we
finalize that part of design/code. Comments/suggestions on those is
welcome.
Idea of using mmap to reserve address space in shared memory segments
and allocating memory within that address space on demand was proposed
by Dmitry Dolgov <9erthalion6@gmail.com>.
Author: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Reviewed-by: Matthias van de Meent <boekewurm+postgres@gmail.com>
---
configure.ac | 4 +
doc/src/sgml/config.sgml | 102 +++-
doc/src/sgml/ref/postgres-ref.sgml | 8 +-
doc/src/sgml/runtime.sgml | 9 +-
doc/src/sgml/system-views.sgml | 57 +-
doc/src/sgml/xfunc.sgml | 92 +++
meson.build | 16 +
src/backend/port/sysv_shmem.c | 259 +++++++-
src/backend/port/win32_shmem.c | 64 +-
src/backend/postmaster/auxprocess.c | 46 +-
src/backend/storage/ipc/ipci.c | 114 ++--
src/backend/storage/ipc/shmem.c | 555 ++++++++++++++++--
src/backend/storage/lmgr/proc.c | 16 +
src/backend/utils/init/postinit.c | 10 +
src/backend/utils/misc/guc_parameters.dat | 57 +-
src/backend/utils/misc/guc_tables.c | 15 +-
src/include/catalog/pg_proc.dat | 4 +-
src/include/pg_config.h.in | 8 +
src/include/pg_config_manual.h | 14 +
src/include/storage/ipc.h | 2 +-
src/include/storage/pg_shmem.h | 6 +-
src/include/storage/shmem.h | 19 +
src/include/storage/shmem_internal.h | 2 +-
src/test/modules/test_shmem/meson.build | 3 +-
...mem_alloc.pl => 001_fixed_shmem_struct.pl} | 31 +
.../t/002_resizable_shmem_struct.pl | 387 ++++++++++++
.../modules/test_shmem/test_shmem--1.0.sql | 55 ++
src/test/modules/test_shmem/test_shmem.c | 504 +++++++++++++++-
src/test/regress/expected/rules.out | 7 +-
src/tools/pgindent/typedefs.list | 1 +
30 files changed, 2299 insertions(+), 168 deletions(-)
rename src/test/modules/test_shmem/t/{001_late_shmem_alloc.pl => 001_fixed_shmem_struct.pl} (58%)
create mode 100644 src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
diff --git a/configure.ac b/configure.ac
index a331749fcb5..bba12b8c8f4 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1912,6 +1912,10 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux-specific madvise constants needed for resizable shared memory. See similar checks in meson.build for explanation of why these checks are here.
+AC_CHECK_DECLS([MADV_POPULATE_WRITE], [], [], [#include <sys/mman.h>])
+AC_CHECK_DECLS([MADV_REMOVE], [], [], [#include <sys/mman.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 236ee067f40..b84e7d6c799 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -12381,6 +12381,20 @@ dynamic_library_path = '/usr/local/lib/postgresql:$libdir'
</listitem>
</varlistentry>
+ <varlistentry id="guc-have-resizable-shmem" xreflabel="have_resizable_shmem">
+ <term><varname>have_resizable_shmem</varname> (<type>boolean</type>)
+ <indexterm>
+ <primary><varname>have_resizable_shmem</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports whether <productname>PostgreSQL</productname> supports <link
+ linkend="xfunc-shared-addin-resizable">Resizable shared memory structures</link>.
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry id="guc-huge-pages-status" xreflabel="huge_pages_status">
<term><varname>huge_pages_status</varname> (<type>enum</type>)
<indexterm>
@@ -12560,32 +12574,100 @@ dynamic_library_path = '/usr/local/lib/postgresql:$libdir'
</listitem>
</varlistentry>
- <varlistentry id="guc-shared-memory-size" xreflabel="shared_memory_size">
- <term><varname>shared_memory_size</varname> (<type>integer</type>)
+ <varlistentry id="guc-shared-memory-initial-size" xreflabel="shared_memory_initial_size">
+ <term><varname>shared_memory_initial_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_initial_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the amount of memory, rounded up to the nearest megabyte, allocated at the server startup in main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-minimum-size" xreflabel="shared_memory_minimum_size">
+ <term><varname>shared_memory_minimum_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_minimum_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the minimum amount of memory, rounded up to the nearest megabyte, required in the main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-maximum-size" xreflabel="shared_memory_maximum_size">
+ <term><varname>shared_memory_maximum_size</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_maximum_size</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the maximum size of the main shared memory area, rounded up
+ to the nearest megabyte. This is the amount of address space that
+ must be reserved for the main shared memory area. This is also the maximum amount of memory that may be allocated in the main shared memory area.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-initial-size-in-huge-pages" xreflabel="shared_memory_initial_size_in_huge_pages">
+ <term><varname>shared_memory_initial_size_in_huge_pages</varname> (<type>integer</type>)
<indexterm>
- <primary><varname>shared_memory_size</varname> configuration parameter</primary>
+ <primary><varname>shared_memory_initial_size_in_huge_pages</varname> configuration parameter</primary>
</indexterm>
</term>
<listitem>
<para>
- Reports the size of the main shared memory area, rounded up to the
- nearest megabyte.
+ Reports the number of huge pages needed in the main shared memory area
+ at the startup based on the specified <xref
+ linkend="guc-huge-page-size"/>. If huge pages are not supported, this
+ will be <literal>-1</literal>.
+ </para>
+ <para>
+ This setting is supported only on <productname>Linux</productname>. It
+ is always set to <literal>-1</literal> on other platforms. For more
+ details about using huge pages on <productname>Linux</productname>, see
+ <xref linkend="linux-huge-pages"/>.
</para>
</listitem>
</varlistentry>
- <varlistentry id="guc-shared-memory-size-in-huge-pages" xreflabel="shared_memory_size_in_huge_pages">
- <term><varname>shared_memory_size_in_huge_pages</varname> (<type>integer</type>)
+ <varlistentry id="guc-shared-memory-minimum-size-in-huge-pages" xreflabel="shared_memory_minimum_size_in_huge_pages">
+ <term><varname>shared_memory_minimum_size_in_huge_pages</varname> (<type>integer</type>)
<indexterm>
- <primary><varname>shared_memory_size_in_huge_pages</varname> configuration parameter</primary>
+ <primary><varname>shared_memory_minimum_size_in_huge_pages</varname> configuration parameter</primary>
</indexterm>
</term>
<listitem>
<para>
- Reports the number of huge pages that are needed for the main shared
- memory area based on the specified <xref linkend="guc-huge-page-size"/>.
+ Reports the minimum number of huge pages needed in the main shared memory area based on the specified <xref linkend="guc-huge-page-size"/>.
If huge pages are not supported, this will be <literal>-1</literal>.
</para>
+ <para>
+ This setting is supported only on <productname>Linux</productname>. It
+ is always set to <literal>-1</literal> on other platforms.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-shared-memory-maximum-size-in-huge-pages" xreflabel="shared_memory_maximum_size_in_huge_pages">
+ <term><varname>shared_memory_maximum_size_in_huge_pages</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>shared_memory_maximum_size_in_huge_pages</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Reports the maximum number of huge pages needed in the
+ main shared memory area based on the specified
+ <xref linkend="guc-huge-page-size"/>. If huge pages are not
+ supported, this will be <literal>-1</literal>.
+ </para>
<para>
This setting is supported only on <productname>Linux</productname>. It
is always set to <literal>-1</literal> on other platforms. For more
diff --git a/doc/src/sgml/ref/postgres-ref.sgml b/doc/src/sgml/ref/postgres-ref.sgml
index b13a16a117f..3eefd629e06 100644
--- a/doc/src/sgml/ref/postgres-ref.sgml
+++ b/doc/src/sgml/ref/postgres-ref.sgml
@@ -143,8 +143,12 @@ PostgreSQL documentation
<para>
This can be used on a running server for most parameters. However,
the server must be shut down for some runtime-computed parameters
- (e.g., <xref linkend="guc-shared-memory-size"/>,
- <xref linkend="guc-shared-memory-size-in-huge-pages"/>, and
+ (e.g., <xref linkend="guc-shared-memory-initial-size"/>,
+ <xref linkend="guc-shared-memory-minimum-size"/>,
+ <xref linkend="guc-shared-memory-maximum-size"/>,
+ <xref linkend="guc-shared-memory-initial-size-in-huge-pages"/>,
+ <xref linkend="guc-shared-memory-minimum-size-in-huge-pages"/>,
+ <xref linkend="guc-shared-memory-maximum-size-in-huge-pages"/>, and
<xref linkend="guc-wal-segment-size"/>).
</para>
diff --git a/doc/src/sgml/runtime.sgml b/doc/src/sgml/runtime.sgml
index d9984910cc4..8ba4d233509 100644
--- a/doc/src/sgml/runtime.sgml
+++ b/doc/src/sgml/runtime.sgml
@@ -1450,11 +1450,12 @@ export PG_OOM_ADJUST_VALUE=0
<varname>CONFIG_HUGETLB_PAGE=y</varname>. You will also have to configure
the operating system to provide enough huge pages of the desired size.
The runtime-computed parameter
- <xref linkend="guc-shared-memory-size-in-huge-pages"/> reports the number
- of huge pages required. This parameter can be viewed before starting the
+ <xref linkend="guc-shared-memory-maximum-size-in-huge-pages"/> reports the
+ number of huge pages that must be available in the pool for postmaster
+ startup to succeed. This parameter can be viewed before starting the
server with a <command>postgres</command> command like:
<programlisting>
-$ <userinput>postgres -D $PGDATA -C shared_memory_size_in_huge_pages</userinput>
+$ <userinput>postgres -D $PGDATA -C shared_memory_maximum_size_in_huge_pages</userinput>
3170
$ <userinput>grep ^Hugepagesize /proc/meminfo</userinput>
Hugepagesize: 2048 kB
@@ -1465,7 +1466,7 @@ hugepages-1048576kB hugepages-2048kB
In this example the default is 2MB, but you can also explicitly request
either 2MB or 1GB with <xref linkend="guc-huge-page-size"/> to adapt
the number of pages calculated by
- <varname>shared_memory_size_in_huge_pages</varname>.
+ <varname>shared_memory_maximum_size_in_huge_pages</varname>.
While we need at least <literal>3170</literal> huge pages in this example,
a larger setting would be appropriate if other programs on the machine
diff --git a/doc/src/sgml/system-views.sgml b/doc/src/sgml/system-views.sgml
index 6b905498337..7c4f4aeba14 100644
--- a/doc/src/sgml/system-views.sgml
+++ b/doc/src/sgml/system-views.sgml
@@ -4333,8 +4333,46 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
Size of the allocation in bytes including padding. For anonymous
allocations, no information about padding is available, so the
<literal>size</literal> and <literal>allocated_size</literal> columns
- will always be equal. Padding is not meaningful for free memory, so
- the columns will be equal in that case also.
+ will always be equal. Padding is not meaningful for free memory, so the
+ columns will be equal in that case also. For resizable allocations which
+ may span multiple memory pages, the padding includes the padding due to
+ page alignment.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>minimum_size</structfield> <type>int8</type>
+ </para>
+ <para>
+ Minimum size in bytes that the resizable allocation can shrink to. Equals
+ <structfield>size</structfield> for fixed-size allocations, anonymous
+ allocations, and free memory.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>maximum_size</structfield> <type>int8</type>
+ </para>
+ <para>
+ Maximum size in bytes that the resizable allocation can grow to. Equals
+ <structfield>size</structfield> for fixed-size allocations, anonymous
+ allocations, and free memory.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>reserved_space</structfield> <type>int8</type>
+ </para>
+ <para>
+ Address space reserved for the allocation in bytes. For resizable
+ structures, this is the total address space reserved to accommodate
+ growth up to <structfield>maximum_size</structfield>, and is greater
+ than or equal to <structfield>allocated_size</structfield>. For
+ fixed-size allocations, anonymous allocations, and free memory this
+ is same as <structfield>allocated_size</structfield>.
</para></entry>
</row>
</tbody>
@@ -4348,6 +4386,21 @@ SELECT * FROM pg_locks pl LEFT JOIN pg_prepared_xacts ppx
<literal>ShmemRequestHash()</literal>.
</para>
+ <para>
+ Resizable structures are allocations whose size can change while the server
+ is running. <xref linkend="guc-have-resizable-shmem"/> indicates whether the
+ running server supports resizable allocations. Each such allocation has a
+ minimum size give by <structfield>minimum_size</structfield> it can shrink to
+ and a maximum size given by <structfield>maximum_size</structfield> it can
+ grow to; enough address space is reserved up front to accommodate its
+ <structfield>maximum_size</structfield>. The aggregate initial, minimum and
+ maximum sizes of all shared memory allocations are reported by <xref
+ linkend="guc-shared-memory-initial-size"/>, <xref
+ linkend="guc-shared-memory-minimum-size"/> and <xref
+ linkend="guc-shared-memory-maximum-size"/> respectively (and their
+ <literal>_in_huge_pages</literal> counterparts).
+ </para>
+
<para>
By default, the <structname>pg_shmem_allocations</structname> view can be
read only by superusers or roles with privileges of the
diff --git a/doc/src/sgml/xfunc.sgml b/doc/src/sgml/xfunc.sgml
index 2b8a11e7ad0..2bfc664d7ca 100644
--- a/doc/src/sgml/xfunc.sgml
+++ b/doc/src/sgml/xfunc.sgml
@@ -3744,6 +3744,98 @@ my_shmem_init(void *arg)
</para>
</sect3>
+ <sect3 id="xfunc-shared-addin-resizable">
+ <title>Resizable shared memory structures</title>
+
+ <para>
+ A resizable memory structure can be requested using
+ <function>ShmemRequestStruct</function> by passing
+ <parameter>.maximum_size</parameter> along with
+ <parameter>.size</parameter>. <parameter>.maximum_size</parameter> is
+ maximum size upto which the structure can grow where as
+ <parameter>.size</parameter> is the initial size of the structure.
+ Optionally, <parameter>.minimum_size</parameter> can be set to the minimum
+ size that the structure can shrink to. While
+ contiguous address space worth <parameter>maximum_size</parameter> is
+ allocated to the structure, only memory worth <parameter>size</parameter>
+ bytes is allocated initially. The <function>init_fn</function> should only
+ initialize the <parameter>size</parameter> amount of memory. The actual
+ memory allocated to this structure at any point in time is given by <link
+ linkend="view-pg-shmem-allocations"><structname>pg_shmem_allocations</structname>.<structfield>allocated_size</structfield></link>
+ and the address space reserved for this structure is given by <link
+ linkend="view-pg-shmem-allocations"><structname>pg_shmem_allocations</structname>.<structfield>reserved_space</structfield></link>.
+ </para>
+
+ <para>
+ The structure can be resized using <function>ShmemResizeStruct</function> by
+ passing it the structure's name and the
+ new size which can be anywhere between <parameter>minimum_size</parameter>
+ and <parameter>maximum_size</parameter>. If the new size is smaller than the
+ current size of the structure, the memory between the new size and current
+ size is freed while keeping the contents of the memory upto new size intact.
+ If the new size is greater than the current size, memory is allocated upto
+ new size while keeping the current contents of the structure intact. The
+ starting address of the structure does not change because of resizing
+ operation.
+ </para>
+
+ <para>
+ <function>ShmemResizeStruct</function> returns <literal>true</literal> on
+ success. If the operating system cannot supply the additional memory
+ required to expand the structure, it emits a <literal>WARNING</literal>
+ naming the structure and the number of bytes that could not be allocated,
+ leaves the structure at its previous size, and returns
+ <literal>false</literal>.
+ </para>
+
+ <para>
+ <function>ShmemResizeStruct</function> does not coordinate with other
+ backends that may be accessing the shared structure. The caller is
+ responsible for ensuring that no backend is accessing the parts of the
+ structure between old size and the new size while
+ <function>ShmemResizeStruct</function> is executing. Similarly after the
+ function has finished, the caller must ensure that each backend knows the
+ new size before accessing the parts of the structure between the old size
+ and the new size. Accessing the range <literal>[current_size,
+ maximum_size)</literal> results in an undefined operating-system dependent
+ behaviour.
+ </para>
+
+ <para>
+ <function>ShmemProtectStruct</function> can be used to make the range beyond
+ the current size inaccessible in a backend's address space, so that stray
+ accesses cause a fault instead of silently succeeding. When a backend
+ starts, protections for all resizable structures are adjusted according to
+ their current sizes, so subsystems need not call
+ <function>ShmemProtectStruct</function> from their per-backend
+ initialization routines. During runtime after calling
+ <function>ShmemResizeStruct</function> from one backend,
+ <function>ShmemProtectStruct</function> should be called from all the
+ backends that may access the structure.
+ </para>
+
+ <para>
+ <productname>PostgreSQL</productname>'s
+ <function>ProcSignalBarrier</function> mechanism (see
+ <filename>src/include/storage/procsignal.h</filename>) may be used to
+ coordinate the resize across backends and adjusting access to the address
+ space.
+ </para>
+
+ <para>
+ This functionality is available only on the platforms which provide the APIs
+ necessary to reserve contiguous address space and to allocate or free memory
+ in that address space on demand. Macro <symbol>HAVE_RESIZABLE_SHMEM</symbol>
+ is defined on such platforms. It can be used to guard code related to
+ resizing a shared memory structure. The functionality is available on with
+ mmap'ed memory, so subsystems which use resizable structures may have to
+ addtionally disable resizable memory usage when <symbol>shared_memory_type</symbol> is not
+ <symbol>SHMEM_TYPE_MMAP</symbol>. A GUC <xref linkend="guc-have-resizable-shmem"/> is set to
+ <literal>on</literal> when this functionality is available in a running
+ server, <literal>off</literal> otherwise.
+ </para>
+ </sect3>
+
<sect3 id="xfunc-shared-addin-dynamic">
<title>Allocating Dynamic Shared Memory after Startup</title>
diff --git a/meson.build b/meson.build
index f4cde249242..4e0c736470c 100644
--- a/meson.build
+++ b/meson.build
@@ -2894,6 +2894,22 @@ decl_checks = [
['timingsafe_bcmp', 'string.h'],
]
+# Linux-specific madvise constants needed for resizable shared memory.
+# Usually we use AC_CHECK_DECLS to check for function declarations, but in this
+# case we are using it to detect existence of constants. These constants are
+# used to define HAVE_RESIZABLE_SHMEM which is used in storage/pg_shmem.h as
+# well as storage/shmem.h. The first abstracts the APIs to allocate shared
+# memory segments from the operating system whereas the second abstracts APIs to
+# allocate shared memory to various subsystems. Since they are related but
+# orthogonal to each other, including any one of them in the other file doesn't
+# make sense. pg_config_manual.h is the only place where HAVE_RESIZABLE_SHMEM
+# can be defined and made available to both without including sys/mman.h. But
+# for that we need constants that indicate the existence of following defines.
+decl_checks += [
+ ['MADV_POPULATE_WRITE', 'sys/mman.h'],
+ ['MADV_REMOVE', 'sys/mman.h'],
+]
+
# Need to check for function declarations for these functions, because
# checking for library symbols wouldn't handle deployment target
# restrictions on macOS
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 2e3886cf9fe..c052776e94c 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -589,44 +589,113 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
return true;
}
+/*
+ * Get the page size being used by the shared memory.
+ *
+ * The function should be called only after the shared memory has been setup.
+ */
+size_t
+GetOSPageSize(void)
+{
+ size_t os_page_size;
+
+ Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
+
+ os_page_size = sysconf(_SC_PAGESIZE);
+
+ /* If huge pages are actually in use, use huge page size */
+ if (huge_pages_status == HUGE_PAGES_ON)
+ GetHugePageSize(&os_page_size, NULL);
+
+ return os_page_size;
+}
+
/*
* Creates an anonymous mmap()ed shared memory segment.
*
- * Pass the requested size in *size. This function will modify *size to the
- * actual size of the allocation, if it ends up allocating a segment that is
- * larger than requested.
+ * initial_size is the amount of memory required at the start of the server.
+ *
+ * *size is the size of the anonymous memory mapping to create. It will be
+ * updated to the actual size of the allocation, if it ends up allocating a
+ * segment that is larger than the requested. When *size > initial_size we want
+ * to reserve a large memory segment while allocating only initial_size amount of
+ * memory at the server start.
+ *
+ * When using huge pages, we make sure that there are enough huge pages
+ * configured to cover the initial_size worth of memory.
*/
static void *
-CreateAnonymousSegment(Size *size)
+CreateAnonymousSegment(Size initial_size, Size *size)
{
Size allocsize = *size;
void *ptr = MAP_FAILED;
int mmap_errno = 0;
+ bool resizable = (initial_size < allocsize);
int mmap_flags = MAP_SHARED | MAP_ANONYMOUS | MAP_HASSEMAPHORE;
+ Assert(initial_size <= allocsize);
+
#ifndef MAP_HUGETLB
/* PGSharedMemoryCreate should have dealt with this case */
Assert(huge_pages != HUGE_PAGES_ON);
#else
if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY)
{
- /*
- * Round up the request size to a suitable large value.
- */
Size hugepagesize;
int huge_mmap_flags;
+ bool probe_ok = true;
+ /* Round up the request size to a suitable large value. */
GetHugePageSize(&hugepagesize, &huge_mmap_flags);
-
if (allocsize % hugepagesize != 0)
allocsize = add_size(allocsize, hugepagesize - (allocsize % hugepagesize));
+ if (initial_size % hugepagesize != 0)
+ initial_size = add_size(initial_size, hugepagesize - (initial_size % hugepagesize));
- ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- mmap_flags | huge_mmap_flags, -1, 0);
- mmap_errno = errno;
- if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
- elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
- allocsize);
+ /*
+ * When the total amount of space requested is larger than the initial
+ * memory requirement, we do not allocate memory worth the entire
+ * requested space upfront. But we will need to make sure that there
+ * are enough huge pages configured to cover the initial memory
+ * requirement. Otherwise, we will choose huge pages map instead of
+ * falling back to normal pages and the server will not start. Hence
+ * we first try to map and allocate the initial memory requirement. If
+ * it fails we fall back to normal pages. If it succeeds, we unmap the
+ * memory and then try to map the entire requested space without
+ * allocating memory.
+ */
+ if (resizable)
+ {
+ void *probe;
+
+ probe = mmap(NULL, initial_size, PROT_READ | PROT_WRITE,
+ mmap_flags | huge_mmap_flags,
+ -1, 0);
+ if (probe == MAP_FAILED)
+ {
+ mmap_errno = errno;
+ probe_ok = false;
+ if (huge_pages == HUGE_PAGES_TRY)
+ elog(DEBUG1,
+ "huge-page probe mmap(%zu) failed, huge pages disabled: %m",
+ initial_size);
+ }
+ else if (munmap(probe, initial_size) != 0)
+ elog(ERROR, "munmap(%p, %zu) huge page probe failed: %m", probe, initial_size);
+ }
+
+ if (probe_ok)
+ {
+ if (resizable)
+ mmap_flags |= MAP_NORESERVE;
+ ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
+ mmap_flags | huge_mmap_flags, -1, 0);
+ mmap_errno = errno;
+ if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
+ elog(DEBUG1,
+ "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
+ allocsize);
+ }
}
#endif
@@ -645,6 +714,10 @@ CreateAnonymousSegment(Size *size)
* to non-huge pages.
*/
allocsize = *size;
+
+ /* Use MAP_NORESERVE when we do not need all the memory up front. */
+ if (resizable)
+ mmap_flags |= MAP_NORESERVE;
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
mmap_flags, -1, 0);
mmap_errno = errno;
@@ -693,13 +766,18 @@ AnonymousShmemDetach(int status, Datum arg)
* standard header. Also, register an on_shmem_exit callback to release
* the storage.
*
+ * initial_size is the amount of memory required at the start of the server.
+ * Used only when we want to reserve a large memory segment while allocating
+ * only initial_size amount of memory at the server start. That facility is only
+ * available when using anonymous shared memory segments.
+ *
* Dead Postgres segments pertinent to this DataDir are recycled if found, but
* we do not fail upon collision with foreign shmem segments. The idea here
* is to detect and re-use keys that may have been assigned by a crashed
* postmaster or backend.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(Size initial_size, Size size,
PGShmemHeader **shim)
{
IpcMemoryKey NextShmemSegID;
@@ -738,7 +816,7 @@ PGSharedMemoryCreate(Size size,
if (shared_memory_type == SHMEM_TYPE_MMAP)
{
- AnonymousShmem = CreateAnonymousSegment(&size);
+ AnonymousShmem = CreateAnonymousSegment(initial_size, &size);
AnonymousShmemSize = size;
/* Register on-exit routine to unmap the anonymous segment */
@@ -754,6 +832,10 @@ PGSharedMemoryCreate(Size size,
/* huge pages are only available with mmap */
SetConfigOption("huge_pages_status", "off",
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ /* resizable shared memory is only available with mmap */
+ SetConfigOption("have_resizable_shmem", "off",
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
}
/*
@@ -991,3 +1073,148 @@ PGSharedMemoryDetach(void)
AnonymousShmem = NULL;
}
}
+
+/*
+ * Make sure that the memory of given size from the given address is released.
+ *
+ * The address and size are expected to be page aligned.
+ *
+ * Only supported on platforms that support anonymous shared memory.
+ *
+ * On success returns true. false if system call fails.
+ */
+bool
+PGSharedMemoryEnsureFreed(void *addr, Size size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be freed"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(addr == (void *) TYPEALIGN(os_page_size, addr));
+ Assert(size == TYPEALIGN(os_page_size, size));
+ }
+#endif
+
+ if (madvise(addr, size, MADV_REMOVE) == -1)
+ {
+ ereport(WARNING, errmsg("could not free shared memory: %m"));
+ return false;
+ }
+
+ return true;
+#endif
+}
+
+/*
+ * Make sure that the memory of given size from the given address is allocated.
+ *
+ * The address and size are expected to be page aligned. The caller is
+ * responsible for ensuring that the range is already writable; this function
+ * only populates it.
+ *
+ * Only supported on platforms that support anonymous shared memory.
+ *
+ * On success returns true, false if system call fails.
+ */
+bool
+PGSharedMemoryEnsureAllocated(void *addr, Size size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be allocated at runtime"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(addr == (void *) TYPEALIGN(os_page_size, addr));
+ Assert(size == TYPEALIGN(os_page_size, size));
+ }
+#endif
+
+ if (madvise(addr, size, MADV_POPULATE_WRITE) == -1)
+ {
+ ereport(WARNING, errmsg("could not allocate shared memory: %m"));
+ return false;
+ }
+
+ return true;
+#endif
+}
+
+/*
+ * Set memory protection on the given region of shared memory.
+ *
+ * Makes [rw_start, rw_end) readable and writable, and [rw_end, prot_end)
+ * inaccessible.
+ *
+ * All addresses are expected to be page aligned.
+ *
+ * Only supported on platforms that support resizable shared memory.
+ *
+ * On success returns true, false if system call fails.
+ */
+bool
+PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+
+ if (!AnonymousShmem)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("only anonymous shared memory can be protected at runtime"));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ size_t os_page_size = GetOSPageSize();
+
+ Assert(rw_start == (void *) TYPEALIGN(os_page_size, rw_start));
+ Assert(rw_end == (void *) TYPEALIGN(os_page_size, rw_end));
+ Assert(prot_end == (void *) TYPEALIGN(os_page_size, prot_end));
+ }
+#endif
+ Assert(rw_end >= rw_start);
+
+ if (rw_end > rw_start)
+ {
+ if (mprotect(rw_start, (char *) rw_end - (char *) rw_start,
+ PROT_READ | PROT_WRITE) != 0)
+ {
+ ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ return false;
+ }
+ }
+
+ if (prot_end > rw_end)
+ {
+ if (mprotect(rw_end, (char *) prot_end - (char *) rw_end,
+ PROT_NONE) != 0)
+ {
+ ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ return false;
+ }
+ }
+
+ return true;
+#endif
+}
diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c
index 794e4fcb2ad..d266cdbe226 100644
--- a/src/backend/port/win32_shmem.c
+++ b/src/backend/port/win32_shmem.c
@@ -200,11 +200,13 @@ EnableLockPagesPrivilege(int elevel)
/*
* PGSharedMemoryCreate
*
- * Create a shared memory segment of the given size and initialize its
- * standard header.
+ * Create a shared memory segment of the given size and initialize its standard
+ * header. initial_size is only relevant when we want to create a larger memory
+ * segment with only a part of it being allocated initially. We don't support
+ * that on Windows, so we ignore it.
*/
PGShmemHeader *
-PGSharedMemoryCreate(Size size,
+PGSharedMemoryCreate(Size initial_size, Size size,
PGShmemHeader **shim)
{
void *memAddress;
@@ -648,3 +650,59 @@ check_huge_page_size(int *newval, void **extra, GucSource source)
}
return true;
}
+
+/*
+ * Get the page size used by the shared memory.
+ *
+ * The function should be called only after the shared memory has been setup.
+ */
+size_t
+GetOSPageSize(void)
+{
+ SYSTEM_INFO sysinfo;
+ size_t os_page_size;
+
+ Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
+
+ GetSystemInfo(&sysinfo);
+ os_page_size = sysinfo.dwPageSize;
+
+ /* If huge pages are actually in use, use huge page size */
+ if (huge_pages_status == HUGE_PAGES_ON)
+ GetHugePageSize(&os_page_size, NULL);
+
+ return os_page_size;
+}
+
+/*
+ * PGSharedMemoryEnsureFreed / PGSharedMemoryEnsureAllocated
+ *
+ * Not supported on Windows. These are only meaningful on platforms with
+ * resizable shared memory (mmap + madvise).
+ */
+bool
+PGSharedMemoryEnsureFreed(void *addr, Size size)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
+
+bool
+PGSharedMemoryEnsureAllocated(void *addr, Size size)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
+
+bool
+PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
+{
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ return false; /* keep compiler quiet */
+}
diff --git a/src/backend/postmaster/auxprocess.c b/src/backend/postmaster/auxprocess.c
index ad4bf4bd2a8..07a3b5c5923 100644
--- a/src/backend/postmaster/auxprocess.c
+++ b/src/backend/postmaster/auxprocess.c
@@ -69,33 +69,43 @@ AuxiliaryProcessMainCommon(void)
BaseInit();
/*
- * Prevent consuming interrupts between setting ProcSignalInit and setting
- * the initial local data checksum value. If a barrier is emitted, and
- * absorbed, before local cached state is initialized the state transition
- * can be invalid.
- */
- HOLD_INTERRUPTS();
-
- ProcSignalInit(NULL, 0);
-
- /*
- * Initialize a local cache of the data_checksum_version, to be updated by
- * the procsignal-based barriers.
+ * Prevent consuming interrupts between ProcSignalInit() and the
+ * initialization of local states that are kept in sync with shared memory
+ * via procsignal-based barriers. If a barrier is emitted, and absorbed,
+ * before local cached state is initialized the state transition can be
+ * invalid.
*
- * This intentionally happens after initializing the procsignal, otherwise
- * we might miss a state change. This means we can get a barrier for the
- * state we've just initialized - but it can happen only once.
+ * These initializations intentionally happen after ProcSignalInit(),
+ * otherwise we might miss a state change. This means we may also receive
+ * a barrier for the state we've just initialized, but it can happen only
+ * once.
*
* The postmaster (which is what gets forked into the new child process)
- * does not handle barriers, therefore it may not have the current value
- * of LocalDataChecksumState value (it'll have the value read from the
- * control file, which may be arbitrarily old).
+ * does not handle barriers, therefore its local states may not reflect
+ * the current state of the shared memory.
*
* NB: Even if the postmaster handled barriers, the value might still be
* stale, as it might have changed after this process forked.
*/
+ HOLD_INTERRUPTS();
+
+ ProcSignalInit(NULL, 0);
+
+ /*
+ * LocalDataChecksumState inherited from Postmaster will have the value
+ * read from the control file, which may be arbitrarily old. Update it.
+ */
InitLocalDataChecksumState();
+ /*
+ * Refresh per-backend protections for resizable shmem structures, in case
+ * these structures have been resized since the startup. Structures are
+ * expected to be kept in sync by respective subsystems using
+ * procsignal-based barriers. But we modify their protections en-masse
+ * here, rather than relying on individual subsystems to do it.
+ */
+ ShmemReprotectResizableStructs();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index e149a738c8d..e5f20c4603e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -52,31 +52,50 @@ RequestAddinShmemSpace(Size size)
/*
* CalculateShmemSize
* Calculates the amount of shared memory needed.
+ *
+ * - `initial` is the amount of memory needed when the server startup.
+ * - `min` is the minimum amount of memory needed when all the resizable
+ * structures are shrunk to their respective minimum sizes.
+ * - `max` is the maximum amount of memory needed when all the resizable
+ * structures are expanded to their respective maximum sizes. It is also the
+ * size of address space that must be reserved for the shared memory segment.
+ *
+ * When no resizable structures are requested, all three totals are identical.
+ *
+ * We take some care to ensure that the total size request doesn't overflow
+ * size_t. If this gets through, we don't need to be so careful during the
+ * actual allocation phase.
*/
-Size
-CalculateShmemSize(void)
+void
+CalculateShmemSize(size_t *initial, size_t *min, size_t *max)
{
- Size size;
+ size_t initial_req;
+ size_t min_req;
+ size_t max_req;
+ size_t fixed_addins;
+
+ ShmemGetRequestedSize(&initial_req, &min_req, &max_req);
/*
* Size of the Postgres shared-memory block is estimated via moderately-
* accurate estimates for the big hogs, plus 100K for the stuff that's too
* small to bother with estimating.
*
- * We take some care to ensure that the total size request doesn't
- * overflow size_t. If this gets through, we don't need to be so careful
- * during the actual allocation phase.
+ * Also include additional requested shmem from preload libraries.
+ *
+ * These are not resizable, so they contribute equally to all three
+ * totals.
*/
- size = 100000;
- size = add_size(size, ShmemGetRequestedSize());
-
- /* include additional requested shmem from preload libraries */
- size = add_size(size, total_addin_request);
+ fixed_addins = add_size(100000, total_addin_request);
- /* might as well round it off to a multiple of a typical page size */
- size = add_size(size, 8192 - (size % 8192));
+ *initial = add_size(initial_req, fixed_addins);
+ *min = add_size(min_req, fixed_addins);
+ *max = add_size(max_req, fixed_addins);
- return size;
+ /* might as well round each off to a multiple of a typical page size */
+ *initial = add_size(*initial, 8192 - (*initial % 8192));
+ *min = add_size(*min, 8192 - (*min % 8192));
+ *max = add_size(*max, 8192 - (*max % 8192));
}
#ifdef EXEC_BACKEND
@@ -121,18 +140,24 @@ CreateSharedMemoryAndSemaphores(void)
{
PGShmemHeader *shim;
PGShmemHeader *seghdr;
- Size size;
+ size_t initial_size;
+ size_t min_size;
+ size_t max_size;
Assert(!IsUnderPostmaster);
/* Compute the size of the shared-memory block */
- size = CalculateShmemSize();
- elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size);
+ CalculateShmemSize(&initial_size, &min_size, &max_size);
+ elog(DEBUG3, "invoking IpcMemoryCreate(initial size=%zu, minimum size=%zu, maximum size=%zu)",
+ initial_size, min_size, max_size);
/*
- * Create the shmem segment
+ * Create the shmem segment.
+ *
+ * Reserve enough shared address space to accommodate every requested
+ * structure grown to its maximum size.
*/
- seghdr = PGSharedMemoryCreate(size, &shim);
+ seghdr = PGSharedMemoryCreate(initial_size, max_size, &shim);
/*
* Make sure that huge pages are never reported as "unknown" while the
@@ -189,32 +214,49 @@ void
InitializeShmemGUCs(void)
{
char buf[64];
- Size size_b;
- Size size_mb;
- Size hp_size;
+ size_t initial_b;
+ size_t min_b;
+ size_t max_b;
+ size_t hp_size;
+ size_t size_mb;
- /*
- * Calculate the shared memory size and round up to the nearest megabyte.
- */
- size_b = CalculateShmemSize();
- size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024);
+ CalculateShmemSize(&initial_b, &min_b, &max_b);
+
+ /* Round each size up to the nearest megabyte. */
+ size_mb = add_size(initial_b, (1024 * 1024) - 1) / (1024 * 1024);
sprintf(buf, "%zu", size_mb);
- SetConfigOption("shared_memory_size", buf,
+ SetConfigOption("shared_memory_initial_size", buf,
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
- /*
- * Calculate the number of huge pages required.
- */
+ size_mb = add_size(min_b, (1024 * 1024) - 1) / (1024 * 1024);
+ sprintf(buf, "%zu", size_mb);
+ SetConfigOption("shared_memory_minimum_size", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ size_mb = add_size(max_b, (1024 * 1024) - 1) / (1024 * 1024);
+ sprintf(buf, "%zu", size_mb);
+ SetConfigOption("shared_memory_maximum_size", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ /* Calculate the number of huge pages required for each size. */
GetHugePageSize(&hp_size, NULL);
if (hp_size != 0)
{
- Size hp_required;
+ size_t hp_required;
+
+ hp_required = initial_b / hp_size + (initial_b % hp_size != 0);
+ sprintf(buf, "%zu", hp_required);
+ SetConfigOption("shared_memory_initial_size_in_huge_pages", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
+
+ hp_required = min_b / hp_size + (min_b % hp_size != 0);
+ sprintf(buf, "%zu", hp_required);
+ SetConfigOption("shared_memory_minimum_size_in_huge_pages", buf,
+ PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
- hp_required = size_b / hp_size;
- if (size_b % hp_size != 0)
- hp_required = add_size(hp_required, 1);
+ hp_required = max_b / hp_size + (max_b % hp_size != 0);
sprintf(buf, "%zu", hp_required);
- SetConfigOption("shared_memory_size_in_huge_pages", buf,
+ SetConfigOption("shared_memory_maximum_size_in_huge_pages", buf,
PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT);
}
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index a3d56cf55dd..fbc05f0dc13 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -19,11 +19,11 @@
* methods). The routines in this file are used for allocating and
* binding to shared memory data structures.
*
- * This module provides facilities to allocate fixed-size structures in shared
- * memory, for things like variables shared between all backend processes.
- * Each such structure has a string name to identify it, specified when it is
- * requested. shmem_hash.c provides a shared hash table implementation on top
- * of that.
+ * This module provides facilities to allocate fixed-size as well as resizable
+ * structures in shared memory, for things like variables shared between all
+ * backend processes. Each such structure has a string name to identify it,
+ * specified when it is requested. shmem_hash.c provides a shared hash table
+ * implementation on top of fixed-size structures.
*
* Shared memory areas should usually not be allocated after postmaster
* startup, although we do allow small allocations later for the benefit of
@@ -102,6 +102,23 @@
* (*options->ptr), and calls the attach_fn callback, if any, for additional
* per-backend setup.
*
+ * Resizable shared memory structures
+ * ----------------------------------
+ *
+ * In order to allocate resizable shared memory structures, set
+ * ShmemRequestStructOpts::maximum_size to the maximum size that the structure
+ * can grow to. The address space for the maximum size will be reserved at
+ * startup, but memory is allocated or freed as the structure grows or shrinks
+ * respectively. ShmemRequestStructOpts::size should be set to the initial size
+ * of the structure, which is the amount of memory allocated at the startup.
+ * Optionally, ShmemRequestStructOpts::minimum_size can be set to the minimum
+ * size that the structure can shrink to. After startup, the structure can be
+ * resized by calling ShmemResizeStruct(). ShmemResizeStruct() enforces that the
+ * new size is within [minimum_size, maximum_size].
+ *
+ * While resizable structures can be created after the startup, the memory
+ * available for them is quite limited.
+ *
* Legacy ShmemInitStruct()/ShmemInitHash() functions
* --------------------------------------------------
*
@@ -268,6 +285,10 @@ typedef struct
void *location; /* location in shared mem */
Size size; /* # bytes requested for the structure */
Size allocated_size; /* # bytes actually allocated */
+ Size minimum_size; /* the minimum size the structure can shrink
+ * to */
+ Size maximum_size; /* the maximum size the structure can grow to */
+ Size reserved_space; /* the total address space reserved */
} ShmemIndexEnt;
/* To get reliable results for NUMA inquiry we need to "touch pages" once */
@@ -276,6 +297,8 @@ static bool firstNumaTouch = true;
static void CallShmemCallbacksAfterStartup(const ShmemCallbacks *callbacks);
static void InitShmemIndexEntry(ShmemRequest *request);
static bool AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok);
+static Size EstimateAllocatedSize(ShmemIndexEnt *entry);
+static void ShmemProtectStructInternal(ShmemIndexEnt *entry);
Datum pg_numa_available(PG_FUNCTION_ARGS);
@@ -341,25 +364,63 @@ ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind)
if (options->name == NULL)
elog(ERROR, "shared memory request is missing 'name' option");
+#ifndef HAVE_RESIZABLE_SHMEM
+ if (options->maximum_size > 0)
+ elog(ERROR, "resizable shared memory is not supported on this platform");
+ if (options->minimum_size > 0)
+ elog(ERROR, "resizable shared memory is not supported on this platform");
+#else
+ if (options->maximum_size > 0 && shared_memory_type != SHMEM_TYPE_MMAP)
+ elog(ERROR, "resizable shared memory requires shared_memory_type = mmap");
+#endif
+
if (IsUnderPostmaster)
{
if (options->size <= 0 && options->size != SHMEM_ATTACH_UNKNOWN_SIZE)
elog(ERROR, "invalid size %zd for shared memory request for \"%s\"",
options->size, options->name);
+ if (options->minimum_size < 0 && options->minimum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ elog(ERROR, "invalid minimum_size %zd for shared memory request for \"%s\"",
+ options->minimum_size, options->name);
+ if (options->maximum_size < 0 && options->maximum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ elog(ERROR, "invalid maximum_size %zd for shared memory request for \"%s\"",
+ options->maximum_size, options->name);
}
else
{
- if (options->size == SHMEM_ATTACH_UNKNOWN_SIZE)
+ if (options->size == SHMEM_ATTACH_UNKNOWN_SIZE ||
+ options->minimum_size == SHMEM_ATTACH_UNKNOWN_SIZE ||
+ options->maximum_size == SHMEM_ATTACH_UNKNOWN_SIZE)
elog(ERROR, "SHMEM_ATTACH_UNKNOWN_SIZE cannot be used during startup");
if (options->size <= 0)
elog(ERROR, "invalid size %zd for shared memory request for \"%s\"",
options->size, options->name);
+ if (options->minimum_size < 0)
+ elog(ERROR, "invalid minimum_size %zd for shared memory request for \"%s\"",
+ options->minimum_size, options->name);
+ if (options->maximum_size < 0)
+ elog(ERROR, "invalid maximum_size %zd for shared memory request for \"%s\"",
+ options->maximum_size, options->name);
}
if (options->alignment != 0 && pg_nextpower2_size_t(options->alignment) != options->alignment)
elog(ERROR, "invalid alignment %zu for shared memory request for \"%s\"",
options->alignment, options->name);
+ if (options->minimum_size > 0 && options->size != SHMEM_ATTACH_UNKNOWN_SIZE &&
+ options->minimum_size > options->size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have minimum size (%zd) less than or equal to size (%zd)",
+ options->name, options->minimum_size, options->size);
+
+ if (options->maximum_size > 0 && options->size > options->maximum_size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have maximum size (%zd) greater than size (%zd)",
+ options->name, options->maximum_size, options->size);
+
+ if (options->minimum_size > 0 && options->maximum_size > 0 &&
+ options->minimum_size > options->maximum_size)
+ elog(ERROR, "resizable shared memory structure \"%s\" should have minimum size (%zd) less than or equal to maximum size (%zd)",
+ options->name, options->minimum_size, options->maximum_size);
+
/* Check that we're in the right state */
if (shmem_request_state != SRS_REQUESTING)
elog(ERROR, "ShmemRequestStruct can only be called from a shmem_request callback");
@@ -381,36 +442,70 @@ ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind)
}
/*
- * ShmemGetRequestedSize() --- estimate the total size of all registered shared
- * memory structures.
+ * ShmemGetRequestedSize() --- estimate total size of all registered shared
+ * memory structures.
+ *
+ * Returns three totals:
+ * - initial - the sum of initial sizes of the requested structures which is the
+ * total amount of memory required at the startup.
+ * - min - the total of minimum sizes of structures
+ * - max - the sum of maximum sizes of the structures, which is the address
+ * space that must be reserved.
+ *
+ * When there are no resizable structures or on the platforms that do not
+ * support resizable structures, all three totals are the same.
*
* This is called at postmaster startup, before the shared memory segment has
* been created.
*/
-size_t
-ShmemGetRequestedSize(void)
+void
+ShmemGetRequestedSize(size_t *initial, size_t *min, size_t *max)
{
- size_t size;
+ size_t initial_size;
+ size_t min_size;
+ size_t max_size;
/* memory needed for the ShmemIndex */
- size = hash_estimate_size(list_length(pending_shmem_requests) + SHMEM_INDEX_ADDITIONAL_SIZE,
- sizeof(ShmemIndexEnt));
- size = CACHELINEALIGN(size);
+ initial_size = hash_estimate_size(list_length(pending_shmem_requests) + SHMEM_INDEX_ADDITIONAL_SIZE,
+ sizeof(ShmemIndexEnt));
+ initial_size = CACHELINEALIGN(initial_size);
+ min_size = initial_size;
+ max_size = initial_size;
/* memory needed for all the requested areas */
foreach_ptr(ShmemRequest, request, pending_shmem_requests)
{
size_t alignment = request->options->alignment;
+ size_t req_min;
+ size_t req_max;
+ size_t req_initial = request->options->size;
+
+ if (request->options->maximum_size > 0)
+ {
+ req_min = request->options->minimum_size;
+ req_max = request->options->maximum_size;
+ }
+ else
+ {
+ req_min = req_initial;
+ req_max = req_initial;
+ }
/* pad the start address for alignment like ShmemAllocRaw() does */
if (alignment < PG_CACHE_LINE_SIZE)
alignment = PG_CACHE_LINE_SIZE;
- size = TYPEALIGN(alignment, size);
+ initial_size = TYPEALIGN(alignment, initial_size);
+ min_size = TYPEALIGN(alignment, min_size);
+ max_size = TYPEALIGN(alignment, max_size);
- size = add_size(size, request->options->size);
+ initial_size = add_size(initial_size, req_initial);
+ min_size = add_size(min_size, req_min);
+ max_size = add_size(max_size, req_max);
}
- return size;
+ *initial = initial_size;
+ *min = min_size;
+ *max = max_size;
}
/*
@@ -431,6 +526,23 @@ ShmemInitRequested(void)
* Initialize the ShmemIndex entries and perform basic initialization of
* all the requested memory areas. There are no concurrent processes yet,
* so no need for locking.
+ *
+ * TODO: If we have resizable structures, we will mmap with MAP_NORESERVE.
+ * In case there is not enough memory to cover the initial sizes of the
+ * shared structures, the mmap will succeed but the initialization will
+ * fail with SIGBUS. Instead we should allocate the memory worth the
+ * initial size of each structure using madvise(MADV_WRITE_POPOULATE). But
+ * instead of doing that for each structure, we should combine contiguous
+ * structures and do it once for the whole range. The algorithm to use is
+ * as follows 1. start from the first request 2. note the start address of
+ * the first structure in the range 3. walk upto a resizable structure or
+ * the last request, noting the initial end address of the last structure
+ * in the range 4. call madvise(MADV_WRITE_POPULATE) for the range (start
+ * address, end address) 5. Next structure becomes the first structure in
+ * the next range, repeat from step 2
+ *
+ * If there are no resizable structure, no need to do anything, as all the
+ * memory is reserved during mmap itself.
*/
foreach_ptr(ShmemRequest, request, pending_shmem_requests)
{
@@ -514,6 +626,7 @@ InitShmemIndexEntry(ShmemRequest *request)
ShmemIndexEnt *index_entry;
bool found;
size_t allocated_size;
+ size_t requested_size;
void *structPtr;
/* look it up in the shmem index */
@@ -531,10 +644,19 @@ InitShmemIndexEntry(ShmemRequest *request)
}
/*
- * We inserted the entry to the shared memory index. Allocate requested
- * amount of shared memory for it, and initialize the index entry.
+ * We inserted the entry to the shared memory index. Allocate requested
+ * amount of address space in the shared memory segment for it, and do
+ * basic initializion. The memory gets allocated during initialization as
+ * the corresponding memory pages are written to. Allocate enough space
+ * for a resizable structure to grow to its maximum size. It is expected
+ * that the initialization callback will use only as much memory as the
+ * initial size of the resizable structure. (Well, if it doesn't, more
+ * memory will be allocated initially than expected, no further harm is
+ * done.)
*/
- structPtr = ShmemAllocRaw(request->options->size,
+ requested_size = request->options->maximum_size > 0 ?
+ request->options->maximum_size : request->options->size;
+ structPtr = ShmemAllocRaw(requested_size,
request->options->alignment,
&allocated_size);
if (structPtr == NULL)
@@ -543,13 +665,36 @@ InitShmemIndexEntry(ShmemRequest *request)
hash_search(ShmemIndex, name, HASH_REMOVE, NULL);
ereport(ERROR,
(errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("not enough shared memory for data structure"
+ errmsg("not enough shared memory space for data structure"
" \"%s\" (%zd bytes requested)",
- name, request->options->size)));
+ name, requested_size)));
}
index_entry->size = request->options->size;
index_entry->allocated_size = allocated_size;
index_entry->location = structPtr;
+ index_entry->reserved_space = allocated_size;
+ if (request->options->maximum_size > 0)
+ {
+ index_entry->minimum_size = request->options->minimum_size;
+ index_entry->maximum_size = request->options->maximum_size;
+
+ /* Adjust allocated size of a resizable structure. */
+ index_entry->allocated_size = EstimateAllocatedSize(index_entry);
+
+ /*
+ * Protect the unused part of the reserved address space for a
+ * resizable structure. If the structure has same minimum and maximum
+ * size, it is effectively a fixed-size structure without any unused
+ * space; no protection is required.
+ */
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ ShmemProtectStructInternal(index_entry);
+ }
+ else
+ {
+ index_entry->minimum_size = request->options->size;
+ index_entry->maximum_size = request->options->size;
+ }
/* Initialize depending on the kind of shmem area it is */
switch (request->kind)
@@ -594,7 +739,7 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
return false;
}
- /* Check that the size in the index matches the request */
+ /* Check that the sizes in the index match the request. */
if (index_entry->size != request->options->size &&
request->options->size != SHMEM_ATTACH_UNKNOWN_SIZE)
{
@@ -604,6 +749,43 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
name, index_entry->size, request->options->size)));
}
+ /*
+ * For resizable structures, also check that minimum_size and maximum_size
+ * match. For fixed-size structures, these are derived (set to size) in
+ * the index entry and not meaningful in the request.
+ */
+ if (request->options->maximum_size != 0)
+ {
+ if (index_entry->minimum_size != request->options->minimum_size &&
+ request->options->minimum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ {
+ ereport(ERROR,
+ errmsg("shared memory struct \"%s\" was created with"
+ " different minimum_size: existing %zu, requested %zu",
+ name, index_entry->minimum_size,
+ request->options->minimum_size));
+ }
+
+ if (index_entry->maximum_size != request->options->maximum_size &&
+ request->options->maximum_size != SHMEM_ATTACH_UNKNOWN_SIZE)
+ {
+ ereport(ERROR,
+ errmsg("shared memory struct \"%s\" was created with"
+ " different maximum_size: existing %zu, requested %zu",
+ name, index_entry->maximum_size,
+ request->options->maximum_size));
+ }
+ }
+ else
+ {
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ elog(ERROR, "shared memory struct \"%s\" was created as resizable, but requested as fixed-size",
+ name);
+ }
+
+ if (index_entry->minimum_size != index_entry->maximum_size)
+ ShmemProtectStructInternal(index_entry);
+
/*
* Re-establish the caller's pointer variable, or do other actions to
* attach depending on the kind of shmem area it is.
@@ -625,6 +807,291 @@ AttachShmemIndexEntry(ShmemRequest *request, bool missing_ok)
return true;
}
+/*
+ * Estimate the actual memory allocated for a resizable structure.
+ *
+ * ... based on the assumption that the memory is allocated in pages.
+ *
+ * The memory pages covered by the current size of a resizable structure are
+ * considered to be allocated. The memory page where the maximal structure ends
+ * also hosts the next structure, unless the maximal structure ends on a page
+ * boundary. Hence that page is allocated because of the next structure. The
+ * memory pages between the page where the current structure ends and the page
+ * where the next structure starts remain unallocated. Thus the memory allocated
+ * for a resizable structure can be estimated as the total address space
+ * reserved for the structure minus the unallocated memory pages between the
+ * current end and the next structure.
+ */
+static Size
+EstimateAllocatedSize(ShmemIndexEnt *entry)
+{
+ Size page_size = GetOSPageSize();
+ char *align_end = (char *) TYPEALIGN(page_size, (char *) entry->location + entry->size);
+ char *floor_max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) entry->location + entry->maximum_size);
+
+ Assert(entry->maximum_size >= entry->size);
+ Assert(entry->reserved_space >= entry->maximum_size);
+
+ if (align_end < floor_max_end)
+ return entry->reserved_space - (floor_max_end - align_end);
+
+ return entry->reserved_space;
+}
+
+/*
+ * ShmemResizeStruct() --- resize a resizable shared memory structure.
+ *
+ * The new size must be within [minimum_size, maximum_size]. If the structure
+ * is being shrunk, the memory pages that are no longer needed are freed. If
+ * the structure is being expanded, the memory pages that are needed for the
+ * new size are allocated. See EstimateAllocatedSize() for explanation of which
+ * pages are allocated for a resizable structure.
+ *
+ * The caller must ensure that no other backend is accessing the part of the
+ * structure between the old size and the new size while this function is
+ * running. It should also ensure that all backends that may access the
+ * structure have observed the new size before they access the range between
+ * the old size and the new size.
+ *
+ * If we can not allocate memory pages when expanding the structure, this
+ * function will return false. On success it returns true. Instead
+ * of returning false, an error is raised if we can not free the memory pages
+ * when shrinking the structure, which should not happen in practice.
+ */
+bool
+ShmemResizeStruct(const char *name, Size new_size)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+ pg_unreachable();
+#else
+ ShmemIndexEnt *result;
+ bool found;
+ Size page_size = GetOSPageSize();
+ char *new_end;
+ bool success = true;
+
+ Assert(new_size > 0);
+
+ /*
+ * Resizable shared memory structures are only supported with mmap'ed
+ * memory.
+ */
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
+
+ /* look it up in the shmem index */
+ LWLockAcquire(ShmemIndexLock, LW_EXCLUSIVE);
+ result = (ShmemIndexEnt *) hash_search(ShmemIndex, name, HASH_FIND, &found);
+ if (!found)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shmem struct \"%s\" is not initialized", name));
+
+ Assert(result);
+
+ if (result->minimum_size == result->maximum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shared memory struct \"%s\" is not resizable", name));
+
+ if (new_size < result->minimum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("cannot shrink shared memory structure \"%s\" below minimum size"
+ "(requested %zu bytes, minimum %zu bytes)",
+ name, new_size, result->minimum_size));
+
+ if (result->maximum_size < new_size)
+ ereport(ERROR,
+ errcode(ERRCODE_INSUFFICIENT_RESOURCES),
+ errmsg("not enough address space is reserved for resizing structure \"%s\""
+ " (required %zu bytes, reserved %zu bytes)",
+ name, new_size, result->maximum_size));
+
+ /*
+ * A structure requires memory pages from the page containing the start of
+ * the structure and current end of the structure to be allocated. When
+ * expanding, we make sure that pages are allocated up to the new end.
+ * When shrinking, release memory pages beyond the new end, but not the
+ * page containing maximal end of the structure, as it may be used by the
+ * next structure.
+ *
+ * We do not consider the current end of the structure as it simplifies
+ * the calculations. Instead we rely on the underlying APIs not to touch
+ * the memory pages that will not be affected by the change in size.
+ */
+ new_end = (char *) TYPEALIGN(page_size, (char *) result->location + new_size);
+ if (new_size < result->size)
+ {
+ char *max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+
+ if (max_end > new_end)
+ {
+ if (!PGSharedMemoryEnsureFreed(new_end, max_end - new_end))
+ ereport(ERROR,
+ errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not free %zu bytes of shared memory from structure \"%s\"",
+ result->size - new_size, name));
+ }
+ }
+ else if (new_size > result->size)
+ {
+ char *struct_start = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location);
+
+ if (new_end > struct_start)
+ {
+ ShmemIndexEnt entry_copy = *result;
+
+ /*
+ * Allocating memory pages in the expanded range may require the
+ * corresponding address space to have read-write access.
+ */
+ entry_copy.size = new_size;
+ ShmemProtectStructInternal(&entry_copy);
+
+ if (!PGSharedMemoryEnsureAllocated(struct_start, new_end - struct_start))
+ {
+ ShmemProtectStructInternal(result);
+ ereport(WARNING,
+ errcode(ERRCODE_OUT_OF_MEMORY),
+ errmsg("could not allocate %zu bytes of shared memory to structure \"%s\"",
+ new_size - result->size, name));
+ success = false;
+ }
+ }
+ }
+
+ /* Update shmem index entry. */
+ if (success)
+ {
+ result->size = new_size;
+ result->allocated_size = EstimateAllocatedSize(result);
+ }
+
+ LWLockRelease(ShmemIndexLock);
+
+ return success;
+#endif
+}
+
+/*
+ * ShmemProtectStruct() --- protect the unused portion of the given resizable
+ * structure.
+ *
+ * Makes the region beyond the current size up to maximum_size inaccessible, and
+ * ensures the region up to the current size is readable and writable. Depending
+ * upon the platform, the protection honours the page boundaries. So it may be
+ * more permissible than strictly needed.
+ *
+ * This function only affects the calling backend's address space. After each
+ * ShmemResizeStruct(), every backend that may access the structure should call
+ * this function before its next access. When backends are still in the middle
+ * of that round of ShmemProtectStruct() calls, ShmemResizeStruct() should not
+ * be called.
+ */
+void
+ShmemProtectStruct(const char *name)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ ShmemIndexEnt *result;
+ bool found;
+
+ LWLockAcquire(ShmemIndexLock, LW_SHARED);
+
+ result = (ShmemIndexEnt *) hash_search(ShmemIndex, name, HASH_FIND, &found);
+ if (!found)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shmem struct \"%s\" is not initialized", name));
+
+ if (result->minimum_size == result->maximum_size)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("shared memory struct \"%s\" is not resizable", name));
+
+ ShmemProtectStructInternal(result);
+
+ LWLockRelease(ShmemIndexLock);
+#endif
+}
+
+/*
+ * ShmemProtectStructInternal --- same as ShmemProtectStruct()
+ *
+ * ..., but called when the ShmemIndexEnt of the struct is available. The caller
+ * should hold ShmemIndexLock if required.
+ */
+static void
+ShmemProtectStructInternal(ShmemIndexEnt *entry)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizable shared memory is not supported on this platform"));
+#else
+ Size page_size = GetOSPageSize();
+ char *rw_start;
+ char *rw_end;
+ char *prot_end;
+
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
+ Assert(entry->minimum_size != entry->maximum_size);
+
+ /* Make at least [location, location+size) readable and writable */
+ rw_start = (char *) TYPEALIGN_DOWN(page_size, entry->location);
+ rw_end = (char *) TYPEALIGN(page_size,
+ (char *) entry->location + entry->size);
+
+ /*
+ * Make remaining portion inaccessible while making sure that the portion
+ * after maximum_size is not affected since it may be used by other
+ * structures.
+ */
+ prot_end = (char *) TYPEALIGN_DOWN(page_size,
+ (char *) entry->location + entry->maximum_size);
+
+ if (!PGSharedMemoryProtect(rw_start, rw_end, prot_end))
+ ereport(ERROR,
+ errcode(ERRCODE_SYSTEM_ERROR),
+ errmsg("could not protect shared memory structure \"%s\"", entry->key));
+#endif
+}
+
+/*
+ * ShmemReprotectResizableStructs() --- re-apply per-backend protection to every
+ * resizable shmem structure.
+ *
+ * On !EXEC_BACKEND platforms, a new backend inherits shmem mappings and their
+ * per-process protections from the postmaster. If any resizable structure has
+ * been resized since postmaster start, the inherited protections no longer
+ * match the current size. This is called early in per-backend startup so the
+ * new backend sees the protections according to the current sizes of the
+ * resizable structures.
+ */
+void
+ShmemReprotectResizableStructs(void)
+{
+#ifdef HAVE_RESIZABLE_SHMEM
+ HASH_SEQ_STATUS status;
+ ShmemIndexEnt *entry;
+
+ LWLockAcquire(ShmemIndexLock, LW_SHARED);
+ hash_seq_init(&status, ShmemIndex);
+ while ((entry = (ShmemIndexEnt *) hash_seq_search(&status)) != NULL)
+ {
+ if (entry->minimum_size != entry->maximum_size)
+ ShmemProtectStructInternal(entry);
+ }
+ LWLockRelease(ShmemIndexLock);
+#endif
+}
+
/*
* InitShmemAllocator() --- set up basic pointers to shared memory.
*
@@ -731,6 +1198,9 @@ InitShmemAllocator(PGShmemHeader *seghdr)
Assert(!found);
result->size = ShmemAllocator->index_size;
result->allocated_size = ShmemAllocator->index_size;
+ result->minimum_size = result->size;
+ result->maximum_size = result->size;
+ result->reserved_space = result->allocated_size;
result->location = ShmemAllocator->index;
}
}
@@ -1047,7 +1517,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr)
Datum
pg_get_shmem_allocations(PG_FUNCTION_ARGS)
{
-#define PG_GET_SHMEM_SIZES_COLS 4
+#define PG_GET_SHMEM_SIZES_COLS 7
ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
HASH_SEQ_STATUS hstat;
ShmemIndexEnt *ent;
@@ -1069,7 +1539,20 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr);
values[2] = Int64GetDatum(ent->size);
values[3] = Int64GetDatum(ent->allocated_size);
- named_allocated += ent->allocated_size;
+ values[4] = Int64GetDatum(ent->minimum_size);
+ values[5] = Int64GetDatum(ent->maximum_size);
+ values[6] = Int64GetDatum(ent->reserved_space);
+
+ /*
+ * Anonymous areas are allocated in the area remaining after all named
+ * areas have been allocated. Thus the amount of shared memory
+ * allocated for anonymous areas can be calculated as the total amount
+ * of space allocated minus the amount of space allocated for named
+ * areas. The amount of free shared memory at the end of the segment
+ * can be calculated as the total size of the segment minus the total
+ * amount of space allocated.
+ */
+ named_allocated += ent->reserved_space;
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
values, nulls);
@@ -1080,6 +1563,9 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
nulls[1] = true;
values[2] = Int64GetDatum(ShmemAllocator->free_offset - named_allocated);
values[3] = values[2];
+ values[4] = values[2];
+ values[5] = values[2];
+ values[6] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
/* output as-of-yet unused shared memory */
@@ -1088,6 +1574,9 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS)
nulls[1] = false;
values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemAllocator->free_offset);
values[3] = values[2];
+ values[4] = values[2];
+ values[5] = values[2];
+ values[6] = values[2];
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
LWLockRelease(ShmemIndexLock);
@@ -1274,23 +1763,9 @@ pg_get_shmem_allocations_numa(PG_FUNCTION_ARGS)
Size
pg_get_shmem_pagesize(void)
{
- Size os_page_size;
-#ifdef WIN32
- SYSTEM_INFO sysinfo;
-
- GetSystemInfo(&sysinfo);
- os_page_size = sysinfo.dwPageSize;
-#else
- os_page_size = sysconf(_SC_PAGESIZE);
-#endif
-
Assert(IsUnderPostmaster);
- Assert(huge_pages_status != HUGE_PAGES_UNKNOWN);
-
- if (huge_pages_status == HUGE_PAGES_ON)
- GetHugePageSize(&os_page_size, NULL);
- return os_page_size;
+ return GetOSPageSize();
}
Datum
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 9d6e69175a5..780cdb5e9ff 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -575,6 +575,14 @@ InitProcess(void)
if (IsUnderPostmaster)
AttachSharedMemoryStructs();
#endif
+
+ /*
+ * Update access to the address space occupied by the resizable shared
+ * structures. We do it here so that the structures can be accessed safely
+ * by this backend. But we will do this again after ProcSignalInit() for
+ * the reasons mentioned there.
+ */
+ ShmemReprotectResizableStructs();
}
/*
@@ -756,6 +764,14 @@ InitAuxiliaryProcess(void)
if (IsUnderPostmaster)
AttachSharedMemoryStructs();
#endif
+
+ /*
+ * Update access to the address space occupied by the resizable shared
+ * structures. We do it here so that the structures can be accessed safely
+ * by this backend. But we will do this again after ProcSignalInit() for
+ * the reasons mentioned there.
+ */
+ ShmemReprotectResizableStructs();
}
/*
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 3d8c9bdebd5..815b865aa38 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -787,6 +787,16 @@ InitPostgres(const char *in_dbname, Oid dboid,
*/
InitLocalDataChecksumState();
+ /*
+ * Refresh per-backend protections for resizable shmem structures. Usually
+ * the subsystems using resizable shared structures will use
+ * ProcSignalBarrier mechanism to coordinate resizing which would involve
+ * adjusting the protections as well. Like InitLocalDataChecksumState()
+ * above, this must run after ProcSignalInit so as not to miss a barrier
+ * for protection change.
+ */
+ ShmemReprotectResizableStructs();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index 66ea4afe920..7a9bee3ed3f 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -1226,6 +1226,13 @@
max => '1000.0',
},
+{ name => 'have_resizable_shmem', type => 'bool', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows whether the running server supports resizable shared memory.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE',
+ variable => 'have_resizable_shmem_enabled',
+ boot_val => 'HAVE_RESIZABLE_SHMEM_ENABLED',
+},
+
{ name => 'hba_file', type => 'string', context => 'PGC_POSTMASTER', group => 'FILE_LOCATIONS',
short_desc => 'Sets the server\'s "hba" configuration file.',
flags => 'GUC_SUPERUSER_ONLY',
@@ -2722,20 +2729,58 @@
max => 'INT_MAX / 2',
},
-{ name => 'shared_memory_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
- short_desc => 'Shows the size of the server\'s main shared memory area (rounded up to the nearest MB).',
+{ name => 'shared_memory_initial_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the amount of memory allocated at server startup in the main shared memory area (rounded up to the nearest MB).',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_initial_size_mb',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_initial_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the number of huge pages needed in the main shared memory area at server startup.',
+ long_desc => '-1 means huge pages are not supported.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_initial_size_in_huge_pages',
+ boot_val => '-1',
+ min => '-1',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_maximum_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the size of the main shared memory area and maximum memory that can be allocated in that area (rounded up to the nearest MB).',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_maximum_size_mb',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_maximum_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the maximum number of huge pages needed in the main shared memory area.',
+ long_desc => '-1 means huge pages are not supported.',
+ flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
+ variable => 'shared_memory_maximum_size_in_huge_pages',
+ boot_val => '-1',
+ min => '-1',
+ max => 'INT_MAX',
+},
+
+{ name => 'shared_memory_minimum_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the minimum amount of memory required in the main shared memory area (rounded up to the nearest MB).',
flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED',
- variable => 'shared_memory_size_mb',
+ variable => 'shared_memory_minimum_size_mb',
boot_val => '0',
min => '0',
max => 'INT_MAX',
},
-{ name => 'shared_memory_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
- short_desc => 'Shows the number of huge pages needed for the main shared memory area.',
+{ name => 'shared_memory_minimum_size_in_huge_pages', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
+ short_desc => 'Shows the minimum number of huge pages needed in the main shared memory area.',
long_desc => '-1 means huge pages are not supported.',
flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_RUNTIME_COMPUTED',
- variable => 'shared_memory_size_in_huge_pages',
+ variable => 'shared_memory_minimum_size_in_huge_pages',
boot_val => '-1',
min => '-1',
max => 'INT_MAX',
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 1ec460b6a82..17d50ffde38 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -643,8 +643,12 @@ static int max_index_keys;
static int max_identifier_length;
static int block_size;
static int segment_size;
-static int shared_memory_size_mb;
-static int shared_memory_size_in_huge_pages;
+static int shared_memory_initial_size_mb;
+static int shared_memory_minimum_size_mb;
+static int shared_memory_maximum_size_mb;
+static int shared_memory_initial_size_in_huge_pages;
+static int shared_memory_minimum_size_in_huge_pages;
+static int shared_memory_maximum_size_in_huge_pages;
static int wal_block_size;
static int num_os_semaphores;
static int effective_wal_level = WAL_LEVEL_REPLICA;
@@ -664,6 +668,13 @@ static bool assert_enabled = DEFAULT_ASSERT_ENABLED;
#endif
static bool exec_backend_enabled = EXEC_BACKEND_ENABLED;
+#ifdef HAVE_RESIZABLE_SHMEM
+#define HAVE_RESIZABLE_SHMEM_ENABLED true
+#else
+#define HAVE_RESIZABLE_SHMEM_ENABLED false
+#endif
+static bool have_resizable_shmem_enabled = HAVE_RESIZABLE_SHMEM_ENABLED;
+
static char *recovery_target_timeline_string;
static char *recovery_target_string;
static char *recovery_target_xid_string;
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index f8a021987b5..712172760b3 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8692,8 +8692,8 @@
{ oid => '5052', descr => 'allocations from the main shared memory segment',
proname => 'pg_get_shmem_allocations', prorows => '50', proretset => 't',
provolatile => 'v', prorettype => 'record', proargtypes => '',
- proallargtypes => '{text,int8,int8,int8}', proargmodes => '{o,o,o,o}',
- proargnames => '{name,off,size,allocated_size}',
+ proallargtypes => '{text,int8,int8,int8,int8,int8,int8}', proargmodes => '{o,o,o,o,o,o,o}',
+ proargnames => '{name,off,size,allocated_size,minimum_size,maximum_size,reserved_space}',
prosrc => 'pg_get_shmem_allocations',
proacl => '{POSTGRES=X,pg_read_all_stats=X}' },
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 661c4a9b168..661c12e9bf3 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,14 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `MADV_POPULATE_WRITE', and to 0
+ if you don't. */
+#undef HAVE_DECL_MADV_POPULATE_WRITE
+
+/* Define to 1 if you have the declaration of `MADV_REMOVE', and to 0 if you
+ don't. */
+#undef HAVE_DECL_MADV_REMOVE
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/include/pg_config_manual.h b/src/include/pg_config_manual.h
index 521b49b8888..ab944babe2b 100644
--- a/src/include/pg_config_manual.h
+++ b/src/include/pg_config_manual.h
@@ -131,6 +131,20 @@
#define EXEC_BACKEND
#endif
+/*
+ * HAVE_RESIZABLE_SHMEM indicates whether resizable shared memory structures are
+ * supported. The implementation requires Linux-specific madvise constants
+ * (MADV_REMOVE and MADV_POPULATE_WRITE) and existence of mprotect() API.
+ *
+ * TODO: We may want to remove EXEC_BACKEND from the condition to test attaching
+ * and resizing resizable shared memory structures in EXEC_BACKEND mode. Windows
+ * will anyway won't have HAVE_RESIZABLE_SHMEM defined since it won't have
+ * MADV_REMOVE and MADV_POPULATE_WRITE.
+ */
+#if HAVE_DECL_MADV_REMOVE && HAVE_DECL_MADV_POPULATE_WRITE && !defined(EXEC_BACKEND)
+#define HAVE_RESIZABLE_SHMEM
+#endif
+
/*
* USE_POSIX_FADVISE controls whether Postgres will attempt to use the
* posix_fadvise() kernel call. Usually the automatic configure tests are
diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h
index b205b00e7a1..46ae87fe863 100644
--- a/src/include/storage/ipc.h
+++ b/src/include/storage/ipc.h
@@ -78,7 +78,7 @@ extern void check_on_shmem_exit_lists_are_empty(void);
extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook;
extern void RegisterBuiltinShmemCallbacks(void);
-extern Size CalculateShmemSize(void);
+extern void CalculateShmemSize(size_t *initial, size_t *min, size_t *max);
extern void CreateSharedMemoryAndSemaphores(void);
#ifdef EXEC_BACKEND
extern void AttachSharedMemoryStructs(void);
diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h
index 10c7b065861..50474c2d579 100644
--- a/src/include/storage/pg_shmem.h
+++ b/src/include/storage/pg_shmem.h
@@ -85,10 +85,14 @@ extern void PGSharedMemoryReAttach(void);
extern void PGSharedMemoryNoReAttach(void);
#endif
-extern PGShmemHeader *PGSharedMemoryCreate(Size size,
+extern PGShmemHeader *PGSharedMemoryCreate(Size initial_size, Size max_size,
PGShmemHeader **shim);
extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2);
extern void PGSharedMemoryDetach(void);
+extern bool PGSharedMemoryEnsureFreed(void *addr, Size size);
+extern bool PGSharedMemoryEnsureAllocated(void *addr, Size size);
+extern bool PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end);
extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags);
+extern size_t GetOSPageSize(void);
#endif /* PG_SHMEM_H */
diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h
index 43b636868c7..947debcede4 100644
--- a/src/include/storage/shmem.h
+++ b/src/include/storage/shmem.h
@@ -57,6 +57,22 @@ typedef struct ShmemStructOpts
*/
size_t alignment;
+ /*
+ * Minimum size this structure can shrink to. Should be set to 0 for
+ * fixed-size structures.
+ */
+ ssize_t minimum_size;
+
+ /*
+ * Maximum size this structure can grow upto in future. The memory is not
+ * allocated right away but the corresponding address space is reserved so
+ * that memory can be mapped to it when the structure grows. Typically
+ * should be used for large resizable structures which need several pages
+ * worth of contiguous memory. Should be set to 0 for fixed-size
+ * structures.
+ */
+ ssize_t maximum_size;
+
/*
* When the shmem area is initialized or attached to, pointer to it is
* stored in *ptr. It usually points to a global variable, used to access
@@ -168,6 +184,9 @@ typedef struct ShmemCallbacks
extern void RegisterShmemCallbacks(const ShmemCallbacks *callbacks);
extern bool ShmemAddrIsValid(const void *addr);
+extern bool ShmemResizeStruct(const char *name, Size new_size);
+extern void ShmemProtectStruct(const char *name);
+extern void ShmemReprotectResizableStructs(void);
/*
* These macros provide syntactic sugar for calling the underlying functions
diff --git a/src/include/storage/shmem_internal.h b/src/include/storage/shmem_internal.h
index 8746b614fa3..6c8c81812a8 100644
--- a/src/include/storage/shmem_internal.h
+++ b/src/include/storage/shmem_internal.h
@@ -36,7 +36,7 @@ extern void ResetShmemAllocator(void);
extern void ShmemRequestInternal(ShmemStructOpts *options, ShmemRequestKind kind);
-extern size_t ShmemGetRequestedSize(void);
+extern void ShmemGetRequestedSize(size_t *initial, size_t *min, size_t *max);
extern void ShmemInitRequested(void);
#ifdef EXEC_BACKEND
extern void ShmemAttachRequested(void);
diff --git a/src/test/modules/test_shmem/meson.build b/src/test/modules/test_shmem/meson.build
index fb4bf328b8f..1f795aae7eb 100644
--- a/src/test/modules/test_shmem/meson.build
+++ b/src/test/modules/test_shmem/meson.build
@@ -27,7 +27,8 @@ tests += {
'bd': meson.current_build_dir(),
'tap': {
'tests': [
- 't/001_late_shmem_alloc.pl',
+ 't/001_fixed_shmem_struct.pl',
+ 't/002_resizable_shmem_struct.pl',
],
},
}
diff --git a/src/test/modules/test_shmem/t/001_late_shmem_alloc.pl b/src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
similarity index 58%
rename from src/test/modules/test_shmem/t/001_late_shmem_alloc.pl
rename to src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
index 5cf07d071ec..d231821e0b4 100644
--- a/src/test/modules/test_shmem/t/001_late_shmem_alloc.pl
+++ b/src/test/modules/test_shmem/t/001_fixed_shmem_struct.pl
@@ -56,5 +56,36 @@ else
);
}
+###
+# Test that a fixed-size shared memory structure cannot be resized.
+# Only relevant on platforms that support resizable shmem.
+###
+my $have_resizable_shmem =
+ $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+if ($have_resizable_shmem)
+{
+ # Try expanding the fixed-size structure
+ my ($ret, $stdout, $stderr) =
+ $node->psql("postgres", "SELECT test_shmem_resize_fixed(1000);");
+ isnt($ret, 0, "expanding a fixed-size structure fails");
+ like($stderr, qr/is not resizable/, "expand error message mentions not resizable");
+
+ # Try shrinking the fixed-size structure
+ ($ret, $stdout, $stderr) =
+ $node->psql("postgres", "SELECT test_shmem_resize_fixed(1);");
+ isnt($ret, 0, "shrinking a fixed-size structure fails");
+ like($stderr, qr/is not resizable/, "shrink error message mentions not resizable");
+}
+
+###
+# Test that minimum_size and maximum_size equal size for a fixed-size structure
+# in pg_shmem_allocations.
+###
+is($node->safe_psql('postgres',
+ "SELECT minimum_size = size AND maximum_size = size FROM pg_shmem_allocations WHERE name = 'test_shmem area';"),
+ 't', "fixed-size structure has minimum_size = maximum_size = size");
+
$node->stop;
+
done_testing();
diff --git a/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl b/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
new file mode 100644
index 00000000000..d35a5d7e031
--- /dev/null
+++ b/src/test/modules/test_shmem/t/002_resizable_shmem_struct.pl
@@ -0,0 +1,387 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Test resizable shared memory functionality, both when loaded at startup via
+# shared_preload_libraries and when loaded after startup (late allocation).
+
+# Verify that enough shared memory is allocated to cover the resizable_shmem
+# structure at its current size but does not exceed the memory required by the
+# current sizes of all shared memory structures. We expect that the backend
+# where we run the query will have touched the entire resizable_shmem structure,
+# so that all the memory pages covering the resizable structure are mapped to
+# the backend's address space.
+#
+# Since we have configured the server so that resizable shared struture
+# dominates the main shared memory segment, the total memory allocated to other
+# shared memory structures does not result in false positive tests below.
+sub check_shmem_usage
+{
+ my ($session, $label, $node) = @_;
+
+ my $shmem_usage = $session->query_safe('SELECT test_shmem_usage();', verbose => 0);
+ my $total_alloc = $node->safe_psql('postgres', "SELECT sum(allocated_size) FROM pg_shmem_allocations;");
+ my $resizable_alloc = $node->safe_psql('postgres',
+ "SELECT allocated_size FROM pg_shmem_allocations WHERE name = 'resizable_shmem';");
+
+ diag "$label: shmem_usage=$shmem_usage, resizable_shmem allocated=$resizable_alloc, sum(allocated_size)=$total_alloc";
+ ok($shmem_usage <= $total_alloc,
+ "$label: allocated shared memory does not exceed total allocated size");
+ ok($shmem_usage >= $resizable_alloc,
+ "$label: shared memory usage covers the resizable_shmem allocation");
+}
+
+# Test a resize operation: resize, verify old data, write new data, verify
+# new data, and check shmem usage. Returns updated ($num_entries, $value).
+sub test_resize
+{
+ my ($node, $prefix, $old_num_entries, $old_value, $new_num_entries, $new_value, $label) = @_;
+
+ $label = "$prefix: $label";
+
+ my $session1 = $node->background_psql('postgres');
+ my $session2 = $node->background_psql('postgres');
+
+ $session1->query_safe("SELECT resizable_shmem_resize($new_num_entries);",
+ verbose => 0);
+
+ # Old data should still be intact in the (possibly smaller) area
+ my $readable_entries = ($new_num_entries < $old_num_entries) ? $new_num_entries : $old_num_entries;
+ is($session1->query_safe("SELECT resizable_shmem_read($readable_entries, $old_value);",
+ verbose => 0),
+ 't', "old data readable after $label");
+
+ $session2->query_safe("SELECT resizable_shmem_write($new_value);",
+ verbose => 0);
+ is($session1->query_safe("SELECT resizable_shmem_read($new_num_entries, $new_value);",
+ verbose => 0),
+ 't', "new data readable after $label");
+
+ check_shmem_usage($session1, "$label (session 1)", $node);
+ check_shmem_usage($session2, "$label (session 2)", $node);
+
+ $session1->quit;
+ $session2->quit;
+
+ return ($new_num_entries, $new_value);
+}
+
+# Verify that reads or writes past the current size, but within the reserved
+# maximum, fault when they reach the protected region.
+sub test_fault_beyond_size
+{
+ my ($node, $initial_entries, $prefix) = @_;
+
+ # Enable restart_after_crash to test postmaster driven restart with
+ # resizable shared memory.
+ $node->safe_psql('postgres',
+ 'ALTER SYSTEM SET restart_after_crash = on;');
+ $node->reload;
+
+ for my $mode ('write', 'read')
+ {
+ my ($ret, $stdout, $stderr) = $node->psql('postgres',
+ "SELECT resizable_shmem_access_beyond_size('$mode');");
+ ok($ret != 0, "$prefix: $mode past current size crashes the backend");
+ like($stderr,
+ qr/server closed the connection unexpectedly|connection to server was lost/,
+ "$prefix: $mode crash reports lost connection");
+
+ $node->poll_query_until('postgres', 'SELECT 1', '1')
+ or die "server did not come back after $mode crash";
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT resizable_shmem_read($initial_entries, 0);"),
+ 't', "$prefix: read succeeds after crash recovery");
+
+ $node->safe_psql('postgres', 'ALTER SYSTEM RESET restart_after_crash;');
+ $node->reload;
+}
+
+# Run the full suite of resizable shared memory tests on the given node.
+sub run_resizable_tests
+{
+ my ($node, $initial_entries, $max_entries, $prefix) = @_;
+ my $have_resizable_shmem = $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+ my $num_entries = $initial_entries;
+
+ # Basic read/write should work on all platforms
+ my $value = 100;
+ $node->safe_psql('postgres', "SELECT resizable_shmem_write($value);");
+ is($node->safe_psql('postgres', "SELECT resizable_shmem_read($num_entries, $value);"),
+ 't', "$prefix: data read after write successful");
+
+ if ($have_resizable_shmem)
+ {
+ # Initial structure state
+ my $session1 = $node->background_psql('postgres');
+ my $session2 = $node->background_psql('postgres');
+
+ $value = 100;
+ # Write and read the initial set of entries.
+ $session1->query_safe("SELECT resizable_shmem_write($value);", verbose => 0);
+ is($session2->query_safe("SELECT resizable_shmem_read($num_entries, $value);",
+ verbose => 0),
+ 't', "$prefix: data read after write successful");
+ check_shmem_usage($session1, "$prefix: initial write (session 1)", $node);
+ check_shmem_usage($session2, "$prefix: initial write (session 2)", $node);
+ $session1->quit;
+ $session2->quit;
+
+ # Verify no other structure is resizable
+ is($node->safe_psql('postgres', "SELECT count(*) FROM pg_shmem_allocations WHERE name <> 'resizable_shmem' AND maximum_size <> minimum_size;"),
+ '0', "$prefix: no other resizable structures");
+
+ # Resize to maximum
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ $max_entries, 500, 'resize to maximum');
+
+ # Shrink to 75% of max
+ my $shrink_entries = int($max_entries * 3 / 4);
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ $shrink_entries, 999, 'shrinking');
+
+ # Resize to the same size (no-op)
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ $num_entries, 1999, 'no-op resize');
+
+ # Shrink to minimum i.e. zero entries and grow back
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ 0, 0, 'shrink to minimum');
+ ($num_entries, $value) = test_resize($node, $prefix, $num_entries, $value,
+ $initial_entries, 2999,
+ 'grow back from minimum');
+
+ # Test resize failure (attempt to resize beyond max - should fail)
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', "SELECT resizable_shmem_resize(" . ($max_entries * 2) . ");");
+ ok($ret != 0 || $stderr =~ /ERROR/, "$prefix: Resize beyond maximum fails");
+
+ # Resize to a size below minimum_size must fail.
+ ($ret, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT resizable_shmem_resize(-1);');
+ ok($ret != 0, "$prefix: resize below minimum_size fails");
+ like($stderr,
+ qr/cannot shrink shared memory structure "resizable_shmem" below minimum size/,
+ "$prefix: resize-below-minimum error comes from ShmemResizeStruct");
+
+ # The fault test relies on a hole being present between the current end
+ # of the structure and its maximal end. Skip when the structure does not
+ # span multiple pages.
+ my $spans_pages = $node->safe_psql('postgres', qq{
+ SELECT (maximum_size - minimum_size) >= test_shmem_pagesize()
+ FROM pg_shmem_allocations WHERE name = 'resizable_shmem';
+ });
+ if ($spans_pages ne 't')
+ {
+ diag "$prefix: skipping fault-beyond-size test: resizable_shmem does not span multiple shmem pages";
+ }
+ else
+ {
+ test_fault_beyond_size($node, $initial_entries, $prefix);
+ }
+ }
+ else
+ {
+ # On unsupported platforms, resizing should fail with a clear error
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', "SELECT resizable_shmem_resize($num_entries);");
+ ok($ret != 0, "$prefix: resize fails on unsupported platform");
+ like($stderr, qr/not supported/, "$prefix: resize error mentions not supported");
+ }
+}
+
+# Check the runtime-computed shared_memory_{initial,minimum,maximum}_size GUC
+# invariants. min <= initial <= max must always hold. When a resizable
+# structure has been registered on a server that supports resizable shared
+# memory structures, min must additionally be strictly less than max;
+# otherwise all three GUCs must be equal.
+sub check_shmem_size_gucs
+{
+ my ($node, $label) = @_;
+ my $pgdata = $node->data_dir;
+
+ my $get = sub {
+ my ($guc) = @_;
+ my ($stdout, $stderr) = run_command([ 'postgres', '-D' => $pgdata, '-C' => $guc ]);
+
+ return $stdout;
+ };
+
+ my $have_resizable_shmem = $get->('have_resizable_shmem');
+ my $ini = 0 + $get->('shared_memory_initial_size');
+ my $min = 0 + $get->('shared_memory_minimum_size');
+ my $max = 0 + $get->('shared_memory_maximum_size');
+ my $have_resizable_struct = ($get->('resizable_shmem.max_entries') ne '');
+
+ ok($min <= $ini && $ini <= $max, "$label: shared_memory size GUCs in expected order");
+
+ if ($have_resizable_struct && $have_resizable_shmem eq 'on')
+ {
+ ok($min < $max, "$label: min < max with resizable structures");
+ }
+ else
+ {
+ ok($min == $ini && $ini == $max,
+ "$label: all shared_memory size GUCs equal when no resizable structures");
+ }
+}
+
+# Log the runtime shared_memory_{initial,minimum,maximum}_size GUCs and huge
+# pages usage information for easier debugging.
+sub diag_shmem_sizes
+{
+ my ($node, $label) = @_;
+
+ my $vals = $node->safe_psql('postgres', q{
+ SELECT format('initial=%s minimum=%s maximum=%s huge_pages_status=%s huge_page_size=%s shmem_page_size=%s',
+ current_setting('shared_memory_initial_size'),
+ current_setting('shared_memory_minimum_size'),
+ current_setting('shared_memory_maximum_size'),
+ current_setting('huge_pages_status'),
+ current_setting('huge_page_size'),
+ test_shmem_pagesize());
+ });
+ diag "$label: $vals";
+}
+
+### Set up a test node.
+#
+# Configure minimal shared memory so that the resizable_shmem structure dominates
+# and any unexpected increase is easy to detect.
+#
+# If we turn on huge pages and the machine where the test is running does not
+# have huge pages available, the test will fail midway because it will not be
+# able to allocate memory pages when expanding the resizable_shmem structure.
+# Hence we turn off huge pages for this test. The test outputs the GUCs
+# shared_memory_{initial,minimum,maximum}_size and information about huge pages.
+# By provisioning enough huge pages, and by changing huge_pages = try/on, the
+# test can be run with huge pages enabled.
+###
+my $node = PostgreSQL::Test::Cluster->new('resizable_shmem');
+$node->init;
+
+$node->append_conf('postgresql.conf', 'huge_pages = off');
+$node->append_conf('postgresql.conf', 'shared_buffers = 128kB');
+$node->append_conf('postgresql.conf', 'max_connections = 5');
+$node->append_conf('postgresql.conf', 'max_worker_processes = 0');
+$node->append_conf('postgresql.conf', 'max_wal_senders = 0');
+$node->append_conf('postgresql.conf', 'max_prepared_transactions = 0');
+$node->append_conf('postgresql.conf', 'max_locks_per_transaction = 10');
+$node->append_conf('postgresql.conf', 'max_pred_locks_per_transaction = 10');
+$node->append_conf('postgresql.conf', 'wal_buffers = 32kB');
+
+###
+# Test 1: Startup allocation via shared_preload_libraries
+###
+my $startup_initial = 25 * 1024 * 1024;
+my $startup_max = 100 * 1024 * 1024;
+
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = test_shmem');
+$node->append_conf('postgresql.conf', "resizable_shmem.initial_entries = $startup_initial");
+$node->append_conf('postgresql.conf', "resizable_shmem.max_entries = $startup_max");
+
+check_shmem_size_gucs($node, 'startup preload');
+
+$node->start;
+$node->safe_psql('postgres', 'CREATE EXTENSION test_shmem;');
+diag_shmem_sizes($node, 'startup');
+run_resizable_tests($node, $startup_initial, $startup_max, 'startup');
+
+my $have_resizable_shmem = $node->safe_psql('postgres', 'SHOW have_resizable_shmem;') eq 'on';
+
+###
+# Test 2: Late allocation (loaded after startup, not in shared_preload_libraries).
+# Use much smaller sizes since only ~100KB of shared memory is available for
+# structures allocated after startup.
+###
+my $late_initial = 5 * 1024;
+my $late_max = 12 * 1024;
+
+$node->safe_psql('postgres', qq{
+ ALTER SYSTEM RESET shared_preload_libraries;
+ ALTER SYSTEM SET resizable_shmem.initial_entries = $late_initial;
+ ALTER SYSTEM SET resizable_shmem.max_entries = $late_max;
+});
+$node->safe_psql('postgres', 'DROP EXTENSION test_shmem;');
+$node->restart;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION test_shmem;');
+diag_shmem_sizes($node, 'late');
+run_resizable_tests($node, $late_initial, $late_max, 'late');
+
+###
+# Test sysv shared memory does not support resizable shmem. Only relevant on
+# platforms that support resizable shmem (HAVE_RESIZABLE_SHMEM), since the
+# module only sets maximum_size in that case.
+###
+if ($have_resizable_shmem)
+{
+ ###
+ # Test 3: Verify that CREATE EXTENSION fails with sysv shared memory
+ # when loaded after startup (not in shared_preload_libraries).
+ ###
+ $node->safe_psql('postgres', 'DROP EXTENSION test_shmem;');
+
+ # Remove settings that would cause the library to auto-load at startup:
+ # shared_preload_libraries and module-prefixed GUCs. ALTER SYSTEM RESET
+ # only affects postgresql.auto.conf, so we must use adjust_conf to remove
+ # from postgresql.conf.
+ $node->adjust_conf('postgresql.conf', 'shared_preload_libraries', undef);
+ $node->adjust_conf('postgresql.conf', 'resizable_shmem.initial_entries', undef);
+ $node->adjust_conf('postgresql.conf', 'resizable_shmem.max_entries', undef);
+ $node->adjust_conf('postgresql.auto.conf', 'shared_preload_libraries', undef);
+ $node->adjust_conf('postgresql.auto.conf', 'resizable_shmem.initial_entries', undef);
+ $node->adjust_conf('postgresql.auto.conf', 'resizable_shmem.max_entries', undef);
+ $node->safe_psql('postgres', qq{
+ ALTER SYSTEM SET shared_memory_type = 'sysv';
+ });
+
+ $node->stop;
+
+ check_shmem_size_gucs($node, 'sysv');
+
+ $node->start;
+
+ is($node->safe_psql('postgres', 'SHOW have_resizable_shmem;'),
+ 'off',
+ 'have_resizable_shmem reports off with shared_memory_type = sysv');
+
+ my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', 'CREATE EXTENSION test_shmem;');
+ ok($ret != 0, 'CREATE EXTENSION fails with resizable shmem on sysv');
+ like($stderr, qr/resizable shared memory requires shared_memory_type = mmap/,
+ 'CREATE EXTENSION error mentions shared_memory_type = mmap requirement');
+
+ ###
+ # Test 4: Verify that resizable structures are also rejected with sysv
+ # shared memory when loaded at startup via shared_preload_libraries.
+ ###
+ $node->safe_psql('postgres', qq{
+ ALTER SYSTEM SET shared_preload_libraries = 'test_shmem';
+ ALTER SYSTEM SET resizable_shmem.initial_entries = $startup_initial;
+ ALTER SYSTEM SET resizable_shmem.max_entries = $startup_max;
+ });
+ $node->stop;
+
+ ok(!$node->start(fail_ok => 1),
+ 'server fails to start with resizable shmem on sysv');
+
+ my $log = slurp_file($node->logfile);
+ like($log, qr/resizable shared memory requires shared_memory_type = mmap/,
+ 'log mentions shared_memory_type = mmap requirement');
+}
+
+done_testing();
+
+#TODO: Add a test to test the behavior of resizable shared memory when memory
+#allocation fails by simulating a memory allocation failure through injection
+#points. Add an injection point in ShmemResizeStruct in expansion branch or in
+#the underlying platform specific memory allocation function.
diff --git a/src/test/modules/test_shmem/test_shmem--1.0.sql b/src/test/modules/test_shmem/test_shmem--1.0.sql
index 2d01fd9256c..eb695604b1d 100644
--- a/src/test/modules/test_shmem/test_shmem--1.0.sql
+++ b/src/test/modules/test_shmem/test_shmem--1.0.sql
@@ -4,6 +4,61 @@
\echo Use "CREATE EXTENSION test_shmem" to load this file. \quit
+-- ===================================================================
+-- Fixed-size shared memory structure
+-- ===================================================================
+
CREATE FUNCTION get_test_shmem_attach_count()
RETURNS pg_catalog.int4 STRICT
AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION test_shmem_resize_fixed(pg_catalog.int4)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+-- ===================================================================
+-- Resizable shared memory structure
+-- ===================================================================
+
+-- Function to resize the test structure in the shared memory
+CREATE FUNCTION resizable_shmem_resize(new_entries integer)
+RETURNS bool
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to write data to all entries in the test structure in shared memory
+-- Writing all the entries makes sure that the memory is actually allocated and
+-- mapped to the process, so that we can later measure the memory usage.
+CREATE FUNCTION resizable_shmem_write(entry_value integer)
+RETURNS void
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to verify that specified number of initial entries have expected value.
+-- Reading all the entries makes sure that the memory is actually mapped to the
+-- process, so that we can later measure the memory usage.
+CREATE FUNCTION resizable_shmem_read(entry_count integer, entry_value integer)
+RETURNS boolean
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to report memory mapped against the main shared memory segment in
+-- the backend where this function runs.
+CREATE FUNCTION test_shmem_usage()
+RETURNS bigint
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to get the shared memory page size
+CREATE FUNCTION test_shmem_pagesize()
+RETURNS integer
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
+
+-- Function to crash the backend by walking entries past the current size up to
+-- the reserved maximum, reading or writing each one as decided by mode.
+CREATE FUNCTION resizable_shmem_access_beyond_size(mode text)
+RETURNS integer
+AS 'MODULE_PATHNAME'
+LANGUAGE C STRICT;
diff --git a/src/test/modules/test_shmem/test_shmem.c b/src/test/modules/test_shmem/test_shmem.c
index 9bd4012b435..83f8ea9cc3c 100644
--- a/src/test/modules/test_shmem/test_shmem.c
+++ b/src/test/modules/test_shmem/test_shmem.c
@@ -1,11 +1,10 @@
/*-------------------------------------------------------------------------
*
* test_shmem.c
- * Helpers to test shmem allocation routines
+ * Helpers to test shmem management routines
*
- * Test basic memory allocation in an extension module. One notable feature
- * that is not exercised by any other module in the repository is the
- * allocating (non-DSM) shared memory after postmaster startup.
+ * Test fixed-size and resizable shared memory structures created during
+ * postmaster startup and after startup respectively.
*
* Copyright (c) 2020-2026, PostgreSQL Global Development Group
*
@@ -17,13 +16,26 @@
#include "postgres.h"
+#include <limits.h>
+
+#include "commands/extension.h"
#include "fmgr.h"
#include "miscadmin.h"
+#include "storage/fd.h"
+#include "storage/pg_shmem.h"
#include "storage/shmem.h"
+#include "utils/builtins.h"
+#include "utils/guc.h"
PG_MODULE_MAGIC;
+
+/* ----------------------------------------------------------------
+ * Fixed-size shared memory structure
+ * ----------------------------------------------------------------
+ */
+
typedef struct TestShmemData
{
int value;
@@ -35,17 +47,6 @@ static TestShmemData *TestShmem;
static bool attached_or_initialized = false;
-static void test_shmem_request(void *arg);
-static void test_shmem_init(void *arg);
-static void test_shmem_attach(void *arg);
-
-static const ShmemCallbacks TestShmemCallbacks = {
- .flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP,
- .request_fn = test_shmem_request,
- .init_fn = test_shmem_init,
- .attach_fn = test_shmem_attach,
-};
-
static void
test_shmem_request(void *arg)
{
@@ -60,6 +61,17 @@ static void
test_shmem_init(void *arg)
{
elog(LOG, "init callback called");
+
+ /*
+ * Reset the per-process flag and the shared "initialized" marker during
+ * postmaster induced restart.
+ */
+ if (!IsUnderPostmaster)
+ {
+ attached_or_initialized = false;
+ TestShmem->initialized = false;
+ }
+
if (TestShmem->initialized)
elog(ERROR, "shmem area already initialized");
TestShmem->initialized = true;
@@ -82,12 +94,12 @@ test_shmem_attach(void *arg)
attached_or_initialized = true;
}
-void
-_PG_init(void)
-{
- elog(LOG, "test_shmem module's _PG_init called");
- RegisterShmemCallbacks(&TestShmemCallbacks);
-}
+static const ShmemCallbacks TestShmemCallbacks = {
+ .flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP,
+ .request_fn = test_shmem_request,
+ .init_fn = test_shmem_init,
+ .attach_fn = test_shmem_attach,
+};
PG_FUNCTION_INFO_V1(get_test_shmem_attach_count);
Datum
@@ -99,3 +111,453 @@ get_test_shmem_attach_count(PG_FUNCTION_ARGS)
elog(ERROR, "shmem area not yet initialized");
PG_RETURN_INT32(TestShmem->attach_count);
}
+
+/*
+ * Attempt to resize the fixed-size shared memory structure. This should
+ * fail because the structure was not allocated with a maximum_size.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_resize_fixed);
+Datum
+test_shmem_resize_fixed(PG_FUNCTION_ARGS)
+{
+ int32 new_size = PG_GETARG_INT32(0);
+
+ ShmemResizeStruct("test_shmem area", new_size);
+ PG_RETURN_VOID();
+}
+
+
+/* ----------------------------------------------------------------
+ * Resizable shared memory structure
+ * ----------------------------------------------------------------
+ */
+
+/*
+ * The test module may be loaded after postmaster startup in which case only
+ * 100K of shared memory is available for the extension. Keep the default
+ * initial and maximum sizes small enough to fit in that space.
+ */
+#define TEST_INITIAL_ENTRIES_DEFAULT 1
+#define TEST_MAX_ENTRIES_DEFAULT 1024
+
+#define TEST_ENTRY_SIZE sizeof(int32) /* Size of each entry */
+
+/*
+ * Resizable test data structure stored in shared memory.
+ *
+ * The test performs resizing, reads or writes, only one at a time and never
+ * concurrently. Hence, there is no need for locks in the test structure.
+ */
+typedef struct TestResizableShmemStruct
+{
+ /* Metadata */
+ int32 num_entries; /* Number of entries that can fit */
+
+ /* Data area - variable size */
+ int32 data[FLEXIBLE_ARRAY_MEMBER];
+} TestResizableShmemStruct;
+
+static TestResizableShmemStruct *resizable_shmem = NULL;
+
+/* GUC variables controlling the size of the test structure */
+static int test_initial_entries;
+static int test_max_entries;
+
+/* Whether to use SHMEM_ATTACH_UNKNOWN_SIZE when attaching to the shared memory */
+/* TODO: We may use opaque_arg to pass this value to the request function.*/
+static bool use_unknown_size = false;
+
+/*
+ * Request shared memory resources.
+ */
+static void
+resizable_shmem_request(void *arg)
+{
+ Size initial_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(test_initial_entries, TEST_ENTRY_SIZE));
+
+/*
+ * Create resizable structure on the platforms which support it. Otherwise create
+ * as a fixed-size structure. Other way would be to conditionally include
+ * .maximum_size in the call to ShmemRequestStruct().
+ */
+#ifdef HAVE_RESIZABLE_SHMEM
+ Size max_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(test_max_entries, TEST_ENTRY_SIZE));
+ Size min_size = offsetof(TestResizableShmemStruct, data);
+#else
+ Size max_size = 0;
+ Size min_size = 0;
+#endif
+
+ ShmemRequestStruct(.name = "resizable_shmem",
+ .size = use_unknown_size ? SHMEM_ATTACH_UNKNOWN_SIZE : initial_size,
+ .minimum_size = min_size,
+ .maximum_size = max_size,
+ .ptr = (void **) &resizable_shmem,
+ );
+}
+
+/*
+ * Initialize shared memory structure.
+ */
+static void
+resizable_shmem_shmem_init(void *arg)
+{
+ Assert(resizable_shmem != NULL);
+
+ resizable_shmem->num_entries = test_initial_entries;
+ memset(resizable_shmem->data, 0, mul_size(test_initial_entries, TEST_ENTRY_SIZE));
+}
+
+/*
+ * Attach to the already-allocated shared memory structure.
+ */
+static void
+resizable_shmem_shmem_attach(void *arg)
+{
+ Assert(resizable_shmem != NULL);
+}
+
+static ShmemCallbacks resizable_shmem_callbacks = {
+ .request_fn = resizable_shmem_request,
+ .init_fn = resizable_shmem_shmem_init,
+ .attach_fn = resizable_shmem_shmem_attach,
+};
+
+/*
+ * Resize the shared memory structure to accommodate the specified number of
+ * entries.
+ *
+ * Negative value for new_entries can be used to test resizing below the
+ * minimum size.
+ *
+ * Returns true if the resize was successful, false if ShmemResizeStruct()
+ * could not allocate the requested memory. On platforms that do not support
+ * resizable shared memory, ShmemResizeStruct() raises an error.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_resize);
+Datum
+resizable_shmem_resize(PG_FUNCTION_ARGS)
+{
+ int32 new_entries = PG_GETARG_INT32(0);
+ Size new_size;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (new_entries < 0)
+ new_size = 1;
+ else
+ new_size = add_size(offsetof(TestResizableShmemStruct, data),
+ mul_size(new_entries, TEST_ENTRY_SIZE));
+ if (!ShmemResizeStruct("resizable_shmem", new_size))
+ PG_RETURN_BOOL(false);
+
+ ShmemProtectStruct("resizable_shmem");
+ resizable_shmem->num_entries = new_entries;
+
+ PG_RETURN_BOOL(true);
+}
+
+/*
+ * Write the given integer value to all entries in the data array.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_write);
+Datum
+resizable_shmem_write(PG_FUNCTION_ARGS)
+{
+ int32 entry_value = PG_GETARG_INT32(0);
+ int32 i;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (i = 0; i < resizable_shmem->num_entries; i++)
+ resizable_shmem->data[i] = entry_value;
+
+ PG_RETURN_VOID();
+}
+
+/*
+ * Check whether the first 'entry_count' entries all have the expected 'entry_value'.
+ * Returns true if all match, false otherwise.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_read);
+Datum
+resizable_shmem_read(PG_FUNCTION_ARGS)
+{
+ int32 entry_count = PG_GETARG_INT32(0);
+ int32 entry_value = PG_GETARG_INT32(1);
+ int32 i;
+
+ if (resizable_shmem == NULL)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (entry_count < 0 || entry_count > resizable_shmem->num_entries)
+ ereport(ERROR,
+ errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("entry_count %d is out of range (0..%d)", entry_count, resizable_shmem->num_entries));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (i = 0; i < entry_count; i++)
+ {
+ if (resizable_shmem->data[i] != entry_value)
+ PG_RETURN_BOOL(false);
+ }
+
+ PG_RETURN_BOOL(true);
+}
+
+/*
+ * Return the memory mapped against the main shared memory segment in this
+ * backend.
+ *
+ * The VMA containing our resizable_shmem pointer identifies the start of the
+ * main shared-memory segment.
+ *
+ * mprotect() calls issued when the resizable structure grows and shrinks can
+ * split the original mmap into several adjacent VMAs, so we sum the accounting
+ * fields across the base VMA and any VMAs contiguous with it.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_usage);
+Datum
+test_shmem_usage(PG_FUNCTION_ARGS)
+{
+ FILE *f;
+ char line[256];
+ uintptr_t target = (uintptr_t) resizable_shmem;
+ bool in_target_vma = false;
+ bool use_hugetlb = (huge_pages_status == HUGE_PAGES_ON);
+ unsigned long prev_end = 0;
+ int64 total_rss_kb = 0;
+ int64 total_swap_kb = 0;
+ int64 total_shared_hugetlb_kb = 0;
+ int64 val;
+ size_t result;
+
+ f = AllocateFile("/proc/self/smaps", "r");
+ if (f == NULL)
+ ereport(ERROR,
+ errcode_for_file_access(),
+ errmsg("could not open /proc/self/smaps: %m"));
+
+ while (fgets(line, sizeof(line), f) != NULL)
+ {
+ unsigned long start;
+ unsigned long end;
+
+ if (sscanf(line, "%lx-%lx", &start, &end) == 2)
+ {
+ if (in_target_vma)
+ {
+ /*
+ * Continue accumulating only across VMAs that are contiguous
+ * with the previous one; stop as soon as we hit a gap or a
+ * different mapping.
+ */
+ if (start != prev_end)
+ break;
+ }
+ else
+ in_target_vma = (target >= start && target < end);
+
+ prev_end = end;
+ }
+ else if (in_target_vma)
+ {
+ if (use_hugetlb)
+ {
+ if (sscanf(line, "Shared_Hugetlb: %ld kB", &val) == 1)
+ total_shared_hugetlb_kb += val;
+ }
+ else
+ {
+ if (sscanf(line, "Rss: %ld kB", &val) == 1)
+ total_rss_kb += val;
+ else if (sscanf(line, "Swap: %ld kB", &val) == 1)
+ total_swap_kb += val;
+ }
+ }
+ }
+
+ FreeFile(f);
+
+ if (use_hugetlb)
+ result = mul_size(total_shared_hugetlb_kb, 1024);
+ else
+ {
+ result = mul_size(total_rss_kb, 1024);
+ result = add_size(result, mul_size(total_swap_kb, 1024));
+ }
+
+ PG_RETURN_INT64(result);
+}
+
+/*
+ * Return the shared memory page size.
+ */
+PG_FUNCTION_INFO_V1(test_shmem_pagesize);
+Datum
+test_shmem_pagesize(PG_FUNCTION_ARGS)
+{
+ PG_RETURN_INT32(pg_get_shmem_pagesize());
+}
+
+/*
+ * Walk the entries between the current size and the reserved maximum, accessing
+ * each one. Ideally, this function should (seg)fault the moment we try to access
+ * the entry outside the currently allocated size, but the memory allocation and
+ * protection mechanisms work on page basis. Hence it may only (seg)fault when a
+ * page boundary is crossed. The mode argument selects between "read" and
+ * "write" access.
+ *
+ * When the current end of the structure and end of maximal structure are on the
+ * same page, this function may not (seg)fault at all.
+ */
+PG_FUNCTION_INFO_V1(resizable_shmem_access_beyond_size);
+Datum
+resizable_shmem_access_beyond_size(PG_FUNCTION_ARGS)
+{
+ text *mode_txt = PG_GETARG_TEXT_PP(0);
+ const char *mode = text_to_cstring(mode_txt);
+ bool do_write;
+ int32 sink = 0;
+
+ if (!resizable_shmem)
+ ereport(ERROR,
+ errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("resizable_shmem is not initialized"));
+
+ if (strcmp(mode, "read") == 0)
+ do_write = false;
+ else if (strcmp(mode, "write") == 0)
+ do_write = true;
+ else
+ ereport(ERROR,
+ errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("mode must be \"read\" or \"write\""));
+
+#ifdef HAVE_RESIZABLE_SHMEM
+
+ /*
+ * Ideally the structure should be protected through a synchronization
+ * cycle across all the backends that may access the structure. But we
+ * don't implement any such synchronization in this test module to keep it
+ * simple. Given that ProcSignalBarrier mechanism is not extensible, we
+ * may not be able to do that as well here. Hence add protect just before
+ * accessing the structure.
+ */
+ ShmemProtectStruct("resizable_shmem");
+#endif
+
+ for (int i = resizable_shmem->num_entries; i < test_max_entries; i++)
+ {
+ if (do_write)
+ resizable_shmem->data[i] = 0xdead;
+ else
+ sink = resizable_shmem->data[i];
+ }
+
+ /*
+ * Return the last read value so that compiler doesn't optimize away the
+ * assignment to sink.
+ */
+ PG_RETURN_INT32(sink);
+}
+
+
+/* ----------------------------------------------------------------
+ * Module initialization
+ * ----------------------------------------------------------------
+ */
+
+void
+_PG_init(void)
+{
+ int guc_context;
+
+ elog(LOG, "test_shmem module's _PG_init called");
+
+ RegisterShmemCallbacks(&TestShmemCallbacks);
+
+ /*
+ * Use PGC_POSTMASTER when loaded at startup so the values are fixed once
+ * the shared memory segment is created. When loaded after startup
+ * PGC_POSTMASTER is not allowed, so we use PGC_SIGHUP instead. Although
+ * we do not intend to change these values at config reload, PGC_SIGHUP is
+ * the least permissive context that allows defining the GUC after startup
+ * and still prevents it from being changed via SET.
+ */
+ if (process_shared_preload_libraries_in_progress)
+ guc_context = PGC_POSTMASTER;
+ else
+ {
+ guc_context = PGC_SIGHUP;
+ resizable_shmem_callbacks.flags = SHMEM_CALLBACKS_ALLOW_AFTER_STARTUP;
+ }
+
+ DefineCustomIntVariable("resizable_shmem.initial_entries",
+ "Initial number of entries in the test structure.",
+ NULL,
+ &test_initial_entries,
+ TEST_INITIAL_ENTRIES_DEFAULT,
+ 1,
+ INT_MAX,
+ guc_context,
+ 0,
+ NULL, NULL, NULL);
+
+ DefineCustomIntVariable("resizable_shmem.max_entries",
+ "Maximum number of entries in the test structure.",
+ NULL,
+ &test_max_entries,
+ TEST_MAX_ENTRIES_DEFAULT,
+ 1,
+ INT_MAX,
+ guc_context,
+ 0,
+ NULL, NULL, NULL);
+
+ /*
+ * When loaded after startup by a backend that is not creating the
+ * extension, the shared memory might have been resized to a size other
+ * than the initial size. Use SHMEM_ATTACH_UNKNOWN_SIZE to attach without
+ * knowing the exact size.
+ */
+ if (!process_shared_preload_libraries_in_progress && !creating_extension)
+ use_unknown_size = true;
+
+ RegisterShmemCallbacks(&resizable_shmem_callbacks);
+}
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 6a3341356da..f2c47ac1c28 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1770,8 +1770,11 @@ pg_shadow| SELECT pg_authid.rolname AS usename,
pg_shmem_allocations| SELECT name,
off,
size,
- allocated_size
- FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size);
+ allocated_size,
+ minimum_size,
+ maximum_size,
+ reserved_space
+ FROM pg_get_shmem_allocations() pg_get_shmem_allocations(name, off, size, allocated_size, minimum_size, maximum_size, reserved_space);
pg_shmem_allocations_numa| SELECT name,
numa_node,
size
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 85d989f395d..a823271262c 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -3186,6 +3186,7 @@ TestDSMRegistryHashEntry
TestDSMRegistryStruct
TestDecodingData
TestDecodingTxnData
+TestResizableShmemStruct
TestShmemData
TestSpec
TestValueType
[application/octet-stream] v20260917-0008-Follow-up-changes-since-last-email-on-hackers.patch (69.3K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/9-v20260917-0008-Follow-up-changes-since-last-email-on-hackers.patch)
download | inline diff:
From 7fd0f0019ee668fac4b6ec3af7da53286081596f Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Mon, 27 Jul 2026 11:14:32 +0530
Subject: [PATCH] Follow-up changes since last email on hackers
Please note that the stress tests added by this or the earlier commits
are not necessarily meant to be committed to the core. But they are in
the patch so that reviewers have some readily available stress tests.
Move the SQL wrapper around pg_resize_shared_buffers() to an extension.
Address stress test utility common for all buffer pool resizing stress
tests.
Use randomised sequence of sizes when resizing.
Test synchronization between buffer pool resizing and
DropRelationBuffers, DropRelationsAllBuffers, DropDatabaseBuffers,
CHECKPOINT, FlushRelationBuffers, FlushRelationsAllBuffers, monitoring
and diagnostic functions in pg_buffercache and pg_prewarm.
FlushDatabaseBuffers() is not covered by any stress test: it is only
called during WAL replay of xl_dbase_create_file_copy_rec (in
dbcommands.c). The primary CREATE DATABASE and ALTER DATABASE SET
TABLESPACE paths use RequestCheckpoint() instead, so the function is
not reachable from a live SQL workload.
A bug fix in BufferSync because of which checkpoint didn't advance post
a buffer pool shrink and blocked any other activity that involved
ProcSignalBarrier.
Use separate number ranges for stress tests and other tests.
Add database checker at the end of all stress tests.
Enable running tests only when PG_TEST_EXTRA has bufmgr_stress. Also
skip all tests if have_resizable_shmem is OFF.
Change pg_resize_shared_buffers() to throw an error when the server
does not support resizable shmem.
Fix when CFI is called in some pg_buffercache functions that scan the
buffer pool.
Skip pg_resize_shared_buffers tests when resizable shmem is not
supported. Add a test to make sure that pg_resize_shared_buffers()
causes an error when invoked in a server which does not support
resizable shmem.
---
contrib/pg_buffercache/pg_buffercache_pages.c | 24 +-
doc/src/sgml/regress.sgml | 11 +
src/backend/catalog/storage.c | 4 +
src/backend/storage/buffer/buf_resize.c | 16 +
src/backend/storage/buffer/bufmgr.c | 63 +-
src/test/buffermgr/Makefile | 7 +-
src/test/buffermgr/README | 28 +-
src/test/buffermgr/buffermgr_test--1.0.sql | 52 ++
src/test/buffermgr/buffermgr_test.control | 3 +
src/test/buffermgr/meson.build | 22 +-
src/test/buffermgr/t/001_resize_buffer.pl | 193 ------
...rance.pl => 001_resize_fault_tolerance.pl} | 11 +-
...ze.pl => 002_client_join_buffer_resize.pl} | 15 +-
...ize_failures.pl => 003_resize_failures.pl} | 7 +
...logger.pl => 004_resize_with_syslogger.pl} | 7 +
.../buffermgr/t/005_resize_unsupported.pl | 51 ++
.../buffermgr/t/010_stress_resize_buffer.pl | 34 +
.../t/011_stress_drop_relation_buffers.pl | 99 +++
.../t/012_stress_drop_database_buffers.pl | 82 +++
src/test/buffermgr/t/013_stress_checkpoint.pl | 47 ++
.../t/014_stress_flush_relation_buffers.pl | 68 ++
.../buffermgr/t/015_stress_pg_buffercache.pl | 60 ++
src/test/buffermgr/t/016_stress_pg_prewarm.pl | 73 ++
src/test/buffermgr/t/StressUtil.pm | 628 ++++++++++++++++++
24 files changed, 1361 insertions(+), 244 deletions(-)
create mode 100644 src/test/buffermgr/buffermgr_test--1.0.sql
create mode 100644 src/test/buffermgr/buffermgr_test.control
delete mode 100644 src/test/buffermgr/t/001_resize_buffer.pl
rename src/test/buffermgr/t/{003_resize_fault_tolerance.pl => 001_resize_fault_tolerance.pl} (99%)
rename src/test/buffermgr/t/{004_client_join_buffer_resize.pl => 002_client_join_buffer_resize.pl} (96%)
rename src/test/buffermgr/t/{005_resize_failures.pl => 003_resize_failures.pl} (96%)
rename src/test/buffermgr/t/{006_resize_with_syslogger.pl => 004_resize_with_syslogger.pl} (87%)
create mode 100644 src/test/buffermgr/t/005_resize_unsupported.pl
create mode 100644 src/test/buffermgr/t/010_stress_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/011_stress_drop_relation_buffers.pl
create mode 100644 src/test/buffermgr/t/012_stress_drop_database_buffers.pl
create mode 100644 src/test/buffermgr/t/013_stress_checkpoint.pl
create mode 100644 src/test/buffermgr/t/014_stress_flush_relation_buffers.pl
create mode 100644 src/test/buffermgr/t/015_stress_pg_buffercache.pl
create mode 100644 src/test/buffermgr/t/016_stress_pg_prewarm.pl
create mode 100644 src/test/buffermgr/t/StressUtil.pm
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 312343fd7bf..7335ccae150 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -150,8 +150,6 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
Datum values[NUM_BUFFERCACHE_PAGES_ELEM];
bool nulls[NUM_BUFFERCACHE_PAGES_ELEM];
- CHECK_FOR_INTERRUPTS();
-
bufHdr = GetBufferDescriptor(i);
/* Lock each buffer header before inspecting. */
buf_state = LockBufHdr(bufHdr);
@@ -220,6 +218,12 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
}
tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+
+ /*
+ * Check for interrupts here, at the end of the loop, so that the buffer
+ * index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
return (Datum) 0;
@@ -457,8 +461,6 @@ pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
char *startptr_buff,
*endptr_buff;
- CHECK_FOR_INTERRUPTS();
-
bufHdr = GetBufferDescriptor(i);
/* Lock each buffer header before inspecting. */
@@ -488,6 +490,12 @@ pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
++idx;
++page_num;
}
+
+ /*
+ * Check for interrupts here, at the end of the loop, so that the
+ * buffer index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
Assert(idx <= max_entries);
@@ -599,8 +607,6 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
BufferDesc *bufHdr;
uint64 buf_state;
- CHECK_FOR_INTERRUPTS();
-
/*
* This function summarizes the state of all headers. Locking the
* buffer headers wouldn't provide an improved result as the state of
@@ -623,6 +629,12 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
buffers_pinned++;
+
+ /*
+ * Check for interrupts here, at the end of the loop, so that the buffer
+ * index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
memset(nulls, 0, sizeof(nulls));
diff --git a/doc/src/sgml/regress.sgml b/doc/src/sgml/regress.sgml
index c74941bfbf2..30092fd820d 100644
--- a/doc/src/sgml/regress.sgml
+++ b/doc/src/sgml/regress.sgml
@@ -275,6 +275,17 @@ make check-world PG_TEST_EXTRA='kerberos ldap ssl load_balance libpq_encryption'
</programlisting>
The following values are currently supported:
<variablelist>
+ <varlistentry>
+ <term><literal>bufmgr_stress</literal></term>
+ <listitem>
+ <para>
+ Runs the shared_buffers resize stress tests under
+ <filename>src/test/buffermgr</filename>. Not enabled by default because
+ they are long-running and resource-intensive.
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry>
<term><literal>checksum</literal>, <literal>checksum_extended</literal></term>
<listitem>
diff --git a/src/backend/catalog/storage.c b/src/backend/catalog/storage.c
index e443a4993c5..5c36820d58f 100644
--- a/src/backend/catalog/storage.c
+++ b/src/backend/catalog/storage.c
@@ -33,6 +33,7 @@
#include "storage/proc.h"
#include "storage/smgr.h"
#include "utils/hsearch.h"
+#include "utils/injection_point.h"
#include "utils/memutils.h"
#include "utils/rel.h"
@@ -383,6 +384,9 @@ RelationTruncate(Relation rel, BlockNumber nblocks)
*
* (See also visibilitymap.c if changing this code.)
*/
+
+ /* Load the injection point before entering the critical section */
+ INJECTION_POINT_LOAD("drop-relation-buffers-scan");
START_CRIT_SECTION();
if (RelationNeedsWAL(rel))
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
index c758493f353..90f5fb1c71d 100644
--- a/src/backend/storage/buffer/buf_resize.c
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -36,6 +36,7 @@
static volatile sig_atomic_t safe_exit = true;
+#ifdef HAVE_RESIZABLE_SHMEM
static bool resize_shared_buffers_internal(void);
static void buf_resize_shmem_exit(int code, Datum arg);
@@ -97,6 +98,7 @@ buf_resize_shmem_resize(int currentNBuffers, int targetNBuffers)
elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE barrier");
return true;
}
+#endif
/*
* C implementation of SQL interface to update the shared buffers according to
@@ -188,8 +190,19 @@ buf_resize_shmem_resize(int currentNBuffers, int targetNBuffers)
Datum
pg_resize_shared_buffers(PG_FUNCTION_ARGS)
{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizing shared buffer pool is not supported on this platform"));
+ pg_unreachable();
+#else
bool success = false;
+ if (shared_memory_type != SHMEM_TYPE_MMAP)
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizing shared buffer pool is not supported on this platform"));
+
/*
* Register the exit hook before claiming resizer_pid, so that if we exit
* after claiming resizer_pid, the hook is in place to reset it.
@@ -269,8 +282,10 @@ pg_resize_shared_buffers(PG_FUNCTION_ARGS)
elog(WARNING, "shared buffer resizing to %d buffers failed", NBuffersGUC);
PG_RETURN_BOOL(success);
+#endif
}
+#ifdef HAVE_RESIZABLE_SHMEM
/*
* Workhorse function for the C implementation.
*/
@@ -404,6 +419,7 @@ buf_resize_shmem_exit(int code, Datum arg)
(void) pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
&expected_pid, 0);
}
+#endif
/*
* Process and acknowledge PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC.
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index cebc85624b8..bb436734585 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -64,6 +64,7 @@
#include "storage/read_stream.h"
#include "storage/smgr.h"
#include "storage/standby.h"
+#include "utils/injection_point.h"
#include "utils/memdebug.h"
#include "utils/ps_status.h"
#include "utils/rel.h"
@@ -3793,36 +3794,35 @@ BufferSync(int flags)
/*
* The buffer pool might have been shrunk between the time the
- * checkpoint collected the buffer ids and now. Ignore any buffers
- * that are out of range now. Those buffers must have been written
- * when they were evicted during resizing.
+ * checkpoint collected the buffer ids and now. Skip any buffers that
+ * are out of range now; they were written when they were evicted
+ * during resizing.
*/
- if (buf_id >= NBuffers)
- continue;
-
- bufHdr = GetBufferDescriptor(buf_id);
-
- num_processed++;
-
- /*
- * We don't need to acquire the lock here, because we're only looking
- * at a single bit. It's possible that someone else writes the buffer
- * and clears the flag right after we check, but that doesn't matter
- * since SyncOneBuffer will then do nothing. However, there is a
- * further race condition: it's conceivable that between the time we
- * examine the bit here and the time SyncOneBuffer acquires the lock,
- * someone else not only wrote the buffer but replaced it with another
- * page and dirtied it. In that improbable case, SyncOneBuffer will
- * write the buffer though we didn't need to. It doesn't seem worth
- * guarding against this, though.
- */
- if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+ if (buf_id < NBuffers)
{
- if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
+ bufHdr = GetBufferDescriptor(buf_id);
+
+ /*
+ * We don't need to acquire the lock here, because we're only
+ * looking at a single bit. It's possible that someone else writes
+ * the buffer and clears the flag right after we check, but that
+ * doesn't matter since SyncOneBuffer will then do nothing.
+ * However, there is a further race condition: it's conceivable
+ * that between the time we examine the bit here and the time
+ * SyncOneBuffer acquires the lock, someone else not only wrote
+ * the buffer but replaced it with another page and dirtied it.
+ * In that improbable case, SyncOneBuffer will write the buffer
+ * though we didn't need to. It doesn't seem worth guarding
+ * against this, though.
+ */
+ if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
{
- TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id);
- PendingCheckpointerStats.buffers_written++;
- num_written++;
+ if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
+ {
+ TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id);
+ PendingCheckpointerStats.buffers_written++;
+ num_written++;
+ }
}
}
@@ -3833,6 +3833,7 @@ BufferSync(int flags)
ts_stat->progress += ts_stat->progress_slice;
ts_stat->num_scanned++;
ts_stat->index++;
+ num_processed++;
/* Have all the buffers from the tablespace been processed? */
if (ts_stat->num_scanned == ts_stat->num_to_scan)
@@ -4976,6 +4977,8 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
if (j >= nforks)
UnlockBufHdr(bufHdr);
}
+
+ INJECTION_POINT_CACHED("drop-relation-buffers-scan", NULL);
}
/* ---------------------------------------------------------------------
@@ -5143,6 +5146,8 @@ DropRelationsAllBuffers(SMgrRelation *smgr_reln, int nlocators)
UnlockBufHdr(bufHdr);
}
+ INJECTION_POINT("drop-relations-all-buffers-scan", NULL);
+
pfree(locators);
pfree(rels);
}
@@ -5243,6 +5248,8 @@ DropDatabaseBuffers(Oid dbid)
else
UnlockBufHdr(bufHdr);
}
+
+ INJECTION_POINT("drop-database-buffers-scan", NULL);
}
/* ---------------------------------------------------------------------
@@ -5340,6 +5347,7 @@ FlushRelationBuffers(Relation rel)
else
UnlockBufHdr(bufHdr);
}
+ INJECTION_POINT("flush-relation-buffers-scan", NULL);
}
/* ---------------------------------------------------------------------
@@ -5435,6 +5443,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
else
UnlockBufHdr(bufHdr);
}
+ INJECTION_POINT("flush-relations-all-buffers-scan", NULL);
pfree(srels);
}
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
index 24c245c900a..92a430e736b 100644
--- a/src/test/buffermgr/Makefile
+++ b/src/test/buffermgr/Makefile
@@ -9,10 +9,15 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = contrib/pg_buffercache \
+EXTRA_INSTALL = contrib/amcheck \
+ contrib/pg_buffercache \
+ contrib/pg_prewarm \
src/test/modules/injection_points \
src/test/modules/test_shmem
+EXTENSION = buffermgr_test
+DATA = buffermgr_test--1.0.sql
+
REGRESS = buffer_resize
# Custom configuration for buffer manager tests
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
index c375ad80989..55bcf0a800c 100644
--- a/src/test/buffermgr/README
+++ b/src/test/buffermgr/README
@@ -3,8 +3,34 @@ src/test/buffermgr/README
Regression tests for buffer manager
===================================
-This directory contains a test suite for resizing buffer manager without restarting the server.
+This directory contains a test suite for resizing buffer manager without
+restarting the server.
+Some of the TAP tests rely on a helper extension buffermgr_test. It bundles SQL
+helpers (such as pg_resize_shared_buffers_sql) used by the tests.
+
+Stress tests
+------------------
+These tests exercise the synchronization between the buffer manager resizing and the code that uses buffer manager. They run the SQL commands that exercise the specific subsystem functionality in parallel with the resizing of buffer manager, both in a tight loop so as to increase the chances of hitting race conditions. Optionally they run pgbench to keep the buffer pool busy.
+
+1. simple pgbench: 010_stress_resize_buffer.pl
+
+2. checkpoint stress test: TODO: Palak Chaturvedi. Disable automatic checkpointing. Run pgbench. Run checkpoint and resize buffer pool in a tight loop.
+
+3. pg_buffercache stress test: TODO: Run pgbench. Run pg_buffercache queries and resize buffer pool in a tight loop.
+
+4. pg_prewarm stress test: TODO: Need to figure out exactly what to test.
+
+5. Relation/Database buffer scan stress test: TODO:
+5.a: Create a table with a lot of data (spanning more than NBuffers/32 pages) and truncate it within the same transaction, insert the data again and drop the table. This should exercise both DropRelationBuffers() and DropRelationsAllBuffers(). Run this in a tight loop along with resizing buffer pool. Run pgbench so that the table being created and dropped has its pages interspersed with other pages in the buffer pool.
+5.b: Similarly Create and drop a database in a tight loop along with resizing buffer pool. Run pgbench so that the database being created and dropped has its pages interspersed with other pages in the buffer pool. You may want to populate some tables in the template database so that the database being created and dropped is prepopulated with some data.
+5.c: This is more nuanced. Run commands exercising FlushRelationBuffers(), FlushRelationsAllBuffers() and FlushDatabaseBuffers() in a tight loop along with resizing buffer pool. Run pgbench so that the relations/databases being flushed have their pages interspersed with other pages in the buffer pool. The exact commands to run are TBD and need experimentation.
+
+Injection point tests
+------------------
+Any issues discovered during the stress tests can be converted into injection
+point tests. Injection points allow to simulate a scenario precisely and
+deterministically.
Running the tests
=================
diff --git a/src/test/buffermgr/buffermgr_test--1.0.sql b/src/test/buffermgr/buffermgr_test--1.0.sql
new file mode 100644
index 00000000000..51901508523
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test--1.0.sql
@@ -0,0 +1,52 @@
+-- Helper functions used by TAP tests in src/test/buffermgr/t.
+--
+
+\echo Use "CREATE EXTENSION buffermgr_test" to load this file. \quit
+
+-- Retries pg_resize_shared_buffers() until it succeeds, then confirms the
+-- new value is in effect. Returns the number of retries taken along with
+-- the wall-clock times immediately before and after the retry loop.
+--
+-- The new size is expected to be set in shared_buffers GUC before calling this
+-- function.
+create function pg_resize_shared_buffers_sql(
+ new_size int,
+ out num_tries int,
+ out started_at timestamptz,
+ out ended_at timestamptz)
+returns record as $$
+declare
+ success boolean := false;
+ tries int := 0;
+ cur_setting text;
+ target text := new_size::text;
+ pending_pattern text := '%(pending: ' || target || ')%';
+begin
+ select setting into cur_setting
+ from pg_settings where name = 'shared_buffers';
+ if cur_setting <> target and cur_setting not like pending_pattern then
+ raise exception 'shared_buffers change not visible to this backend: setting is %, expected % or matching %',
+ cur_setting, target, pending_pattern;
+ end if;
+
+ started_at := clock_timestamp();
+ while not success loop
+ tries := tries + 1;
+ select pg_resize_shared_buffers() into success;
+ if not success then
+ perform pg_sleep(0.1);
+ end if;
+ end loop;
+ ended_at := clock_timestamp();
+
+ select setting into cur_setting
+ from pg_settings where name = 'shared_buffers';
+ if cur_setting <> target then
+ raise exception 'shared_buffers resize did not take effect: expected %, got %',
+ target, cur_setting;
+ end if;
+
+ num_tries := tries;
+ return;
+end;
+$$ language plpgsql;
diff --git a/src/test/buffermgr/buffermgr_test.control b/src/test/buffermgr/buffermgr_test.control
new file mode 100644
index 00000000000..2d13d889ec7
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.control
@@ -0,0 +1,3 @@
+comment = 'Helpers for src/test/buffermgr TAP tests'
+default_version = '1.0'
+relocatable = true
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
index 7a6d5e29f8d..3c64d2740ee 100644
--- a/src/test/buffermgr/meson.build
+++ b/src/test/buffermgr/meson.build
@@ -1,5 +1,10 @@
# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+test_install_data += files(
+ 'buffermgr_test.control',
+ 'buffermgr_test--1.0.sql',
+)
+
tests += {
'name': 'buffermgr',
'sd': meson.current_source_dir(),
@@ -15,11 +20,18 @@ tests += {
'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
},
'tests': [
- 't/001_resize_buffer.pl',
- 't/003_resize_fault_tolerance.pl',
- 't/004_client_join_buffer_resize.pl',
- 't/005_resize_failures.pl',
- 't/006_resize_with_syslogger.pl',
+ 't/001_resize_fault_tolerance.pl',
+ 't/002_client_join_buffer_resize.pl',
+ 't/003_resize_failures.pl',
+ 't/004_resize_with_syslogger.pl',
+ 't/005_resize_unsupported.pl',
+ 't/010_stress_resize_buffer.pl',
+ 't/011_stress_drop_relation_buffers.pl',
+ 't/012_stress_drop_database_buffers.pl',
+ 't/013_stress_checkpoint.pl',
+ 't/014_stress_flush_relation_buffers.pl',
+ 't/015_stress_pg_buffercache.pl',
+ 't/016_stress_pg_prewarm.pl',
],
},
}
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
deleted file mode 100644
index fb5a42be26a..00000000000
--- a/src/test/buffermgr/t/001_resize_buffer.pl
+++ /dev/null
@@ -1,193 +0,0 @@
-# Copyright (c) 2025-2025, PostgreSQL Global Development Group
-#
-# Minimal test testing shared_buffer resizing under load
-
-use strict;
-use warnings;
-use IPC::Run;
-use PostgreSQL::Test::Cluster;
-use PostgreSQL::Test::Utils;
-use Test::More;
-
-# Function to check if pgbench is still running.
-#
-# Relying on IPC::Run's pumpable status to check if pgbench is still running has
-# been proven unreliable. Instead we rely on existence of pgbench processes in
-# pg_stat_activity. Since we use -C with pgbench, there can be a non-zero
-# chance that no pgbench process is running even thought pgbench is running. But
-# that's a very rare possibility that can be ignored.
-sub pgbench_processes_active
-{
- my ($node, $application_name) = @_;
-
- my $result = $node->safe_psql('postgres',
- "SELECT count(*) FROM pg_stat_activity WHERE application_name = '$application_name';");
- return int($result) > 0;
-}
-
-my $resize_sql_func_def = q{
-create or replace function pg_resize_shared_buffers_sql(new_size int, out num_tries int) returns int as $$
-declare
- success boolean := false;
- tries int := 0;
- cur_setting text;
- pending_pattern text;
- target text := new_size::text;
-begin
- -- Wait until pg_settings reports the new value as pending,
- -- i.e. "<old value> (pending: <new value>)".
- pending_pattern := '%(pending: ' || target || ')%';
- loop
- select setting into cur_setting
- from pg_settings where name = 'shared_buffers';
- exit when cur_setting like pending_pattern or cur_setting = target;
- perform pg_sleep(0.1);
- raise notice 'Current setting: %', cur_setting;
- end loop;
-
- -- pg_resize_shared_buffers() returns true on success; retry until it succeeds.
- while not success loop
- tries := tries + 1;
- select pg_resize_shared_buffers() into success;
- if not success then
- perform pg_sleep(0.1);
- end if;
- raise notice 'pg_resize_shared_buffers() attempt %: success = %', tries, success;
- end loop;
-
- -- Confirm the new value is in effect (no longer pending).
- select setting into cur_setting
- from pg_settings where name = 'shared_buffers';
- if cur_setting <> target then
- raise exception 'shared_buffers resize did not take effect: expected %, got %',
- target, cur_setting;
- end if;
-
- num_tries := tries;
- return;
-end;
-$$ language plpgsql;
-};
-
-# Function to resize buffer pool and verify the change.
-sub apply_and_verify_buffer_change
-{
- my ($node, $new_size) = @_;
-
- # Use the new pg_resize_shared_buffers() interface which handles everything synchronously
- $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$new_size'");
- $node->safe_psql('postgres', "SELECT pg_reload_conf()");
- $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers_sql($new_size)");
-
- # Any failure in resizing the buffer pool will cause the test to timeout. So
- # if we reach here, the resize was successful. Just declare it as a
- # successful test so that we can see progress in the test output.
- ok(1, "Buffer pool resized to $new_size");
-}
-
-my @buffer_sizes = (128, 28, 16 * 1024, 32 * 1024, 1024, 512, 16, 24, 256, 128 * 1024, 16 * 1204);
-
-# Initialize a cluster and start pgbench in the background for concurrent load.
-my $node = PostgreSQL::Test::Cluster->new('main');
-$node->init;
-
-# Permit resizing up to 1GB for this test and let the server start with 128MB.
-$node->append_conf('postgresql.conf', qq{
-max_shared_buffers = } . (sort { $b <=> $a } @buffer_sizes)[0] . qq{
-shared_buffers = 16
-log_statement = none
-restart_after_crash = off
-});
-
-$node->start;
-$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
-$node->safe_psql('postgres', $resize_sql_func_def);
-
-my $pgb_scale = 10;
-my $pgb_duration = 120;
-my $pgb_num_clients = 10;
-# make it easy to identify pgbench processes in pg_stat_activity
-my $application_name = 'pgbench_buffer_resize_test';
-$node->pgbench(
- "--initialize --init-steps=dtpvg --scale=$pgb_scale --quiet",
- 0,
- [qr{^$}],
- [ # stderr patterns to verify initialization stages
- qr{dropping old tables},
- qr{creating tables},
- qr{done in \d+\.\d\d s }
- ],
- "pgbench initialization (scale=$pgb_scale)"
-);
-my ($pgbench_stdin, $pgbench_stdout, $pgbench_stderr) = ('', '', '');
-# Use --exit-on-abort so that the test stops on the first server crash or error,
-# thus making it easy to debug the failure. Use -C to increase the chances of a
-# new backend being created while resizing the buffer pool.
-my $pgbench_process = IPC::Run::start(
- [
- 'pgbench',
- '-p', $node->port,
- '-h', $node->host,
- '-T', $pgb_duration,
- '-c', $pgb_num_clients,
- '-C',
- '--exit-on-abort',
- '--continue-on-error',
- "dbname=postgres application_name=$application_name"
- ],
- '<' => \$pgbench_stdin,
- '>' => \$pgbench_stdout,
- '2>' => \$pgbench_stderr
-);
-
-ok($pgbench_process, "pgbench started successfully");
-
-# Resize buffer pool to various sizes while pgbench is running in the
-# background. We use smaller sizes to induce frequent buffer eviction and
-# allocation. Also smaller buffer pool means frequent wraparound in background
-# writer, default buffer allocation strategy and checkpointer.
-#
-# TODO: These are pseudo-randomly picked sizes, but we can do better.
-my $tests_completed = 0;
-
-# Reset background writer stats before starting the resize cycle
-$node->safe_psql('postgres', "SELECT pg_stat_reset_shared('bgwriter')");
-
-# Resize as many times as possible while pgbench is running.
-while (pgbench_processes_active($node, $application_name))
-{
- for my $target_size (@buffer_sizes)
- {
- # Stop if pgbench finished
- if (!pgbench_processes_active($node, $application_name))
- {
- last;
- }
-
- apply_and_verify_buffer_change($node, $target_size);
- $tests_completed++;
-
- # Wait for the resized buffer pool to stabilize.
- sleep(1);
- }
-}
-
-ok($tests_completed > scalar(@buffer_sizes), "All buffer size transitions were tested");
-note "Completed $tests_completed buffer resize operations while pgbench was running";
-
-# Check that the background writer did some work during the resize cycle
-is($node->safe_psql('postgres', "SELECT buffers_clean > 0 FROM pg_stat_bgwriter"), 't', "Background writer ran during resize cycle");
-
-# Make sure that pgbench finishes
-$pgbench_process->signal('TERM');
-ok((IPC::Run::finish $pgbench_process), "pgbench finished successfully");
-
-# Log any error output from pgbench for debugging
-diag("pgbench stderr:\n$pgbench_stderr");
-diag("pgbench stdout:\n$pgbench_stdout");
-
-# Ensure database is still functional after all the buffer changes
-$node->connect_ok("dbname=postgres",
- "Database remains accessible after $tests_completed buffer resize operations");
-
-done_testing();
diff --git a/src/test/buffermgr/t/003_resize_fault_tolerance.pl b/src/test/buffermgr/t/001_resize_fault_tolerance.pl
similarity index 99%
rename from src/test/buffermgr/t/003_resize_fault_tolerance.pl
rename to src/test/buffermgr/t/001_resize_fault_tolerance.pl
index 366929f4c45..87c95dbedf1 100644
--- a/src/test/buffermgr/t/003_resize_fault_tolerance.pl
+++ b/src/test/buffermgr/t/001_resize_fault_tolerance.pl
@@ -29,6 +29,13 @@ $node->append_conf('postgresql.conf', 'max_shared_buffers = 32');
$node->append_conf('postgresql.conf', 'restart_after_crash = on');
$node->start;
+# Bail out if this build does not support resizable shared memory, which
+# also means that resizing buffer pool is not supported.
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
# Load injection points extension for test coordination
$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
@@ -832,8 +839,4 @@ else
done_testing();
# Few more tests to add but may be somewhere else
-# TODO: test when there are backends that have not attached to the shared memory
# TODO: test that a non-superuser cannot run pg_resize_shared_buffers()
-# TODO: the resize_sql_func_def in 001_resize_buffer may be useful in other
-# tests (not necessarily this one). Maybe we can use it in other tests where we
-# are looping in TAP test code.
diff --git a/src/test/buffermgr/t/004_client_join_buffer_resize.pl b/src/test/buffermgr/t/002_client_join_buffer_resize.pl
similarity index 96%
rename from src/test/buffermgr/t/004_client_join_buffer_resize.pl
rename to src/test/buffermgr/t/002_client_join_buffer_resize.pl
index fda0f01bb27..6cca03d1816 100644
--- a/src/test/buffermgr/t/004_client_join_buffer_resize.pl
+++ b/src/test/buffermgr/t/002_client_join_buffer_resize.pl
@@ -78,20 +78,21 @@ max_parallel_workers_per_gather = 0
});
$node->start;
+# Bail out if this build does not support resizable shared memory, which
+# also means that resizing buffer pool is not supported.
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
# Enable injection points
$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
# Get the block size (this is fixed for the binary)
my $block_size = $node->safe_psql('postgres', "SHOW block_size");
# Try to create pg_buffercache extension for buffer analysis
-eval {
- $node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
-};
-if ($@) {
- $node->stop;
- plan skip_all => 'pg_buffercache extension not available - cannot verify buffer usage';
-}
# Create a small test table, and fetch its properties for later reference if required.
$node->safe_psql('postgres', qq{
diff --git a/src/test/buffermgr/t/005_resize_failures.pl b/src/test/buffermgr/t/003_resize_failures.pl
similarity index 96%
rename from src/test/buffermgr/t/005_resize_failures.pl
rename to src/test/buffermgr/t/003_resize_failures.pl
index 820aa07f006..21986733f74 100644
--- a/src/test/buffermgr/t/005_resize_failures.pl
+++ b/src/test/buffermgr/t/003_resize_failures.pl
@@ -24,6 +24,13 @@ $node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
$node->append_conf('postgresql.conf', "max_shared_buffers = $max_nbuffers");
$node->start;
+# Bail out if this build does not support resizable shared memory, which
+# also means that resizing buffer pool is not supported.
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
# pg_buffercache lets us locate the bufferid holding a given page.
$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
if ($have_injection_points)
diff --git a/src/test/buffermgr/t/006_resize_with_syslogger.pl b/src/test/buffermgr/t/004_resize_with_syslogger.pl
similarity index 87%
rename from src/test/buffermgr/t/006_resize_with_syslogger.pl
rename to src/test/buffermgr/t/004_resize_with_syslogger.pl
index 75b047ad984..e546b4c390e 100644
--- a/src/test/buffermgr/t/006_resize_with_syslogger.pl
+++ b/src/test/buffermgr/t/004_resize_with_syslogger.pl
@@ -25,6 +25,13 @@ logging_collector = on
});
$node->start;
+# Bail out if this build does not support resizable shared memory, which
+# also means that resizing buffer pool is not supported.
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
# Check that the syslogger is running by writing a log marker and waiting for it
# to appear in the log file.
sub check_syslogger_running
diff --git a/src/test/buffermgr/t/005_resize_unsupported.pl b/src/test/buffermgr/t/005_resize_unsupported.pl
new file mode 100644
index 00000000000..fcaf3311fea
--- /dev/null
+++ b/src/test/buffermgr/t/005_resize_unsupported.pl
@@ -0,0 +1,51 @@
+# Copyright (c) 2026-2026, PostgreSQL Global Development Group
+#
+# Test that pg_resize_shared_buffers() errors out when resizable shared
+# memory is not supported.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $initial_nbuffers = 256;
+my $max_nbuffers = 512;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf(
+ 'postgresql.conf', qq{
+shared_buffers = $initial_nbuffers
+max_shared_buffers = $max_nbuffers
+});
+$node->start;
+
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') eq 'on')
+{
+ # The builds that support resizable shared memory, usually, will not support
+ # the feature when SysV shared memory is used.
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_memory_type = 'sysv'");
+ $node->restart;
+}
+
+is( $node->safe_psql('postgres', 'SHOW have_resizable_shmem'),
+ 'off',
+ 'have_resizable_shmem reports off');
+
+my $target_nbuffers = $initial_nbuffers / 2;
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$target_nbuffers'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+my ($ret, $stdout, $stderr) =
+ $node->psql('postgres', "SELECT pg_resize_shared_buffers()");
+isnt($ret, 0,
+ 'pg_resize_shared_buffers fails when resizable shared memory is unsupported'
+);
+like(
+ $stderr,
+ qr/resizing shared buffer pool is not supported on this platform/,
+ 'error message reports that resizing shared buffer pool is unsupported'
+);
+
+done_testing();
diff --git a/src/test/buffermgr/t/010_stress_resize_buffer.pl b/src/test/buffermgr/t/010_stress_resize_buffer.pl
new file mode 100644
index 00000000000..70b33f83e97
--- /dev/null
+++ b/src/test/buffermgr/t/010_stress_resize_buffer.pl
@@ -0,0 +1,34 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Minimal stress test: resize shared_buffers repeatedly against regular pgbench
+# workload.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios.
+my @buffer_sizes =
+ (128, 28, 16 * 1024, 32 * 1024, 1024, 512, 16, 24, 256, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_buffer_resize_test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+
+$stress->setup;
+
+$stress->run;
+
+done_testing();
diff --git a/src/test/buffermgr/t/011_stress_drop_relation_buffers.pl b/src/test/buffermgr/t/011_stress_drop_relation_buffers.pl
new file mode 100644
index 00000000000..a3053a87d69
--- /dev/null
+++ b/src/test/buffermgr/t/011_stress_drop_relation_buffers.pl
@@ -0,0 +1,99 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Stress test execution of DropRelationBuffers(), DropRelationsAllBuffers() and
+# FlushRelationsAllBuffers() concurrently with shared_buffers resizing.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use List::Util qw(max);
+use PostgreSQL::Test::Utils;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios. At any time during the run buffer pool should be large enough to
+# let a new backend join while other backends are performing COPY, which seems
+# to pin many buffers at a time; avoid too small sizes.
+my @buffer_sizes = (512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+# The injection points verify that the buffer pool scan is exercised as expected
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_drop_relation_buffers_test',
+ injection_points => [
+ 'drop-relation-buffers-scan',
+ 'drop-relations-all-buffers-scan',
+ 'flush-relations-all-buffers-scan',
+ ],
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+$stress->setup;
+
+# Force execution of FlushRelationsAllBuffers() by skipping WAL logging DMLs to
+# a newly created table.
+my $node = $stress->node;
+$node->append_conf(
+ 'postgresql.conf', qq{
+wal_level = minimal
+max_wal_senders = 0
+wal_skip_threshold = 0
+});
+$node->restart;
+
+# Test specific load preparation.
+#
+# DropRelationBuffers() and DropRelationsAllBuffers() scan the buffer pool
+# only when the size of the relation exceeds NBuffers/32. Create a
+# relation larger than max(@buffer_sizes)/32 so a scan always runs. Dump
+# it once so pgbench clients can COPY it back in each iteration instead of
+# regenerating the data.
+my $tempdir = PostgreSQL::Test::Utils::tempdir;
+my $refdata_path = "$tempdir/refdata.bin";
+$node->safe_psql(
+ 'postgres', qq{
+ CREATE UNLOGGED TABLE refdata_source AS
+ SELECT g AS a, repeat('x', 1900)::bytea AS b
+ FROM generate_series(1, 16800) g;
+});
+$node->safe_psql('postgres',
+ "COPY refdata_source TO '$refdata_path' WITH (FORMAT binary)");
+my $max_nbuffers = max @buffer_sizes;
+my $required_pages = int($max_nbuffers / 32) + 1;
+my $refdata_pages = $node->safe_psql('postgres',
+ "SELECT (pg_relation_size('refdata_source') / current_setting('block_size')::bigint)::int"
+);
+cmp_ok($refdata_pages, '>', $required_pages,
+ "refdata_source spans more than NBuffers/32 pages at the largest tested pool size"
+);
+
+# Workload script fed to pgbench.
+#
+# Each client picks table names from a disjoint numeric range so table
+# names from concurrent clients do not collide. The :pgbench_id prefix
+# further disambiguates across the two pgbench flavors (persistent /
+# per_transaction) that StressUtil runs in parallel.
+my $workload_sql = qq{
+\\set tid :pgbench_id * 1000000000 + :client_id * 100000000 + random(1, 100000000)
+BEGIN;
+CREATE TABLE t_:tid (a int, b bytea);
+COPY t_:tid FROM '$refdata_path' WITH (FORMAT binary);
+TRUNCATE t_:tid; -- hit DropRelationBuffers()
+COPY t_:tid FROM '$refdata_path' WITH (FORMAT binary);
+COMMIT; -- hit FlushRelationsAllBuffers()
+DROP TABLE t_:tid; -- hit DropRelationsAllBuffers()
+};
+
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/012_stress_drop_database_buffers.pl b/src/test/buffermgr/t/012_stress_drop_database_buffers.pl
new file mode 100644
index 00000000000..77afb1f950d
--- /dev/null
+++ b/src/test/buffermgr/t/012_stress_drop_database_buffers.pl
@@ -0,0 +1,82 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Exercise the buffer pool scan that drops buffers belonging to a given
+# database (DropDatabaseBuffers()) concurrently with shared_buffers
+# resizing.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use List::Util qw(max);
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety of
+# scenarios. Avoid very small sizes because concurrent CREATE DATABASE clones
+# can pin more buffers than a very small pool provides, which can cause
+# unrelated failures (in particular, per-transaction pgbench connections
+# get "no unpinned buffers available" when opening a fresh backend).
+my @buffer_sizes = (256, 512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_drop_database_buffers_test',
+ injection_points => ['drop-database-buffers-scan'],
+
+ # Concurrent CREATE DATABASE clones pin many buffers per client; raising
+ # this can exhaust the small pool sizes resulting in the
+ # pg_resize_shared_buffers() query failing with and fail with error "no
+ # unpinned buffers available".
+ pgbench_clients => 4,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+
+$stress->setup;
+
+# Populate a template database with enough data. Size the seed table to roughly
+# 1/8 of the largest buffer pool we exercise. At the largest pool it still
+# occupies ~12% so that DropDatabaseBuffers() finds enough fraction of buffers
+# to drop. At smaller pool sizes the template exceeds the pool, thus covering
+# all the combinations of scanning buffer pool and dropping buffers.
+#
+# The 1900-byte payload makes sure that each row remains in the heap.
+my $node = $stress->node;
+my $seed_template = 'seedtemplate';
+my $max_buffer_pool = max @buffer_sizes;
+my $seed_row_count = 4 * $max_buffer_pool / 8;
+$node->safe_psql('postgres', "CREATE DATABASE $seed_template");
+$node->safe_psql(
+ $seed_template, qq{
+ CREATE TABLE seedtab (a int, b bytea);
+ INSERT INTO seedtab
+ SELECT g, repeat('x', 1900)::bytea
+ FROM generate_series(1, $seed_row_count) g;
+});
+$node->safe_psql('postgres',
+ "UPDATE pg_database SET datistemplate = true, datallowconn = false "
+ . "WHERE datname = '$seed_template'");
+
+# Workload script fed to pgbench.
+#
+# Each client picks database names from a disjoint numeric range so
+# database names from concurrent clients do not collide. The :pgbench_id
+# prefix further disambiguates across the two pgbench flavors
+# (persistent / per_transaction) that StressUtil runs in parallel.
+my $workload_sql = qq{
+\\set tid :pgbench_id * 1000000000 + :client_id * 100000000 + random(1, 100000000)
+CREATE DATABASE d_:tid TEMPLATE $seed_template;
+DROP DATABASE d_:tid;
+};
+
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/013_stress_checkpoint.pl b/src/test/buffermgr/t/013_stress_checkpoint.pl
new file mode 100644
index 00000000000..8c7d6bb5841
--- /dev/null
+++ b/src/test/buffermgr/t/013_stress_checkpoint.pl
@@ -0,0 +1,47 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Test synchronization between shared_buffers resize and CHECKPOINT. The test
+# issues explicit CHECKPOINTs so that the frequency of heckpoints can be
+# controlled so as increase the likelihood of concurrent checkpoint with every
+# resize.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios.
+my @buffer_sizes =
+ (128, 28, 16 * 1024, 32 * 1024, 1024, 512, 16, 24, 256, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_checkpoint_stress_test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+
+$stress->setup;
+
+# Disable implicit checkpoints
+$stress->node->append_conf(
+ 'postgresql.conf', qq{
+checkpoint_timeout = 1h
+max_wal_size = 100GB
+});
+$stress->node->reload;
+
+$stress->run(
+ workload_sql => "CHECKPOINT;\n",
+ workload_weight => 1,
+ default_load_weight => 10,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/014_stress_flush_relation_buffers.pl b/src/test/buffermgr/t/014_stress_flush_relation_buffers.pl
new file mode 100644
index 00000000000..52e326ef391
--- /dev/null
+++ b/src/test/buffermgr/t/014_stress_flush_relation_buffers.pl
@@ -0,0 +1,68 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Stress FlushRelationBuffers() concurrently with shared_buffers
+# resizing.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios. At any time during the run the buffer pool must be large
+# enough to let a new backend join while other backends are CLUSTERing a
+# small table, which pins several buffers at once; avoid too small sizes.
+my @buffer_sizes = (512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+# The injection point verifies that the buffer pool scan is exercised in
+# FlushRelationBuffers(). If the function stops scanning the buffer pool
+# the test is useless, so we assert that it fires at least once per
+# workload transaction.
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_flush_relation_buffers_test',
+ injection_points => ['flush-relation-buffers-scan'],
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+
+$stress->setup;
+
+my $node = $stress->node;
+my $ts2_dir = $node->basedir . '/tblsp2';
+mkdir $ts2_dir or die "mkdir $ts2_dir: $!";
+$node->safe_psql('postgres', "CREATE TABLESPACE ts2 LOCATION '$ts2_dir'");
+
+# Workload script fed to pgbench.
+#
+# Each client picks table names from a disjoint numeric range so table
+# names from concurrent clients do not collide. The :pgbench_id prefix
+# further disambiguates across the two pgbench flavors (persistent /
+# per_transaction) that StressUtil runs in parallel.
+my $workload_sql = qq{
+\\set tid :pgbench_id * 1000000000 + :client_id * 100000000 + random(1, 100000000)
+BEGIN;
+CREATE TABLE t_:tid (a int PRIMARY KEY, b bytea);
+INSERT INTO t_:tid
+ SELECT g, repeat('x', 1900)::bytea
+ FROM generate_series(1, 100) g;
+-- hit FlushRelationBuffers()
+ALTER TABLE t_:tid SET TABLESPACE ts2;
+ALTER TABLE t_:tid SET TABLESPACE pg_default;
+DROP TABLE t_:tid;
+COMMIT;
+};
+
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/015_stress_pg_buffercache.pl b/src/test/buffermgr/t/015_stress_pg_buffercache.pl
new file mode 100644
index 00000000000..4d823e1fb9d
--- /dev/null
+++ b/src/test/buffermgr/t/015_stress_pg_buffercache.pl
@@ -0,0 +1,60 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Stress the pg_buffercache diagnostic and monitoring functions
+# concurrently with shared_buffers resizing.
+#
+# Destructive functions like pg_buffercache_evict_* and
+# pg_buffercache_mark_dirty_* are excluded because they might cause failures
+# unrelated to the test.
+#
+# TODO: This test fails because pg_buffercache_os_pages_internal() expects the
+# buffer pool size to be constant. Fix is on the way.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+# A mix of small and large sizes exercises the resize logic in a variety
+# of scenarios. At any time during the run buffer pool should be large
+# enough to let a new backend join while other backends are scanning the
+# buffer pool via pg_buffercache; avoid too small sizes.
+my @buffer_sizes = (512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_pg_buffercache_test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+$stress->setup;
+
+$stress->node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache');
+
+# Test specific workload. Include NUMA view only if the server supports NUMA.
+my $workload_sql = qq{
+SELECT count(*) FROM pg_buffercache;
+SELECT * FROM pg_buffercache_summary();
+SELECT count(*) FROM pg_buffercache_usage_counts();
+SELECT count(*) FROM pg_buffercache_os_pages;
+};
+if ($stress->node->safe_psql('postgres', 'SELECT pg_numa_available()') eq 't')
+{
+ $workload_sql .= "SELECT count(*) FROM pg_buffercache_numa;\n";
+}
+
+# Use the default workload only to populate the buffer pool, but main workload
+# is the pg_buffercache queries.
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+done_testing();
diff --git a/src/test/buffermgr/t/016_stress_pg_prewarm.pl b/src/test/buffermgr/t/016_stress_pg_prewarm.pl
new file mode 100644
index 00000000000..23161f8865a
--- /dev/null
+++ b/src/test/buffermgr/t/016_stress_pg_prewarm.pl
@@ -0,0 +1,73 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Stress test the autoprewarm concurrently with shared_buffers resizing.
+#
+# TODO: This test fails because apw_dump_now() assumes NBuffers is
+# constant. Fix is on the way.
+
+use strict;
+use warnings;
+use FindBin;
+use lib $FindBin::RealBin;
+use Test::More;
+use StressUtil;
+
+if (!$ENV{PG_TEST_EXTRA} || $ENV{PG_TEST_EXTRA} !~ /\bbufmgr_stress\b/)
+{
+ plan skip_all => "bufmgr_stress not enabled in PG_TEST_EXTRA";
+}
+
+my @buffer_sizes = (512, 1024, 4096, 16 * 1024, 32 * 1024, 128 * 1024);
+
+my $stress = StressUtil->new(
+ buffer_sizes => \@buffer_sizes,
+ application_name => 'pgbench_pg_prewarm_test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,);
+$stress->setup;
+
+my $node = $stress->node;
+
+# Configure the autoprewarm background worker to run as frequently as possible.
+$node->append_conf(
+ 'postgresql.conf', qq{
+shared_preload_libraries = 'pg_prewarm'
+pg_prewarm.autoprewarm = on
+pg_prewarm.autoprewarm_interval = 1s
+});
+$node->restart;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_prewarm');
+
+# Dump the buffer pool through additional load to increase the likelihood of it
+# happening concurrently with a resize. Wrap the call in an exception block to
+# swallow "dump file is being used by PID N" errors.
+my $workload_sql = qq{
+DO \$\$
+BEGIN
+ PERFORM autoprewarm_dump_now();
+EXCEPTION WHEN OTHERS THEN
+ NULL;
+END
+\$\$;
+};
+
+$stress->run(
+ workload_sql => $workload_sql,
+ workload_weight => 10,
+ default_load_weight => 1,);
+
+my $dumpfile = $node->data_dir . '/autoprewarm.blocks';
+ok(-s $dumpfile, "buffer pool was dumped at least once");
+
+# Restart to confirm that the dump file can be read and the buffer pool can be
+# prewarmed.
+my $log_offset = -s $node->logfile;
+$node->restart;
+$node->wait_for_log(
+ qr/autoprewarm successfully prewarmed \d+ of \d+ previously-loaded blocks/,
+ $log_offset);
+pass("autoprewarm prewarmed shared buffers after restart");
+
+done_testing();
diff --git a/src/test/buffermgr/t/StressUtil.pm b/src/test/buffermgr/t/StressUtil.pm
new file mode 100644
index 00000000000..6a942421e09
--- /dev/null
+++ b/src/test/buffermgr/t/StressUtil.pm
@@ -0,0 +1,628 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+
+=pod
+
+=head1 NAME
+
+StressUtil - shared driver for the buffermgr shared_buffers resize stress tests
+
+=head1 SYNOPSIS
+
+ use StressUtil;
+
+ # Configure a stress test driver
+ my $stress = StressUtil->new(
+ buffer_sizes => [128, 28, ...],
+ application_name => 'pgbench_..._test',
+ pgbench_clients => 10,
+ pgbench_scale => 10,
+ pgbench_duration => 120,
+ injection_points => [...], # optional
+ );
+
+ # Setup the cluster and pgbench workload
+ $stress->setup;
+
+ # -- per-test prep goes here; may use $stress->node --
+
+ # Run the stress test with optional custom workload and perform post-stress
+ # checks
+ $stress->run(
+ default_load_weight => 1, # required if workload_sql set
+ workload_sql => $sql, # optional
+ workload_weight => 10, # required if workload_sql set
+ );
+
+=head1 DESCRIPTION
+
+StressUtil provides common routines for the shared_buffers resize stress tests.
+These are the routines for setting up the cluster and pgbench database, resizing
+shared_buffers in a tight loop while a pgbench workload runs concurrently, and
+performing post-stress checks.
+
+=cut
+
+package StressUtil;
+
+use strict;
+use warnings FATAL => 'all';
+
+use IPC::Run;
+use List::Util qw(max min shuffle);
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+=pod
+
+=head1 METHODS
+
+=over
+
+=item StressUtil->new(%opts)
+
+Construct a stress-test object which can be used to run the stress test with the
+given specifications. Named options:
+
+=over
+
+=item buffer_sizes
+
+Array reference of shared_buffers values (in number of buffers) that the resize
+loop cycles through. Required.
+
+=item application_name
+
+application_name string set on the pgbench connection; the resize loop
+polls pg_stat_activity for this value to detect when pgbench has
+exited. Required.
+
+=item injection_points
+
+Array reference of injection point names used to detect whether a code path is
+hit during stress test. They are attached with the B<notice> action. Number of
+times the notice message appears in the server error log indicates the number of
+times a certain code path is hit during the stress test. Defaults to the empty
+list. Used only when the build supports injection points.
+
+=item pgbench_clients
+
+Number of pgbench client connections. Split evenly across the persistent and
+per-transaction pgbench processes, so must be at least B<2>. May be overridden
+at run time by the C<PG_TEST_RESIZE_STRESS_CLIENTS> environment variable.
+
+=item pgbench_scale
+
+pgbench scale factor. May be overridden at run time by the
+C<PG_TEST_RESIZE_STRESS_SCALE> environment variable.
+
+=item pgbench_duration
+
+pgbench duration in seconds. May be overridden at run time by the
+C<PG_TEST_RESIZE_STRESS_SECONDS> environment variable so that the test never
+hits the timeout in a successful run.
+
+=back
+
+=cut
+
+sub new
+{
+ my ($class, %opts) = @_;
+ for my $required (
+ qw(buffer_sizes
+ application_name
+ pgbench_clients
+ pgbench_scale
+ pgbench_duration))
+ {
+ defined $opts{$required} or die "$required required";
+ }
+
+ # Env vars override the test-specified values.
+ my $duration =
+ int($ENV{PG_TEST_RESIZE_STRESS_SECONDS} // $opts{pgbench_duration});
+ my $clients =
+ int($ENV{PG_TEST_RESIZE_STRESS_CLIENTS} // $opts{pgbench_clients});
+ my $scale =
+ int($ENV{PG_TEST_RESIZE_STRESS_SCALE} // $opts{pgbench_scale});
+
+ my $timeout_cap = int(($ENV{PG_TEST_TIMEOUT_DEFAULT} // 0) * 0.8);
+ if ($timeout_cap > 0 && $duration > $timeout_cap)
+ {
+ note "clamping pgbench duration from $duration to $timeout_cap";
+ $duration = $timeout_cap;
+ }
+
+ $clients >= 2
+ or die "pgbench_clients must be at least 2 "
+ . "(split between persistent and per-transaction pgbench)";
+
+ my $self = {
+ buffer_sizes => $opts{buffer_sizes},
+ application_name => $opts{application_name},
+ pgbench_clients => $clients,
+ pgbench_scale => $scale,
+ pgbench_duration => $duration,
+ injection_points => $opts{injection_points} // [],
+ node => undef,
+ injection_points_supported => undef,
+ };
+ return bless $self, $class;
+}
+
+=pod
+
+=item $stress->node
+
+Return the underlying C<PostgreSQL::Test::Cluster> node. Valid only
+after setup().
+
+=cut
+
+sub node { return $_[0]->{node}; }
+
+=pod
+
+=item $stress->setup
+
+Create and initialize PostgreSQL cluster and other necessary objects required
+for the stress test.
+
+=cut
+
+sub setup
+{
+ my ($self) = @_;
+
+ my $node = PostgreSQL::Test::Cluster->new('main');
+ $node->init;
+ $self->{node} = $node;
+
+ my $max_buffer_pool = max @{ $self->{buffer_sizes} };
+ my $initial_buffers = min @{ $self->{buffer_sizes} };
+ my $ips_supported = ($ENV{enable_injection_points} // 'no') eq 'yes';
+ my $use_ips = $ips_supported && @{ $self->{injection_points} };
+ $self->{injection_points_supported} = $ips_supported;
+
+ $node->append_conf(
+ 'postgresql.conf', qq{
+max_shared_buffers = $max_buffer_pool
+shared_buffers = $initial_buffers
+log_statement = none
+restart_after_crash = off
+});
+
+ # Route injection-point NOTICEs to the server log, not to the pgbench
+ # client which does not expect them.
+ if ($use_ips)
+ {
+ $node->append_conf(
+ 'postgresql.conf', qq{
+shared_preload_libraries = injection_points
+log_min_messages = notice
+client_min_messages = warning
+});
+ }
+
+ $node->start;
+
+ # Bail out if this build does not support resizable shared memory, which
+ # also means that resizing buffer pool is not supported.
+ if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+ {
+ plan skip_all =>
+ "resizable shared memory not supported by this build";
+ }
+
+ $node->safe_psql('postgres', "CREATE EXTENSION buffermgr_test");
+ $node->safe_psql('postgres', "CREATE EXTENSION amcheck");
+
+ if ($use_ips)
+ {
+ $node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+ }
+
+ # Create a table to capture the history of resizes
+ $node->safe_psql(
+ 'postgres', qq{
+CREATE TABLE resize_log(
+ size int NOT NULL,
+ started_at timestamptz NOT NULL,
+ ended_at timestamptz NOT NULL,
+ num_tries int NOT NULL);
+});
+
+ # Reset the bgwriter stats so we can assert that it ran during the test.
+ $node->safe_psql('postgres', "SELECT pg_stat_reset_shared('bgwriter')");
+}
+
+=pod
+
+=item $stress->run(%opts)
+
+Run the stress test. This resizes shared_buffers repeatedly while a pgbench
+workload runs concurrently. Perform post-stress checks.
+
+Two pgbench processes run in parallel: one keeps its connections open for the
+whole run (persistent), the other reconnects for every transaction. The
+C<pgbench_clients> count is split evenly between them. Each pgbench is passed a
+distinct C<pgbench_id> variable and gets a distinct application_name so custom
+workloads can build names that do not collide across the two.
+
+pgbench always runs its built-in tpcb-like workload; an optional custom
+workload is added if requested.
+
+Named options:
+
+=over
+
+=item workload_sql
+
+Contents of a custom pgbench workload script. Optional.
+
+=item workload_weight
+
+Weight of the custom workload script relative to the built-in
+tpcb-like script. Required when C<workload_sql> is set; must not be
+set otherwise.
+
+=item default_load_weight
+
+Weight of the built-in tpcb-like workload relative to the custom
+workload script. Required when C<workload_sql> is set; must not be
+set otherwise.
+
+=back
+
+=cut
+
+sub run
+{
+ my ($self, %opts) = @_;
+ my $node = $self->{node} or die "setup() must be called first";
+ my $ips = $self->{injection_points};
+ my $use_ips = $self->{injection_points_supported} && @$ips;
+
+ my @pgbench_args;
+
+ my $workload_path;
+ if (defined $opts{workload_sql})
+ {
+ my $default_weight = $opts{default_load_weight}
+ // die "default_load_weight required when workload_sql is set";
+ my $workload_weight = $opts{workload_weight}
+ // die "workload_weight required when workload_sql is set";
+
+ push @pgbench_args, '-b', "tpcb-like\@$default_weight";
+
+ $workload_path = $node->basedir . '/workload.sql';
+ open(my $wfh, '>', $workload_path)
+ or die "cannot write $workload_path: $!";
+ print $wfh $opts{workload_sql};
+ close($wfh);
+ push @pgbench_args, '-f', "$workload_path\@$workload_weight";
+ }
+ elsif (defined $opts{default_load_weight}
+ || defined $opts{workload_weight})
+ {
+ die "default_load_weight and workload_weight require workload_sql";
+ }
+
+ # Attach injection points just before starting the workload.
+ if ($use_ips)
+ {
+ for my $ip (@$ips)
+ {
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('$ip', 'notice')");
+ }
+ }
+
+ my $log_offset = -s $node->logfile;
+
+ my @procs = _start_pgbench_workloads($self, \@pgbench_args);
+
+ _wait_for_pgbench_ready($node, $self->{application_name});
+
+ my $tests_completed = _run_resize_loop($self);
+
+ _run_post_checks($self, \@procs, $log_offset, $workload_path,
+ $tests_completed);
+}
+
+=back
+
+=cut
+
+# Resize the buffer pool and log the outcome to resize_log.
+sub _apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ $node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ # Start a new backend so that it inherits the reloaded GUC directly from the
+ # postmaster.
+ $node->safe_psql(
+ 'postgres', qq{
+ INSERT INTO resize_log(size, started_at, ended_at, num_tries)
+ SELECT $new_size, started_at, ended_at, num_tries
+ FROM pg_resize_shared_buffers_sql($new_size)
+ });
+
+ # A resize failure causes the test to time out, so reaching here means
+ # success.
+ ok(1, "buffer pool resized to $new_size");
+}
+
+# Return true if either pgbench workload is still running, false otherwise.
+#
+# IPC::Run's pumpable status is unreliable; check pg_stat_activity instead.
+sub _pgbench_processes_active
+{
+ my ($node, $application_name) = @_;
+
+ my $result = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_stat_activity "
+ . "WHERE application_name LIKE '${application_name}%'");
+ return int($result) > 0;
+}
+
+# Wait until at least one pgbench workload has registered in
+# pg_stat_activity, so the resize loop's first _pgbench_processes_active
+# check does not race pgbench startup and exit immediately.
+sub _wait_for_pgbench_ready
+{
+ my ($node, $application_name) = @_;
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) >= 1 FROM pg_stat_activity "
+ . "WHERE application_name LIKE '${application_name}%'")
+ or die "timed out waiting for pgbench workloads to connect";
+}
+
+# Initialize pgbench, then start two pgbench workloads in parallel: one
+# persistent, one with -C (per_transaction). Clients are split evenly. Each
+# pgbench gets a distinct application_name and a distinct :pgbench_id script
+# variable so custom workloads can build non-colliding names across the two.
+#
+# Returns a list of per-process hashes with keys process, stdout_ref,
+# stderr_ref, pgbench_id.
+sub _start_pgbench_workloads
+{
+ my ($self, $extra_args) = @_;
+ my $node = $self->{node};
+ my $scale = $self->{pgbench_scale};
+ my $app_name = $self->{application_name};
+ my $total_clients = $self->{pgbench_clients};
+
+ $node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$scale --quiet",
+ 0,
+ [qr{^$}],
+ [
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$scale)");
+
+ my @flavors = (
+ { pgbench_id => 1, extra => [] },
+ { pgbench_id => 2, extra => ['-C'] },);
+ my $clients_1 = int($total_clients / 2);
+ my $clients_2 = $total_clients - $clients_1;
+ $flavors[0]->{clients} = $clients_1;
+ $flavors[1]->{clients} = $clients_2;
+
+ my @procs;
+ for my $f (@flavors)
+ {
+ my $id = $f->{pgbench_id};
+ my $flavor_app = "${app_name}_${id}";
+ my ($stdin, $stdout, $stderr) = ('', '', '');
+ my $process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-h', $node->host,
+ '-T', $self->{pgbench_duration},
+ '-c', $f->{clients},
+ '-D', "pgbench_id=$id",
+ # stop on first server crash, so that conditions at the time of
+ # crash are preserved for diagnosis.
+ '--exit-on-abort',
+ '--continue-on-error',
+ @{ $f->{extra} },
+ @$extra_args,
+ "dbname=postgres application_name=$flavor_app"
+ ],
+ '<' => \$stdin,
+ '>' => \$stdout,
+ '2>' => \$stderr);
+
+ ok($process, "pgbench started successfully (pgbench_id=$id)");
+ push @procs,
+ {
+ process => $process,
+ stdout_ref => \$stdout,
+ stderr_ref => \$stderr,
+ pgbench_id => $id,
+ };
+ }
+ return @procs;
+}
+
+# Resize as many times as possible while pgbench is running, cycling
+# through $self->{buffer_sizes} in a shuffled order without ever picking
+# the same size twice in a row. Returns the number of resizes performed.
+sub _run_resize_loop
+{
+ my ($self) = @_;
+ my $node = $self->{node};
+ my $app_name = $self->{application_name};
+ my $buffer_sizes = $self->{buffer_sizes};
+ my @queue;
+ my $last_picked;
+ my $tests_completed = 0;
+
+ while (_pgbench_processes_active($node, $app_name))
+ {
+ if (!@queue)
+ {
+ @queue = shuffle(@$buffer_sizes);
+ if (defined $last_picked
+ && @queue > 1
+ && $queue[0] == $last_picked)
+ {
+ @queue[0, 1] = @queue[1, 0];
+ }
+ }
+ $last_picked = shift @queue;
+ _apply_and_verify_buffer_change($node, $last_picked);
+ $tests_completed++;
+ }
+ return $tests_completed;
+}
+
+# Assert the resize loop ran through at least one full sequence and
+# every size in @$buffer_sizes was picked at least once.
+sub _assert_all_sizes_used
+{
+ my ($node, $buffer_sizes, $tests_completed) = @_;
+
+ cmp_ok($tests_completed, '>', scalar(@$buffer_sizes),
+ "all buffer size transitions were tested");
+ note
+ "completed $tests_completed buffer resize operations while pgbench was running";
+
+ my $ndistinct = $node->safe_psql('postgres',
+ "SELECT count(DISTINCT size) FROM resize_log");
+ is($ndistinct, scalar(@$buffer_sizes),
+ "every buffer size was exercised at least once");
+}
+
+# Make sure that the pgbench workloads have ended and perform post-stress
+# checks.
+sub _run_post_checks
+{
+ my ($self, $procs, $log_offset, $workload_path, $tests_completed) = @_;
+
+ my $node = $self->{node};
+ my $ips = $self->{injection_points};
+
+ for my $p (@$procs)
+ {
+ my $id = $p->{pgbench_id};
+ $p->{process}->signal('TERM');
+ ok((IPC::Run::finish $p->{process}),
+ "pgbench finished successfully (pgbench_id=$id)");
+ note("pgbench_id=$id stderr:\n" . ${ $p->{stderr_ref} })
+ if ${ $p->{stderr_ref} } ne '';
+ note("pgbench_id=$id stdout:\n" . ${ $p->{stdout_ref} });
+ }
+
+ _assert_all_sizes_used($node, $self->{buffer_sizes}, $tests_completed);
+
+ # Log resize latency distribution and max retry count, for post-mortem
+ note $node->safe_psql(
+ 'postgres',
+ q{SELECT format('resize stats: n=%s, min=%s, avg=%s, max=%s, max_tries=%s',
+ count(*),
+ min(ended_at - started_at),
+ avg(ended_at - started_at),
+ max(ended_at - started_at),
+ max(num_tries))
+ FROM resize_log});
+
+ # Checkpointer activity, for post-mortem.
+ note $node->safe_psql(
+ 'postgres',
+ q{SELECT format('checkpointer stats: timed=%s, requested=%s, buffers_written=%s',
+ num_timed, num_requested, buffers_written)
+ FROM pg_stat_checkpointer});
+
+ is( $node->safe_psql(
+ 'postgres', "SELECT buffers_clean > 0 FROM pg_stat_bgwriter"),
+ 't',
+ "background writer ran during resize cycle");
+
+ # Server error log is expected to be crash free
+ $node->log_check("no PANIC or SIGBUS during stress run",
+ $log_offset, log_unlike => [ qr/PANIC/, qr/signal 7/ ]);
+
+ # pg_dumpall reads every table and catalog in every database; An error free
+ # dump indicates that the database remained non-corrupt after the stress
+ # run. We are not interested in the dump output, so discard it to /dev/null.
+ $node->command_ok(
+ [ 'pg_dumpall', '--no-sync', '-f', '/dev/null' ],
+ "pg_dumpall succeeds after stress run");
+
+ # pg_dumpall does not scan indexes; run bt_index_parent_check over every
+ # btree index to catch index-level corruption.
+ $node->safe_psql(
+ 'postgres', q{
+ SELECT bt_index_parent_check(c.oid, true, true)
+ FROM pg_class c
+ JOIN pg_index i ON i.indexrelid = c.oid
+ WHERE c.relkind = 'i'
+ AND c.relam = (SELECT oid FROM pg_am WHERE amname = 'btree')
+ });
+ ok(1, "all btree indexes verified");
+
+ # Verify tables
+ my $heap_findings = $node->safe_psql(
+ 'postgres', q{
+ SELECT count(*)
+ FROM (SELECT c.oid AS rel
+ FROM pg_class c
+ WHERE c.relkind IN ('r', 'S')
+ AND c.relpersistence = 'p') r,
+ LATERAL verify_heapam(r.rel, check_toast => true) v
+ });
+ is($heap_findings, '0', "verify_heapam found no corruption");
+
+ if (@$ips)
+ {
+ SKIP:
+ {
+ skip "injection points not supported by this build"
+ unless $self->{injection_points_supported};
+
+ my $workload_txns = 0;
+ for my $p (@$procs)
+ {
+ my $n;
+ if (defined $workload_path)
+ {
+ ($n) = ${ $p->{stdout_ref} } =~
+ m{SQL script \d+:\s+\Q$workload_path\E.*?number of transactions actually processed:\s*(\d+)}s;
+ }
+ else
+ {
+ ($n) = ${ $p->{stdout_ref} } =~
+ m{number of transactions actually processed:\s*(\d+)};
+ }
+ ok( defined $n,
+ "transaction count found in pgbench stdout (pgbench_id="
+ . $p->{pgbench_id} . ")");
+ $workload_txns += $n if defined $n;
+ }
+
+ my $log_content = slurp_file($node->logfile, $log_offset);
+ for my $ip (@$ips)
+ {
+ my $count = () = $log_content =~
+ /notice triggered for injection point $ip\b/g;
+ cmp_ok($count, '>=', $workload_txns,
+ "injection point $ip fired at least $workload_txns times"
+ );
+ }
+ }
+ }
+}
+
+1;
[application/octet-stream] v20260917-0007-Allow-to-resize-shared-buffers-without-restart.patch (178.8K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/10-v20260917-0007-Allow-to-resize-shared-buffers-without-restart.patch)
download | inline diff:
From ecfb341052f710714d61cb0f6940fcbc2d644cc4 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Sat, 6 Jun 2026 13:21:50 +0530
Subject: [PATCH] Allow to resize shared buffers without restart
User interface
==============
shared_buffers is now PGC_SIGHUP instead of PGC_POSTMASTER.
When a server is running, the new value of GUC (set using ALTER SYSTEM
... SET shared_buffers = ...; followed by SELECT pg_reload_conf()) does
not come into effect immediately. Instead a function
pg_resize_shared_buffers() is used to resize the buffer pool. The
function uses the current value of GUC in the backend where it is
executed. The function also coordinates the buffer access in other
backends while resizing the buffer pool.
SHOW shared_buffers now shows the current size of the shared buffer pool
but it also shows pending size of shared buffers, if any.
A new GUC max_shared_buffers is introduced to control the maximum value
of shared_buffers that can be set. By default it is 0. When explicitly
set, it needs to be higher than 'shared_buffers'. When max_shared_buffers
is set to 0, it assumes the same value as GUC shared_buffers. This GUC
determines the size of address space reserved for future buffer pool
sizes and the size of buffer look up table.
When shrinking the shared buffers pool, each buffer in the area being
shrunk needs to be flushed if it's dirty so as not to loose the changes
to that buffer after shrinking. Also, each such buffer needs to be
removed from the buffer mapping table so that backends do not access it
after shrinking. If a buffer being evicted is pinned, we abort the
resizing operation. There are other alternative which are not
implemented in the current patches 1. to wait for the pinned buffer to
get unpinned, 2. the backend is killed or it itself cancels the query
or 3. rollback the operation which is implemented currently. Note that
option 1 and 2 would require the pinning related local and shared
records to be accessed. But we need infrastructure to do either of this
right now.
Passing current buffer pool state to a new backend
==================================================
So far the buffer pool metdata (NBuffers and the shared memory segment
address space) is saved in process local heap memory since it's static
for the life of a server. It is passed to a new backend through
Postmaster. But with buffer pool being resized while the server running,
we need Postmaster to update its buffer pool metadata as the resizing
progresses and pass it to the new backend. This has few complications:
1. Postmaster does not receive ProcSignalBarrier. So we need to signal
it separately.
2. Postmaster's local state is inherited by the new backend when
fork()ed. But we need more complex implementation to pass it to an
exec()ed backend.
3. If Postmaster can not attend to its core functionality while it is
busy responding to the resizing signal
Instead, we maintain the buffer manager state in the shared memory.
Every backend maintains its own local copy of the state. When a backend
starts, it fetches syncs its local state with the global state before it
accesses any shared buffers. During run time it updates the local state
in response to the ProcSignalBarriers. The resizing does not advance
until the recent changes to the buffer pool state have been absorbed by
all the concurrent backends. This allows the backends to continue with
their regular activity without bothering about the buffer pool state
becoming inconsistent underneath.
Discussion points
=================
Removing the evicted buffers from buffer ring
---------------------------------------------
If the buffer pool has been shrunk, the buffers in the buffer ring may
not be valid anymore. Modify GetBufferFromRing to check if the buffer is
still valid before using it. This makes GetBufferFromRing() a bit more
expensive because of additional boolean condition and masks any bug that
introduces an invalid buffer into the ring. The alternative fix is more
complex as explained below.
The strategy object is created in CurrentMemoryContext and is not
available in any global structure hence inaccessible when processing
buffer resizing barriers. We may modify GetAccessStrategy() to register
strategy in a global linked list and then arrange to deregister it once
it's no more in use. Looking at the places which use
GetAccessStrategy(), fixing all those may be some work. So defering it
to v2.
Buffer lookup table
-------------------
In order to shrink the buffer lookup table, we need to compact the hash
table directory and the hash table entries so that the free space is
moved to the end of the memory allocated to the buffer lookup table.
This requires exclusively locking the shared hash table for a longer
duration, freezing the server for that duration. Furthermore the
compaction operation itself requires significant code. Hence we setup
the buffer lookup table considering the maximum possible size of the
buffer pool which is MaxAvailableMemory only once at the beginning. It
is not resized even though buffer pool is resized. We will need separate
effort later to implement a hash table which can be resized without
locking it for a longer duration.
BgWriter reset
--------------
The background writer makes sure to free buffers ahead of the clock
hand. For this it keeps track of the clock hand and jump forward if it
fells behind the clock hand. The position of clock hand is maintained as
a couple (size of buffer pool, position of next victim buffer). When
buffer pool size changes, the position of clock hand is adjusted
according to the new size. The background's knowledge of clock hand goes
out of sync when resize happens. Hence when resize happens the
background writer resets its knowledge of clock hand and jumps to the
clock hand's position. In case the background writer was ahead of clock
hand when resize happens, it will loose that advantage causing a
momentary glitch which may not be noticeable. It may be possible to
compare background writer's past knowledge of clock hand with the
current position by adjusting the previous according to the new size of
the buffer pool and avoid reset. But it requires more investigation
into background worker and clock hand synchronization. Hence deferred it
to a future version.
Fault tolerance of pg_resize_shared_buffers()
---------------------------------------------
If the backend executing pg_resize_shared_buffers() is interrupted
because of an ERROR, query cancellation, terminate signal, timeout etc.
the server is restarted to avoid leaving the buffer pool in an
inconsistent state. The server restarts with the buffer pool setup with
the new size. If the chances of such interruption are very low, this
solution might work for the first version. However, the restart can be
avoided by in following ways:
1. roll back the resize operation, if ERRORs are recoverable
2. In case of timeout and query cancellation, leave the resizing
operation with the buffer pool in degenerate but functioning state.
Re-attempting the resize would complete the operation or roll it
back.
3. In case of non-recoverable errors or terminate signal, it's better to
restart the operation.
But this needs more thought and implementation.
GUC type of shared_buffers
--------------------------
With the current code, on the platforms where have_resizable_shmem is
OFF, the users will be able to load new value of shared_buffers but not
resize the buffer pool. Probably we should change that set the type of
GUC to be POSTMASTER on those platforms. However, that still leaves the
server running one of the platforms which usually supports resizable
shared memory structure, but can not do so at run time because the
shared memory type does not support it, we can not do the same. Will
tackle this as the patches get finalized.
Calling BufferManagerInitProc()
-------------------------------
This function gets called twice in a backend startup sequence. Once so that
BufferManagerInitalizeAccess can access the buffer pool and second time after
ProcSignalInit() for the reasons mentioned there. I think we need a better place
so that we avoid calling it twice and set it only once properly. Also it's not
clear whether the function has been called at all the right places.
I am actually not sure why don't we call ProcSignalInit() right after
InitProcess() in a backend startup sequence. If that happens, we don't need to
worry about it.
Updating local activeNBuffers when allocating new buffer
--------------------------------------------------------
Local activeNBuffers is updated in response to ProcSignalBarrier. It is also
updated by ClockSweepTick() when choosing a victim buffer for the reasons
mentioned there. The GetStrategyBuffer() function which allocates a new buffer
from the buffer ring doesn't update the local activeNBuffers. I think it should
be harmless to just rely on the ProcSignalBarrier to update the local
activeNBuffers and everyone uses the same value. But it might interfere with the
way Clock hand is wrapped.
Barrier optimization
--------------------
Each resize operation uses three barriers to synchronize the buffer pool state
across all the backends. As noted at a few places in the code, we may be able to
use lesser number of barriers to reduce waiting time. However, we need to be
sure that we are not introducing any hazard by doing so. Hence we will keep the
current implementation for now and optimize it after we have enough tests to
cover all the scenarios.
Testing
-------
We have added a new test suite to test the buffer pool resizing. There is a
stress test exercising the buffer pool resizing with pgbench. There are also
some white box tests which use injection points to trigger specific scenarios.
However, we require more tests to cover all the scenarios.
Author: Ashutosh Bapat
Initial patches by: Dmitry Dolgov <9erthalion6@gmail.com>,
Inspired by patches from: Haoyu Huang <haoyu.huang.68@gmail.com>
Author of some tests: Palak Chaturvedi <chaturvedipalak1911@gmail.com>
Reviewed-by: Tomas Vondra
---
contrib/pg_buffercache/pg_buffercache_pages.c | 7 +
contrib/pg_prewarm/autoprewarm.c | 5 +
doc/src/sgml/config.sgml | 62 +-
doc/src/sgml/func/func-admin.sgml | 69 ++
src/backend/bootstrap/bootstrap.c | 2 +
src/backend/postmaster/auxprocess.c | 7 +
src/backend/postmaster/postmaster.c | 5 +
src/backend/storage/buffer/Makefile | 3 +-
src/backend/storage/buffer/README | 62 +-
src/backend/storage/buffer/buf_init.c | 299 ++++++-
src/backend/storage/buffer/buf_resize.c | 523 +++++++++++
src/backend/storage/buffer/buf_table.c | 27 +-
src/backend/storage/buffer/bufmgr.c | 176 +++-
src/backend/storage/buffer/freelist.c | 150 +++-
src/backend/storage/buffer/meson.build | 1 +
src/backend/storage/ipc/procsignal.c | 10 +
src/backend/storage/lmgr/proc.c | 31 +
src/backend/tcop/postgres.c | 3 +
src/backend/utils/init/globals.c | 2 +
src/backend/utils/init/postinit.c | 77 +-
src/backend/utils/misc/guc.c | 2 +-
src/backend/utils/misc/guc_parameters.dat | 15 +-
src/include/catalog/pg_proc.dat | 15 +
src/include/miscadmin.h | 3 +
src/include/storage/buf_internals.h | 48 +-
src/include/storage/bufmgr.h | 15 +
src/include/storage/procsignal.h | 5 +
src/include/utils/guc.h | 2 +
src/include/utils/guc_hooks.h | 2 +
src/test/Makefile | 3 +-
src/test/README | 3 +
src/test/buffermgr/Makefile | 35 +
src/test/buffermgr/README | 26 +
src/test/buffermgr/buffermgr_test.conf | 11 +
src/test/buffermgr/expected/buffer_resize.out | 290 ++++++
src/test/buffermgr/meson.build | 25 +
src/test/buffermgr/sql/buffer_resize.sql | 106 +++
src/test/buffermgr/t/001_resize_buffer.pl | 193 ++++
.../buffermgr/t/003_resize_fault_tolerance.pl | 839 ++++++++++++++++++
.../t/004_client_join_buffer_resize.pl | 221 +++++
src/test/buffermgr/t/005_resize_failures.pl | 171 ++++
.../buffermgr/t/006_resize_with_syslogger.pl | 58 ++
src/test/meson.build | 1 +
.../perl/PostgreSQL/Test/BackgroundPsql.pm | 76 ++
src/tools/pgindent/typedefs.list | 1 +
45 files changed, 3564 insertions(+), 123 deletions(-)
create mode 100644 src/backend/storage/buffer/buf_resize.c
create mode 100644 src/test/buffermgr/Makefile
create mode 100644 src/test/buffermgr/README
create mode 100644 src/test/buffermgr/buffermgr_test.conf
create mode 100644 src/test/buffermgr/expected/buffer_resize.out
create mode 100644 src/test/buffermgr/meson.build
create mode 100644 src/test/buffermgr/sql/buffer_resize.sql
create mode 100644 src/test/buffermgr/t/001_resize_buffer.pl
create mode 100644 src/test/buffermgr/t/003_resize_fault_tolerance.pl
create mode 100644 src/test/buffermgr/t/004_client_join_buffer_resize.pl
create mode 100644 src/test/buffermgr/t/005_resize_failures.pl
create mode 100644 src/test/buffermgr/t/006_resize_with_syslogger.pl
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 9512f1efa2f..312343fd7bf 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -289,6 +289,13 @@ pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
HeapTuple tuple;
Datum result;
+ /*
+ * TODO: This allocates memory using NBuffers which may change while this
+ * function is executed. We need to change this function so that it
+ * doesn't rely on NBuffers being static throughout the execution of this
+ * function.
+ */
+
if (SRF_IS_FIRSTCALL())
{
int i,
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index deb4c2671b5..e12323fb7c6 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -702,6 +702,11 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
return 0;
}
+ /*
+ * TODO: we need to modify this function to not rely on NBuffers being
+ * constant.
+ */
+
/*
* With sufficiently large shared_buffers, allocation will exceed 1GB, so
* allow for a huge allocation to prevent outright failure.
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index b84e7d6c799..8418af072f9 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -1802,7 +1802,6 @@ include_dir 'conf.d'
that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
(Non-default values of <symbol>BLCKSZ</symbol> change the minimum
value.)
- This parameter can only be set at server start.
</para>
<para>
@@ -1825,6 +1824,67 @@ include_dir 'conf.d'
appropriate, so as to leave adequate space for the operating system.
</para>
+ <para>
+ The shared memory consumed by the buffer pool is allocated and
+ initialized according to the value of the GUC at the time of starting
+ the server. A desired new value of GUC can be loaded while the server is
+ running using <systemitem>SIGHUP</systemitem>. But the buffer pool will
+ not be resized immediately. Use
+ <function>pg_resize_shared_buffers()</function> to dynamically resize
+ the shared buffer pool (see <xref linkend="functions-admin"/> for
+ details). If the running server has
+ <varname>have_resizable_shmem</varname> set to OFF,
+ <function>pg_resize_shared_buffers()</function> throws an error since
+ resizing a shared memory structure is not supported on that server. In
+ such a case the new value of <varname>shared_buffers</varname> gets
+ loaded using <function>pg_reload_conf</function> but the new size of the
+ pool takes effect only after restarting the server. <command>SHOW
+ shared_buffers</command> shows the currently effective value and any
+ pending value of the GUC. Please note that when the GUC is changed, the
+ other GUCS which use this GUCs value to set their defaults will not be
+ changed. They may still require a server restart to consider new value.
+ </para>
+
+ <para>
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-max-shared-buffers" xreflabel="max_shared_buffers">
+ <term><varname>max_shared_buffers</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>max_shared_buffers</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Sets the upper limit for the <varname>shared_buffers</varname> value.
+ The default value is <literal>0</literal>,
+ which means no explicit limit is set and <varname>max_shared_buffers</varname>
+ will be automatically set to the value of <varname>shared_buffers</varname>
+ at server startup.
+ If this value is specified without units, it is taken as blocks,
+ that is <symbol>BLCKSZ</symbol> bytes, typically 8kB.
+ This parameter can only be set at server start.
+ </para>
+
+ <para>
+ This parameter determines the amount of memory address space to reserve
+ in each backend for expanding the buffer pool in future. While the
+ memory for buffer pool is allocated on demand as it is resized, the
+ memory required for the buffer lookup table and the array used to sort
+ buffers during a checkpoint is allocated at the server start
+ considering the largest buffer pool size allowed by this parameter.
+ <!-- TODO: Provide a numeric example of how much extra memory say max_shared_buffers = 1GB consume. -->
+ </para>
+
+ <para>
+ When <varname>have_resizable_shmem</varname> is OFF, this parameter does
+ not have any effect except limiting the maximum value of
+ <varname>shared_buffers</varname>. It does not decide the address space
+ to be reserved or the size of the buffer lookup table and the array
+ used to sort buffers during a checkpoint.
+ </para>
</listitem>
</varlistentry>
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 0eae1c1f616..2a621fff060 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -99,6 +99,75 @@
<returnvalue>off</returnvalue>
</para></entry>
</row>
+
+ <row>
+ <entry role="func_table_entry"><para role="func_signature">
+ <indexterm>
+ <primary>pg_resize_shared_buffers</primary>
+ </indexterm>
+ <function>pg_resize_shared_buffers</function> ()
+ <returnvalue>boolean</returnvalue>
+ </para>
+ <para>
+ Dynamically resizes the shared buffer pool to match the current value of
+ the <varname>shared_buffers</varname> parameter in the client backend
+ where it is run. This function implements a coordinated resize process
+ that ensures all backend processes continue to operate without causing
+ any hazard. The resize happens in multiple phases to maintain data
+ consistency and system stability. Returns <literal>true</literal> if the
+ resize was successful, otherwise <literal>false</literal>. In the latter
+ case, the buffer pool is left at its previous size; this can happen if a
+ buffer being evicted during shrink could not be released, or if the
+ operating system could not supply enough shared memory to expand the
+ pool. Consult the server log for the underlying reason. This function
+ can only be called by superusers.
+ </para>
+ <para>
+ To resize shared buffers, first update the <varname>shared_buffers</varname>
+ setting and reload the configuration, then verify the new value is loaded
+ before calling this function. For example:
+<programlisting>
+postgres=# ALTER SYSTEM SET shared_buffers = '256MB'; -- Step 1
+ALTER SYSTEM
+postgres=# SELECT pg_reload_conf(); -- Step 2
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers; -- Step 3
+ shared_buffers
+-------------------------
+ 128MB (pending: 256MB)
+(1 row)
+
+postgres=# SELECT pg_resize_shared_buffers(); -- Step 4
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+postgres=# SHOW shared_buffers; -- Step 5 (verification)
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+</programlisting>
+ The <command>SHOW shared_buffers</command> at Step 3 is important to
+ verify that the configuration reload was successful and the new value is
+ available to the current session before attempting the resize. The
+ output shows both the current and pending values when the GUC change is
+ pending to be applied.
+ </para>
+ <para>
+ <!-- TODO: Document behaviour when the function is called on platforms that do not support resizable shared memory -->
+ linkend="functions-admin-signal-table"/> send control signals to
+ other server processes. Use of these functions is restricted to
+ superusers by default but access may be granted to others using
+ <command>GRANT</command>, with noted exceptions.
+ </para>
+ </entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index a678f345230..fa9a01a7ec3 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -376,6 +376,8 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
InitializeFastPathLocks();
+ InitializeMaxNBuffers();
+
ShmemCallRequestCallbacks();
CreateSharedMemoryAndSemaphores();
diff --git a/src/backend/postmaster/auxprocess.c b/src/backend/postmaster/auxprocess.c
index 07a3b5c5923..a09d376d333 100644
--- a/src/backend/postmaster/auxprocess.c
+++ b/src/backend/postmaster/auxprocess.c
@@ -19,6 +19,7 @@
#include "miscadmin.h"
#include "pgstat.h"
#include "postmaster/auxprocess.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/proc.h"
@@ -106,6 +107,12 @@ AuxiliaryProcessMainCommon(void)
*/
ShmemReprotectResizableStructs();
+ /*
+ * Update buffer manager's local state, which might have been changed by
+ * an online resize after startup.
+ */
+ BufferManagerInitProc();
+
RESUME_INTERRUPTS();
/*
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 90c7c4528e8..9a9e60e5fb1 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -959,6 +959,11 @@ PostmasterMain(int argc, char *argv[])
*/
InitializeFastPathLocks();
+ /*
+ * Calculate MaxNBuffers after NBuffersGUC has been set.
+ */
+ InitializeMaxNBuffers();
+
/*
* Also call any legacy shmem request hooks that might've been installed
* by preloaded libraries.
diff --git a/src/backend/storage/buffer/Makefile b/src/backend/storage/buffer/Makefile
index fd7c40dcb08..3bc9aee85de 100644
--- a/src/backend/storage/buffer/Makefile
+++ b/src/backend/storage/buffer/Makefile
@@ -17,6 +17,7 @@ OBJS = \
buf_table.o \
bufmgr.o \
freelist.o \
- localbuf.o
+ localbuf.o \
+ buf_resize.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index b332e002ba1..9ee05d562eb 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -181,8 +181,11 @@ buffer header spinlock, which would have to be taken anyway to increment the
buffer reference count, so it's nearly free.)
The "clock hand" is a buffer index, nextVictimBuffer, that moves circularly
-through all the available buffers. nextVictimBuffer is protected by the
-buffer_strategy_lock.
+through all the available buffers. Usually a victim can be chosen from the whole
+buffer pool, except when resizing the buffer pool. During resizing the victim
+can be chosen from a range of buffer pool which will not be affected by the
+resizing. See "Resizing shared buffers section" below. nextVictimBuffer is
+protected by the buffer_strategy_lock.
The algorithm for a process that needs to obtain a victim buffer is:
@@ -275,3 +278,58 @@ As of 8.4, background writer starts during recovery mode when there is
some form of potentially extended recovery to perform. It performs an
identical service to normal processing, except that checkpoints it
writes are technically restartpoints.
+
+Resizing shared buffers at runtime
+----------------------------------
+
+Before PostgreSQL 20, the size of the shared buffer pool (i.e. the number of
+shared buffers) was given by the global variable NBuffers and was fixed at
+server start time using GUC 'shared_buffers'. In order to change the size of the
+shared buffer pool, one needed to change the GUC and restart the server.
+Starting PostgreSQL 20, PostgreSQL supports resizing the buffer pool without a
+server restart. The value of the GUC 'shared_buffers' is stored in a new GUC
+variable called NBuffersGUC. The old variable NBuffers now strictly reflects the
+number of buffers in the buffer pool. The new GUC variable max_shared_buffers
+defines the maximum size of the shared buffer pool. See configure.sgml for more
+details about these GUCs.
+
+Resizing buffer pool involves resizing the data structures used by the buffer
+manager and coordinating the resize with all the backends.
+
+The buffer manager maintains following data structures in shared memory.
+1. Buffer Descriptors: An array of BufferDesc structures, one per buffer.
+2. Buffer Blocks: An array of buffer blocks, forming the buffer pool.
+3. Buffer Lookup Table: A hash table mapping a page to the buffer containing
+ that page.
+4. IO conditional variables: An array of conditional variables, one per buffer.
+5. Checkpoint buffer ids: An array of buffer ids used during checkpointing.
+6. Buffer Control: A structure containing information about the size of the
+ buffer pool and whether it is being resized.
+
+Except for the Buffer Lookup Table, Checkpoint buffer ids and Buffer Control
+all other structures are registered as resizable structures. We use
+ShmemResizeStruct() described in doc/src/sgml/xfunc.sgml to resize them. The
+hash table contains multiple shared structures, each of which needs to be
+resized separately. Since these substructures are not registered as separate
+structures, they can not be turned into resizable structures and hence the hash
+table is not registered as a resizable structure. Instead we set it up for
+maximal buffer pool at the server startup. The Checkpoint buffer ids array is
+filled in by the checkpointer with buffers to be written out during a
+checkpoint; if it were resized alongside the buffer pool, a shrink concurrent
+with an ongoing checkpoint could discard entries the checkpointer still needs
+to process. To avoid that, it is also sized for the maximal buffer pool at
+server startup.
+
+For ease of resizing, we differentiate between the number of buffers in the
+whole buffer pool (BufferControl::currentNBuffers) and the number of buffers at
+the start of the buffer pool from which a victim can be chosen for replacement
+(BufferControl::activeNBuffers). Each backend maintains a local copy of these
+two numbers. These copies provide a local view of the buffer pool that the
+backend can rely upon without worrying about the concurrent changes happening to
+the buffer pool. These copies are updated through a barrier mechanism when a
+resize is performed.
+
+To resize shared buffers at runtime a user performs the steps mentioned in the
+description of shared_buffers GUC variable in configure.sgml. Actual resizing
+protocol is documented in the prologue of function pg_resize_shared_buffers() in
+src/backend/storage/buffer/buf_resize.c
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 9ddf6551fcd..91671bce888 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,10 +17,15 @@
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/pg_shmem.h"
#include "storage/proclist.h"
#include "storage/shmem.h"
#include "storage/subsystems.h"
+#include "utils/guc.h"
+#include "utils/guc_hooks.h"
+#include "utils/injection_point.h"
+BufferControlBlock *BufferControl;
BufferDescPadded *BufferDescriptors;
char *BufferBlocks;
ConditionVariableMinimallyPadded *BufferIOCVArray;
@@ -37,6 +42,31 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
.attach_fn = BufferManagerShmemAttach,
};
+/*
+ * Resizable shared memory structures backing the buffer pool.
+ */
+static const struct
+{
+ const char *name;
+ size_t element_size;
+ size_t alignment;
+ void **ptr;
+} BufferManagerResizableStructs[] = {
+
+ {"Buffer Descriptors",
+ sizeof(BufferDescPadded),
+ PG_CACHE_LINE_SIZE,
+ (void **) &BufferDescriptors},
+ {"Buffer Blocks",
+ BLCKSZ,
+ PG_IO_ALIGN_SIZE,
+ (void **) &BufferBlocks},
+ {"Buffer IO Condition Variables",
+ sizeof(ConditionVariableMinimallyPadded),
+ PG_CACHE_LINE_SIZE,
+ (void **) &BufferIOCVArray},
+};
+
/*
* Data Structures:
* buffers live in a freelist and a lookup data structure.
@@ -69,6 +99,27 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
* multiple times. Check the PrivateRefCount infrastructure in bufmgr.c.
*/
+/*
+ * Initialize a single buffer.
+ */
+static void
+InitializeBuffer(int buf_id)
+{
+ /*
+ * Do not use GetBufferDescriptor here since it relies on the buffer
+ * descriptor being initialized.
+ */
+ BufferDesc *buf = &(BufferDescriptors[buf_id]).bufferdesc;
+
+ ClearBufferTag(&buf->tag);
+ pg_atomic_init_u64(&buf->state, 0);
+ buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ buf->buf_id = buf_id;
+ pgaio_wref_clear(&buf->io_wref);
+ proclist_init(&buf->lock_waiters);
+ ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+}
+
/*
* Register shared memory area for the buffer pool.
@@ -76,25 +127,21 @@ const ShmemCallbacks BufferManagerShmemCallbacks = {
static void
BufferManagerShmemRequest(void *arg)
{
- ShmemRequestStruct(.name = "Buffer Descriptors",
- .size = NBuffersGUC * sizeof(BufferDescPadded),
- /* Align descriptors to a cacheline boundary. */
- .alignment = PG_CACHE_LINE_SIZE,
- .ptr = (void **) &BufferDescriptors,
- );
-
- ShmemRequestStruct(.name = "Buffer Blocks",
- .size = NBuffersGUC * (Size) BLCKSZ,
- /* Align buffer pool on IO page size boundary. */
- .alignment = PG_IO_ALIGN_SIZE,
- .ptr = (void **) &BufferBlocks,
- );
+ /*
+ * Fall back to fixed sized shared buffer pool if resizable shared memory
+ * is not supported on this platform.
+ */
+#ifdef HAVE_RESIZABLE_SHMEM
+ bool resizable = (shared_memory_type == SHMEM_TYPE_MMAP);
+#else
+ bool resizable = false;
+#endif
+ int min_nbuffers = resizable ? MIN_NUM_BUFFERS : 0;
+ int max_nbuffers = resizable ? MaxNBuffers : 0;
- ShmemRequestStruct(.name = "Buffer IO Condition Variables",
- .size = NBuffersGUC * sizeof(ConditionVariableMinimallyPadded),
- /* Align descriptors to a cacheline boundary. */
- .alignment = PG_CACHE_LINE_SIZE,
- .ptr = (void **) &BufferIOCVArray,
+ ShmemRequestStruct(.name = "Buffer Control",
+ .size = sizeof(BufferControl),
+ .ptr = (void **) &BufferControl,
);
/*
@@ -103,11 +150,29 @@ BufferManagerShmemRequest(void *arg)
* memory at runtime. As that'd be in the middle of a checkpoint, or when
* the checkpointer is restarted, memory allocation failures would be
* painful.
+ *
+ * When the buffer pool is resizable, it is sized for MaxNBuffers up front
+ * so that the entries filled in by the checkpointer are not freed even
+ * when the buffer pool is shrunk.
*/
ShmemRequestStruct(.name = "Checkpoint BufferIds",
- .size = NBuffersGUC * sizeof(CkptSortItem),
+ .size = (size_t) (resizable ? MaxNBuffers : NBuffersGUC) * sizeof(CkptSortItem),
+ .alignment = PG_CACHE_LINE_SIZE,
.ptr = (void **) &CkptBufferIds,
);
+
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
+ {
+ size_t elem_size = BufferManagerResizableStructs[i].element_size;
+
+ ShmemRequestStruct(.name = BufferManagerResizableStructs[i].name,
+ .minimum_size = min_nbuffers * elem_size,
+ .size = NBuffersGUC * elem_size,
+ .maximum_size = max_nbuffers * elem_size,
+ .alignment = BufferManagerResizableStructs[i].alignment,
+ .ptr = BufferManagerResizableStructs[i].ptr,
+ );
+ }
}
/*
@@ -129,34 +194,194 @@ BufferManagerShmemInit(void *arg)
* Initialize all the buffer headers.
*/
for (int i = 0; i < NBuffers; i++)
+ InitializeBuffer(i);
+
+ /* Initialize BufferControl */
+ pg_atomic_init_u32(&BufferControl->currentNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->activeNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->targetNBuffers, NBuffersGUC);
+ pg_atomic_init_u32(&BufferControl->resizer_pid, 0);
+
+ /* Need to perform per backend steps in this backend too. */
+ BufferManagerShmemAttach(arg);
+}
+
+static void
+BufferManagerShmemAttach(void *arg)
+{
+ /* Initialize per-backend file flush context */
+ WritebackContextInit(&BackendWritebackContext,
+ &backend_flush_after);
+
+ BufferManagerInitProc();
+}
+
+/*
+ * Fetch latest buffer pool sizes (NBuffers and activeNBuffers) shared state
+ * into process local globals.
+ */
+void
+BufferManagerInitProc(void)
+{
+ NBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ elog(DEBUG1, "setting process local buffer pool sizes: currentNBuffers = %d, activeNBuffers = %d", NBuffers, activeNBuffers);
+}
+
+/*
+ * Protect unused shared memory reserved address space.
+ *
+ * Protect the parts of the shared memory address space reserved by the buffer
+ * manager which are not used by current structures from being accessed by
+ * backends.
+ *
+ * Unused address spaces of all resizable shared structures, including the
+ * buffer manager ones, are protected at the server startup together using
+ * ShmemProtectResizableStructs(). We need this function only during resizing of
+ * the buffer pool when we specifically adjust protections of buffer manager
+ * structures.
+ */
+void
+BufferManagerShmemProtect(void)
+{
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
{
- BufferDesc *buf = GetBufferDescriptor(i);
+ if (i == 2)
+ INJECTION_POINT("buffer-mgr-protect-struct", NULL);
+ ShmemProtectStruct(BufferManagerResizableStructs[i].name);
+ }
+}
+
+/*
+ * Resize and reinitialize shared buffer manager structures when resizing the
+ * buffer pool.
+ *
+ * Returns true if all the structures were resized successfully, false
+ * otherwise. We expect shrink to always succeed, but expansion may fail if the
+ * system is out of memory.
+ *
+ * The caller will always see all the structures resized consistently. If
+ * expanding a structure fails, all the expanded structures are shrunk back to
+ * their original sizes and no barrier is sent to the other backends.
+ */
+bool
+BufferManagerShmemResize(int currentNBuffers, int targetNBuffers)
+{
+#ifndef HAVE_RESIZABLE_SHMEM
+ ereport(ERROR,
+ errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("resizing shared buffer pool is not supported on this platform"));
+ pg_unreachable();
+#else
+ int resized = 0;
- ClearBufferTag(&buf->tag);
+ Assert(shared_memory_type == SHMEM_TYPE_MMAP);
- pg_atomic_init_u64(&buf->state, 0);
- buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
+ for (int i = 0; i < lengthof(BufferManagerResizableStructs); i++)
+ {
+ const char *name = BufferManagerResizableStructs[i].name;
+ size_t elem_size = BufferManagerResizableStructs[i].element_size;
+ bool resize_ok = true;
+
+#ifdef USE_INJECTION_POINTS
+ if (i == 2)
+ {
+ /* Injection point to simulate an interruption in this function. */
+ INJECTION_POINT("buffer-mgr-resize-struct", NULL);
- buf->buf_id = i;
+ /*
+ * Injection point to simulate a failure in resizing a structure
+ * like memory allocation failure without actually running out of
+ * memory.
+ */
+ if (IS_INJECTION_POINT_ATTACHED("buffer-mgr-resize-struct-fail"))
+ resize_ok = false;
+ }
+#endif
- pgaio_wref_clear(&buf->io_wref);
+ if (resize_ok)
+ resize_ok = ShmemResizeStruct(name, (size_t) targetNBuffers * elem_size);
- proclist_init(&buf->lock_waiters);
- ConditionVariableInit(BufferDescriptorGetIOCV(buf));
+ if (!resize_ok)
+ {
+ Assert(targetNBuffers > currentNBuffers);
+ for (int j = 0; j < resized; j++)
+ ShmemResizeStruct(BufferManagerResizableStructs[j].name,
+ (size_t) currentNBuffers * BufferManagerResizableStructs[j].element_size);
+ return false;
+ }
+ resized++;
}
- /* Initialize per-backend file flush context */
- WritebackContextInit(&BackendWritebackContext,
- &backend_flush_after);
+ /* Initialize the headers for new buffers. */
+ for (int i = currentNBuffers; i < targetNBuffers; i++)
+ InitializeBuffer(i);
+
+ return true;
+#endif
}
-static void
-BufferManagerShmemAttach(void *arg)
+/*
+ * check_shared_buffers
+ * GUC check_hook for shared_buffers
+ *
+ * When reloading the configuration, shared_buffers should not be set to a value
+ * higher than max_shared_buffers fixed at the boot time.
+ */
+bool
+check_shared_buffers(int *newval, void **extra, GucSource source)
{
- /* Update the size of the buffer pool. */
- NBuffers = NBuffersGUC;
+ if (finalMaxNBuffers && *newval > MaxNBuffers)
+ {
+ GUC_check_errdetail("\"shared_buffers\" must be less than \"max_shared_buffers\".");
+ return false;
+ }
+ return true;
+}
- /* Initialize per-backend file flush context */
- WritebackContextInit(&BackendWritebackContext,
- &backend_flush_after);
+/*
+ * show_shared_buffers
+ * GUC show_hook for shared_buffers
+ *
+ * Shows both current and pending buffer counts with proper unit formatting.
+ */
+const char *
+show_shared_buffers(bool use_units)
+{
+ static char buffer[128];
+ int64 current_value;
+ const char *current_unit;
+ int currentNBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+
+ if (use_units)
+ convert_int_from_base_unit(currentNBuffers, GUC_UNIT_BLOCKS, ¤t_value, ¤t_unit);
+ else
+ {
+ current_unit = "";
+ current_value = currentNBuffers;
+ }
+ snprintf(buffer, sizeof(buffer), INT64_FORMAT "%s", current_value, current_unit);
+
+ if (currentNBuffers != NBuffersGUC)
+ {
+ int64 pending_value;
+ const char *pending_unit;
+
+ /*
+ * Shared buffer pool is pending to be resized, show both current and
+ * pending sizes.
+ */
+ if (use_units)
+ convert_int_from_base_unit(NBuffersGUC, GUC_UNIT_BLOCKS, &pending_value, &pending_unit);
+ else
+ {
+ pending_value = NBuffersGUC;
+ pending_unit = "";
+ }
+ snprintf(buffer + strlen(buffer), sizeof(buffer) - strlen(buffer), " (pending: " INT64_FORMAT "%s)",
+ pending_value, pending_unit);
+ }
+
+ return buffer;
}
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
new file mode 100644
index 00000000000..c758493f353
--- /dev/null
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -0,0 +1,523 @@
+/*-------------------------------------------------------------------------
+ *
+ * buf_resize.c
+ * shared buffer pool resizing functionality
+ *
+ * This module contains the implementation of shared buffer pool resizing,
+ * including the main resize coordination function and barrier processing
+ * functions that synchronize all backends during resize operations.
+ *
+ * Portions Copyright (c) 1996-2026, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/storage/buffer/buf_resize.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "access/htup_details.h"
+#include "fmgr.h"
+#include "funcapi.h"
+#include "miscadmin.h"
+#include "postmaster/bgwriter.h"
+#include "storage/bufmgr.h"
+#include "storage/buf_internals.h"
+#include "storage/ipc.h"
+#include "storage/pg_shmem.h"
+#include "storage/pmsignal.h"
+#include "storage/procsignal.h"
+#include "storage/shmem.h"
+#include "utils/builtins.h"
+#include "utils/injection_point.h"
+
+static volatile sig_atomic_t safe_exit = true;
+
+static bool resize_shared_buffers_internal(void);
+static void buf_resize_shmem_exit(int code, Datum arg);
+
+/*
+ * Set the new buffer allocation pool size, broadcast it to all the backends
+ * and wait for them to acknowledge it.
+ */
+static void
+buf_resize_set_new_alloc_size(int alloc_size)
+{
+ uint64 generation;
+
+ pg_atomic_write_u32(&BufferControl->activeNBuffers, alloc_size);
+ StrategyAdjustNewBufAllocSize();
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC);
+ INJECTION_POINT("pgrsb-new-buffer-alloc-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC barrier");
+}
+
+/*
+ * Update the buffer pool size, broadcast it to all the backends and wait for
+ * them to acknowledge the change.
+ */
+static void
+buf_resize_set_buffer_pool_size(int new_size)
+{
+ uint64 generation;
+
+ pg_atomic_write_u32(&BufferControl->currentNBuffers, new_size);
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE);
+ INJECTION_POINT("pgrsb-buffer-pool-size-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE barrier");
+}
+
+/*
+ * Resize the shared buffer manager structures, broadcast the change to all
+ * the backends and wait for them to acknowledge it.
+ *
+ * If memory is not available when expanding the buffer pool, this function
+ * returns false without sending the barrier. When shrinking the buffer pool, we
+ * don't expect any failure, so this function always returns true.
+ */
+static bool
+buf_resize_shmem_resize(int currentNBuffers, int targetNBuffers)
+{
+ uint64 generation;
+
+ if (!BufferManagerShmemResize(currentNBuffers, targetNBuffers))
+ {
+ Assert(targetNBuffers > currentNBuffers);
+ return false;
+ }
+
+ generation = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE);
+ INJECTION_POINT("pgrsb-buffer-pool-resize-barrier-sent", NULL);
+ WaitForProcSignalBarrier(generation);
+ elog(LOG, "all backends acknowledged PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE barrier");
+ return true;
+}
+
+/*
+ * C implementation of SQL interface to update the shared buffers according to
+ * the current values of shared_buffers GUC.
+ *
+ * Atomic BufferControl::resizer_pid holds the PID of the backend currently
+ * performing a resize, or 0 when no resize is in progress. Using
+ * compare-and-exchange to set and reset this field, we make sure that only one
+ * resize is in progress at a time.
+ *
+ * Shrinking the buffer pool involves the following steps:
+ * - s1: Set BufferControl::activeNBuffers to the new size of the buffer pool
+ * and send SHBUF_NEW_BUFFER_ALLOC barrier to all backends. Every backend is
+ * expected to update their local buffer allocation pool size and acknowledge
+ * the barrier.
+ * - s2: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, new buffer allocations will be restricted
+ * to the new size of the buffer pool.
+ * - s3: Evict the buffers beyond the new size. A backend which still requires
+ * a previously allocated buffer which is being evicted, must have pinned it.
+ * If a pinned buffer is encountered, the resize operation is rolled back and
+ * the function returns false.
+ * - s4: If eviction succeeds, no backend should be using the buffers beyond
+ * the new size of the buffer pool and no new buffers can be allocated in
+ * that range. Update BufferControl::currentNBuffers to the new size of the
+ * buffer pool and send SHBUF_BUFFER_POOL_SIZE barrier to all backends. In
+ * response, all the backends should update their local buffer pool size and
+ * acknowledge the barrier.
+ * - s5: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, no backend will be accessing the shared
+ * buffer manager structures beyond the new size of the buffer pool.
+ * - s6: Resize the shared buffer manager structures to the new size (using
+ * ShmemResizeStruct()) and send SHBUF_BUFFER_POOL_RESIZE barrier to all
+ * backends. In response, all the backends should call ShmemProtectStruct()
+ * to update the memory address space protection of the shared buffer
+ * manager structures.
+ * - s7: Wait for all backends to acknowledge the barrier, before returning
+ * true to indicate successful resizing.
+ *
+ * Expanding the buffer pool involves the following steps:
+ * - e1: Resize the shared buffer manager structures to the new size (using
+ * ShmemResizeStruct()) and send SHBUF_BUFFER_POOL_RESIZE barrier to all the
+ * backends. In response, all the backends should call ShmemProtectStruct()
+ * to update the memory address space protection of the shared buffer
+ * manager structures. If expanding the shared buffer manager structures
+ * fails because of lack of memory, the function returns false without
+ * sending the barrier.
+ * - e2: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, every backend will be able to access the
+ * shared buffer manager structures beyond the old size of the buffer pool.
+ * - e3: Update BufferControl::currentNBuffers to the new size of the buffer
+ * pool and send SHBUF_BUFFER_POOL_SIZE barrier to all backends. In response,
+ * all the backends should update their local buffer pool size and
+ * acknowledge the barrier.
+ * - e4: Wait for all backends to acknowledge the barrier. When all backends
+ * have acknowledged the barrier, every backend is setup to use the new size
+ * of the buffer pool.
+ * - e5: Update BufferControl::activeNBuffers to the new size of the buffer
+ * pool, so that backends can start allocating from the new area of the
+ * buffer pool. Send SHBUF_NEW_BUFFER_ALLOC barrier to all backends. In
+ * response, all the backends should update their local buffer allocation
+ * pool size and acknowledge the barrier.
+ * - e6: Wait for all backends to acknowledge the barrier, before returning
+ * true to indicate successful resizing.
+ *
+ * Reason we introduce e4:
+ * Once we expand the new buffer allocation area, all the backends will start
+ * allocating buffers from the new area. Since this happens asynchronously,
+ * there is a chance that some backends may see buffers from outside their
+ * known buffer pool size. To avoid that, first set the new buffer pool size,
+ * broadcast it to all the backends and wait for them to update their
+ * knowledge of buffer pool size. We may be able to avoid sending the barrier
+ * after the first step or avoid them altogether but it's not clear that that
+ * is completely hazard free. It feels safer this way, even though it takes
+ * longer.
+ *
+ * If a timeout happens or request to cancel query arrives while the function is
+ * being executed, we need to abort the operation immediately. The state of the
+ * buffer pool and its state as viewed by the backends may not be consistent at
+ * that point. Hence we escalate it to PANIC to restart the server and avoid
+ * inconsistent state. We may improve this situation by leaving the buffer pool
+ * in a consistent but degenerate state and allowing a subsequent resize
+ * operation to rollback or continue the operation.
+ *
+ * If an ERROR is raised while the function is being executed, we may have
+ * already entered an inconsistent state. Hence we escalate it to PANIC to
+ * restart the server and avoid inconsistent state.
+ */
+Datum
+pg_resize_shared_buffers(PG_FUNCTION_ARGS)
+{
+ bool success = false;
+
+ /*
+ * Register the exit hook before claiming resizer_pid, so that if we exit
+ * after claiming resizer_pid, the hook is in place to reset it.
+ */
+ before_shmem_exit(buf_resize_shmem_exit, 0);
+
+ PG_TRY();
+ {
+ uint32 expected_pid = 0;
+
+ if (!pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, MyProcPid))
+ {
+ /*
+ * Another backend holds resizer_pid; expected_pid was updated by
+ * the CAS to reflect its PID.
+ */
+ elog(LOG, "shared buffer resize already in progress in backend %u",
+ expected_pid);
+ /* No shared memory was touched, so it should be safe to exit. */
+ Assert(safe_exit);
+ }
+ else
+ {
+ INJECTION_POINT("pg-resize-shared-buffers-flag-set", NULL);
+
+ /*
+ * We are about to make changes to the shared memory which can not
+ * be rolled back easily since we need all the backends to
+ * acknowledge these changes. Indicate that a sudden exit in this
+ * state can leave the server in an inconsistent state.
+ */
+ safe_exit = false;
+ success = resize_shared_buffers_internal();
+
+ /*
+ * The changes to shared memory are in a consistent state across
+ * all the backends, so it should be safe to exit.
+ */
+ safe_exit = true;
+ }
+ }
+ PG_FINALLY();
+ {
+ uint32 expected_pid = MyProcPid;
+
+ /*
+ * We are in the middle of resizing and caught an error. Without
+ * knowing the reason and exact state of resizing it's not safe to
+ * continue or to exit. Restarting the server is the safest option
+ * here. Emit the error to the server log and raise PANIC to restart
+ * the server.
+ */
+ if (!safe_exit)
+ {
+ HOLD_INTERRUPTS();
+ errcontext("during shared buffer resize");
+ EmitErrorReport();
+ ereport(PANIC,
+ errmsg("shared buffer resize caught an error when shared memory was in an inconsistent state"));
+ pg_unreachable();
+ }
+
+ /*
+ * Reset the PID, if we set it before removing the shmem_exit hook so
+ * as not to leave it set after the backend has exited.
+ */
+ (void) pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, 0);
+ cancel_before_shmem_exit(buf_resize_shmem_exit, 0);
+ }
+ PG_END_TRY();
+
+ if (success)
+ elog(LOG, "shared buffer resizing to %d buffers completed successfully", NBuffersGUC);
+ else
+ elog(WARNING, "shared buffer resizing to %d buffers failed", NBuffersGUC);
+
+ PG_RETURN_BOOL(success);
+}
+
+/*
+ * Workhorse function for the C implementation.
+ */
+static bool
+resize_shared_buffers_internal(void)
+{
+ int currentNBuffers;
+ int targetNBuffers;
+ bool resize_success;
+
+ currentNBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+ targetNBuffers = NBuffersGUC;
+ if (currentNBuffers == targetNBuffers)
+ {
+ elog(LOG, "shared buffers are already at %d, no need to resize", currentNBuffers);
+ return true;
+ }
+
+ /*
+ * TODO: What if the NBuffersGUC value seen here is not the desired one
+ * because somebody did a pg_reload_conf() between the last
+ * pg_reload_conf() and execution of this function?
+ */
+
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, targetNBuffers);
+ elog(LOG, "resizing shared buffers from %d to %d", currentNBuffers, targetNBuffers);
+
+ if (targetNBuffers < currentNBuffers)
+ {
+ /*
+ * step s1, s2: Restrict new buffer allocations to the new buffer pool
+ * size.
+ *
+ * TODO: Alternate design idea by Andres (as I understand it): Set
+ * BufferControl::activeNBuffers and send the barrier. Instead of
+ * waiting for barrier, start evicting buffers but don't unpin the
+ * evicted buffers so that they will not be considered for new
+ * allocations. Once all the buffers are evicted wait for the barrier
+ * to be acknowledged. This will reduce the time taken to shrink the
+ * buffer pool.
+ */
+ elog(LOG, "shrinking buffer pool, restricting allocations to %d buffers", targetNBuffers);
+ buf_resize_set_new_alloc_size(targetNBuffers);
+
+ /* Step s3: Evict buffers in the area being shrunk */
+ elog(LOG, "evicting buffers %u..%u", targetNBuffers + 1, currentNBuffers);
+ if (!EvictExtraBuffers(targetNBuffers, currentNBuffers))
+ {
+ elog(WARNING, "failed to evict extra buffers during shrinking");
+
+ /* Eviction failed, rollback the buffer resize operation. */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, currentNBuffers);
+ buf_resize_set_new_alloc_size(currentNBuffers);
+ return false;
+ }
+
+ /* Step s4, s5: Update the buffer pool size. */
+ buf_resize_set_buffer_pool_size(targetNBuffers);
+ }
+
+ /* Step s6, s7 or e1, e2: Resize the buffer manager structures. */
+ resize_success = buf_resize_shmem_resize(currentNBuffers, targetNBuffers);
+
+ if (targetNBuffers > currentNBuffers)
+ {
+ if (!resize_success)
+ {
+ elog(WARNING, "failed to expand buffer pool structures");
+
+ /* Revert any changes to the shared memory in this function. */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, currentNBuffers);
+ return false;
+ }
+
+ /* Step e3, e4: Declare new buffer pool size. */
+ buf_resize_set_buffer_pool_size(targetNBuffers);
+
+ /* Step e5, e6: Let expanded buffer pool be used by all backends. */
+ buf_resize_set_new_alloc_size(targetNBuffers);
+ }
+
+ return true;
+}
+
+/*
+ * Function to handle process exit when buffer resizing is in progress.
+ */
+static void
+buf_resize_shmem_exit(int code, Datum arg)
+{
+ uint32 expected_pid;
+
+ /*
+ * If resizer_pid does not match our PID, either we never claimed it or we
+ * have already released it. Nothing to do.
+ */
+ if (pg_atomic_read_u32(&BufferControl->resizer_pid) != MyProcPid)
+ return;
+
+ /*
+ * Resize is in progress and the process crashed. We do not know exactly
+ * at which step of the resizing we are. Just restart the server to be
+ * safe.
+ *
+ * TODO: If we can perform heavy operations in this callback like waiting
+ * for barriers, we could set the current status of resizing in the
+ * process local memory and use this callback to rollback every operation
+ * that was performed, except buffer eviction.
+ */
+ if (!safe_exit)
+ ereport(PANIC,
+ errmsg("buffer resize operation interrupted, restarting to avoid inconsistent state"));
+
+ /*
+ * safe_exit should be set to true when new allocations are not using the
+ * whole buffer pool, so the following condition should never happen. But
+ * be on the safer side.
+ */
+ if (pg_atomic_read_u32(&BufferControl->currentNBuffers) != pg_atomic_read_u32(&BufferControl->activeNBuffers))
+ ereport(PANIC,
+ (errmsg("buffer resize operation interrupted at an unexpected stage, restarting to avoid inconsistent state")));
+
+ /*
+ * Reset targetNBuffers before releasing resizer_pid, so that a backend
+ * claiming ownership immediately afterwards does not have its own
+ * targetNBuffers overwritten by us.
+ */
+ pg_atomic_write_u32(&BufferControl->targetNBuffers, pg_atomic_read_u32(&BufferControl->currentNBuffers));
+
+ expected_pid = MyProcPid;
+ (void) pg_atomic_compare_exchange_u32(&BufferControl->resizer_pid,
+ &expected_pid, 0);
+}
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC.
+ */
+bool
+ProcessBarrierNewBufferAlloc(void)
+{
+ elog(DEBUG2, "processing barrier to restrict new buffer allocations to %d buffers (target = %d)",
+ pg_atomic_read_u32(&BufferControl->activeNBuffers), pg_atomic_read_u32(&BufferControl->targetNBuffers));
+
+ INJECTION_POINT("pgrsb-handle-new-buffer-alloc-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(NBuffers == pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ return true;
+}
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE.
+ */
+bool
+ProcessBarrierBufferPoolResize(void)
+{
+ elog(DEBUG2, "processing barrier to propagate resized shared buffer pool structures");
+
+ INJECTION_POINT("pgrsb-handle-buffer-pool-resize-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(NBuffers == pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ Assert(activeNBuffers == pg_atomic_read_u32(&BufferControl->activeNBuffers));
+
+ /*
+ * Access permissions to address range covered by a resizable structure is
+ * maintained consistently across all the backends right from the time a
+ * backend is started. We maintain that consistency as the buffer pool is
+ * resized.So ideally modifying the access permissions in this backend
+ * should not fail. But if it does, the address space accessible to this
+ * backend may be inconsistent with the new buffer pool size and also with
+ * the other backends. This may cause data corruption and other memory
+ * access issues, if we let this backend continue to run and access the
+ * buffer pool. Better to quit from the faulty backend.
+ */
+ PG_TRY();
+ {
+ BufferManagerShmemProtect();
+ }
+ PG_CATCH();
+ {
+ /*
+ * We don't know what caused the error and so avoid using further
+ * resources. Emit the original error to the server log so that it's
+ * not lost and raise a FATAL to terminate this backend.
+ */
+ HOLD_INTERRUPTS();
+ errcontext("during shared buffer pool resize barrier");
+ EmitErrorReport();
+ ereport(FATAL,
+ (errmsg("shared buffer pool resize barrier caught an error while updating buffer pool protection")));
+ pg_unreachable();
+ }
+ PG_END_TRY();
+
+ return true;
+}
+
+/*
+ * Process and acknowledge PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE.
+ */
+bool
+ProcessBarrierBufferPoolSize(void)
+{
+ elog(DEBUG2, "processing barrier to establish new size of the buffer pool to %d", pg_atomic_read_u32(&BufferControl->currentNBuffers));
+
+ INJECTION_POINT("pgrsb-handle-buffer-pool-size-barrier", NULL);
+
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ Assert(activeNBuffers == pg_atomic_read_u32(&BufferControl->activeNBuffers));
+ NBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+
+ return true;
+}
+
+/*
+ * SQL-callable function reporting the current shared buffer pool resize
+ * status.
+ */
+Datum
+pg_get_buffer_resize_status(PG_FUNCTION_ARGS)
+{
+#define PG_GET_BUFFER_RESIZE_STATUS_COLS 4
+ TupleDesc tupdesc;
+ Datum values[PG_GET_BUFFER_RESIZE_STATUS_COLS];
+ bool nulls[PG_GET_BUFFER_RESIZE_STATUS_COLS] = {0};
+ HeapTuple tuple;
+
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+ tupdesc = BlessTupleDesc(tupdesc);
+
+ values[0] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->activeNBuffers));
+ values[1] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->currentNBuffers));
+ values[2] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->targetNBuffers));
+ values[3] = Int32GetDatum((int32) pg_atomic_read_u32(&BufferControl->resizer_pid));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ PG_RETURN_DATUM(HeapTupleGetDatum(tuple));
+#undef PG_GET_BUFFER_RESIZE_STATUS_COLS
+}
+
+/*
+ * TODO: add progress report facility if required.
+ */
diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c
index c82b71deaa6..85664990b3c 100644
--- a/src/backend/storage/buffer/buf_table.c
+++ b/src/backend/storage/buffer/buf_table.c
@@ -28,6 +28,7 @@
#include "utils/builtins.h"
#include "storage/lwlock.h"
#include "storage/subsystems.h"
+#include "storage/pg_shmem.h"
/* entry for buffer lookup hashtable */
typedef struct
@@ -58,15 +59,25 @@ BufTableShmemRequest(void *arg)
* Request the shared buffer lookup hashtable.
*
* Since we can't tolerate running out of lookup table entries, we must be
- * sure to specify an adequate table size here. The maximum steady-state
- * usage is of course as many entries as the number of buffers in the
- * pool, but BufferAlloc() tries to insert a new entry before deleting the
- * old. In principle this could be happening in each partition
- * concurrently, so we could need as many as (number of buffers in the
- * pool) + NUM_BUFFER_PARTITIONS entries. Since we are still requesting
- * shared memory, use the GUC value instead of the actual size.
+ * sure to specify an adequate table size here. The maximum number of
+ * entries we could need is the number of buffers, but BufferAlloc() tries
+ * to insert a new entry before deleting the old. In principle this could
+ * be happening in each partition concurrently, so we could need as many
+ * as (number of buffers) + NUM_BUFFER_PARTITIONS entries. The number of
+ * buffers can be increased upto MaxNBuffers at run time. So we need to
+ * make sure that the table can accommodate MaxNBuffers +
+ * NUM_BUFFER_PARTITIONS entries.
+ *
+ * See storage/buffer/README for reasons why we don't register the hash
+ * table as a resizable structure.
+ *
*/
- size = NBuffersGUC + NUM_BUFFER_PARTITIONS;
+#ifdef HAVE_RESIZABLE_SHMEM
+ size = (shared_memory_type == SHMEM_TYPE_MMAP) ? MaxNBuffers : NBuffersGUC;
+#else
+ size = NBuffersGUC;
+#endif
+ size = size + NUM_BUFFER_PARTITIONS;
ShmemRequestHash(.name = "Shared Buffer Lookup Table",
.nelems = size,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 9f86c3320e1..cebc85624b8 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -223,7 +223,10 @@ int io_max_combine_limit = DEFAULT_IO_COMBINE_LIMIT;
int checkpoint_flush_after = DEFAULT_CHECKPOINT_FLUSH_AFTER;
int bgwriter_flush_after = DEFAULT_BGWRITER_FLUSH_AFTER;
int backend_flush_after = DEFAULT_BACKEND_FLUSH_AFTER;
+
int NBuffers = 0; /* number of buffers in the buffer pool */
+int activeNBuffers = 0; /* number of buffers at the start of the pool
+ * from which new buffer allocations happnen. */
/* local state for LockBufferForCleanup */
static BufferDesc *PinCountWaitBuf = NULL;
@@ -814,6 +817,15 @@ PrefetchBuffer(Relation reln, ForkNumber forkNum, BlockNumber blockNum)
* Compared to ReadBuffer(), this avoids a buffer mapping lookup when it's
* successful. Return true if the buffer is valid and still has the expected
* tag. In that case, the buffer is pinned and the usage count is bumped.
+ *
+ * The callers of this function should make sure that the buffer is valid even
+ * if the shared buffer pool has undergone a resize. Buffer pool resizing waits
+ * for all backends to acknowledge the barrier before changing the buffer pool
+ * size. Hence the caller should call this function after validation without an
+ * intervening ProcSignalBarrier processing.
+ *
+ * TODO: This function could perform the validation in this function itself
+ * instead of relying on the two callers who do it currently.
*/
bool
ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockNum,
@@ -3655,6 +3667,11 @@ BufferSync(int flags)
TRACE_POSTGRESQL_BUFFER_SYNC_START(NBuffers, num_to_scan);
+ /*
+ * TODO: Test the case when buffer pool is shrunk after CkptBufferIds is
+ * filled and num_to_scan is higher than the new NBuffers?
+ */
+
/*
* Sort buffers that need to be written to reduce the likelihood of random
* IO. The sorting is also important for the implementation of balancing
@@ -3768,6 +3785,21 @@ BufferSync(int flags)
buf_id = CkptBufferIds[ts_stat->index].buf_id;
Assert(buf_id != -1);
+ /*
+ * TODO: We need to test the scenario when the buffer pool is shrunk
+ * after checkpointer has collected the buffer ids and one or more of
+ * the buffer ids is out of range.
+ */
+
+ /*
+ * The buffer pool might have been shrunk between the time the
+ * checkpoint collected the buffer ids and now. Ignore any buffers
+ * that are out of range now. Those buffers must have been written
+ * when they were evicted during resizing.
+ */
+ if (buf_id >= NBuffers)
+ continue;
+
bufHdr = GetBufferDescriptor(buf_id);
num_processed++;
@@ -3843,6 +3875,13 @@ BufferSync(int flags)
/*
* BgBufferSync -- Write out some dirty buffers in the pool.
*
+ * Usually new buffers are allocated from the whole buffer pool. However, when
+ * resizing the buffer pool, the new buffer allocations are restricted to some
+ * initial portion of the buffer pool. The rest of pool is either being evicted
+ * when shrinking or not utilized yet when growing, hence not interesting to the
+ * background writer. Hence the background writer always restricts its activity
+ * to the same portion of the buffer pool as new buffer allocations.
+ *
* This is called periodically by the background writer process.
*
* Returns true if it's appropriate for the bgwriter process to go into
@@ -3894,6 +3933,7 @@ BgBufferSync(WritebackContext *wb_context)
/* Variables for final smoothed_density update */
long new_strategy_delta;
uint32 new_recent_alloc;
+ static int prev_activeNBuffers;
Assert(AmBackgroundWriterProcess());
@@ -3903,6 +3943,24 @@ BgBufferSync(WritebackContext *wb_context)
*/
strategy_buf_id = StrategySyncStart(&strategy_passes, &recent_alloc);
+ if (prev_activeNBuffers != activeNBuffers)
+ {
+ /*
+ * Previous clock sweep position does not make sense if the size of
+ * the active buffer pool has changed.
+ *
+ * TODO: actually we have added a fix in
+ * StrategyAdjustNewBufAllocSize() to adjust the complete_passes so
+ * that we don't have to reset the saved info here, but still the
+ * Assert(StrategyDelta >= 0) below fails without resetting the saved
+ * info. Resetting the saved info means we loose the current position
+ * of the bgwriter and thus may have to unnecessarily scan already
+ * scanned buffers and the new allocations may have to find victim
+ * themselves. Needs further investigation.
+ */
+ saved_info_valid = false;
+ }
+
/* Report buffer alloc counts to pgstat */
PendingBgWriterStats.buf_alloc += recent_alloc;
@@ -3930,7 +3988,7 @@ BgBufferSync(WritebackContext *wb_context)
int32 passes_delta = strategy_passes - prev_strategy_passes;
strategy_delta = strategy_buf_id - prev_strategy_buf_id;
- strategy_delta += (long) passes_delta * NBuffers;
+ strategy_delta += (long) passes_delta * activeNBuffers;
if ((int32) (next_passes - strategy_passes) > 0)
{
@@ -3947,7 +4005,7 @@ BgBufferSync(WritebackContext *wb_context)
next_to_clean >= strategy_buf_id)
{
/* on same pass, but ahead or at least not behind */
- bufs_to_lap = NBuffers - (next_to_clean - strategy_buf_id);
+ bufs_to_lap = activeNBuffers - (next_to_clean - strategy_buf_id);
#ifdef BGW_DEBUG
elog(DEBUG2, "bgwriter ahead: bgw %u-%u strategy %u-%u delta=%ld lap=%d",
next_passes, next_to_clean,
@@ -3969,7 +4027,7 @@ BgBufferSync(WritebackContext *wb_context)
#endif
next_to_clean = strategy_buf_id;
next_passes = strategy_passes;
- bufs_to_lap = NBuffers;
+ bufs_to_lap = activeNBuffers;
}
/*
@@ -3996,12 +4054,13 @@ BgBufferSync(WritebackContext *wb_context)
strategy_delta = 0;
next_to_clean = strategy_buf_id;
next_passes = strategy_passes;
- bufs_to_lap = NBuffers;
+ bufs_to_lap = activeNBuffers;
}
/* Update saved info for next time */
prev_strategy_buf_id = strategy_buf_id;
prev_strategy_passes = strategy_passes;
+ prev_activeNBuffers = activeNBuffers;
saved_info_valid = true;
/*
@@ -4022,7 +4081,7 @@ BgBufferSync(WritebackContext *wb_context)
* strategy point and where we've scanned ahead to, based on the smoothed
* density estimate.
*/
- bufs_ahead = NBuffers - bufs_to_lap;
+ bufs_ahead = activeNBuffers - bufs_to_lap;
reusable_buffers_est = (float) bufs_ahead / smoothed_density;
/*
@@ -4060,7 +4119,7 @@ BgBufferSync(WritebackContext *wb_context)
* the BGW will be called during the scan_whole_pool time; slice the
* buffer pool into that many sections.
*/
- min_scan_buffers = (int) (NBuffers / (scan_whole_pool_milliseconds / BgWriterDelay));
+ min_scan_buffers = (int) (activeNBuffers / (scan_whole_pool_milliseconds / BgWriterDelay));
if (upcoming_alloc_est < (min_scan_buffers + reusable_buffers_est))
{
@@ -4082,13 +4141,24 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
- /* Execute the LRU scan */
+ /*
+ * Execute the LRU scan on the part of the buffer pool from where the new
+ * allocations will happen.
+ *
+ * Note that activeNBuffers may change during this loop if buffer pool
+ * gets resized concurrently. This may invalidate the num_to_scan count.
+ * If the buffer pool was shrunk this would make background writing more
+ * aggressive, which might be desired since the buffer pool is smaller. If
+ * the buffer pool was grown this would make the background writer not
+ * scan the newly added buffers, which should not have any dirty buffers
+ * yet.
+ */
while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
{
int sync_state = SyncOneBuffer(next_to_clean, true,
wb_context);
- if (++next_to_clean >= NBuffers)
+ if (++next_to_clean >= activeNBuffers)
{
next_to_clean = 0;
next_passes++;
@@ -8053,8 +8123,6 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
uint64 buf_state;
bool buffer_flushed;
- CHECK_FOR_INTERRUPTS();
-
buf_state = pg_atomic_read_u64(&desc->state);
if (!(buf_state & BM_VALID))
continue;
@@ -8071,6 +8139,13 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
if (buffer_flushed)
(*buffers_flushed)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8105,8 +8180,6 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
uint64 buf_state = pg_atomic_read_u64(&(desc->state));
bool buffer_flushed;
- CHECK_FOR_INTERRUPTS();
-
/* An unlocked precheck should be safe and saves some cycles. */
if ((buf_state & BM_VALID) == 0 ||
!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
@@ -8133,6 +8206,13 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
if (buffer_flushed)
(*buffers_flushed)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8250,8 +8330,6 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
uint64 buf_state = pg_atomic_read_u64(&(desc->state));
bool buffer_already_dirty;
- CHECK_FOR_INTERRUPTS();
-
/* An unlocked precheck should be safe and saves some cycles. */
if ((buf_state & BM_VALID) == 0 ||
!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
@@ -8277,6 +8355,13 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
(*buffers_already_dirty)++;
else
(*buffers_skipped)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -8304,8 +8389,6 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
uint64 buf_state;
bool buffer_already_dirty;
- CHECK_FOR_INTERRUPTS();
-
buf_state = pg_atomic_read_u64(&desc->state);
if (!(buf_state & BM_VALID))
continue;
@@ -8321,6 +8404,13 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
(*buffers_already_dirty)++;
else
(*buffers_skipped)++;
+
+ /*
+ * Checking interrupt before we are done with the current buffer might
+ * invalidate the buffer itself if a concurrent resizing shrinks
+ * buffer pool below the current buffer id.
+ */
+ CHECK_FOR_INTERRUPTS();
}
}
@@ -9017,3 +9107,57 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+
+/*
+ * When shrinking shared buffers pool, evict the buffers which will not be part
+ * of the shrunk buffer pool.
+ *
+ * If this function encounters a pinned buffer towards the end of the buffer
+ * pool, we would have evicted most of the buffers and yet rollback the resize
+ * operation. If we could find pinned buffers before evicting any, we could save
+ * wasted work and avoid performance impact because of evicted buffers. But
+ * there is no guarantee that a buffer won't be pinned after we check it. So we
+ * have to check for pinned buffers while evicting and rollback if we encounter
+ * any.
+ */
+bool
+EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
+{
+ bool result = true;
+
+ for (Buffer buf = targetNBuffers + 1; buf <= currentNBuffers; buf++)
+ {
+ BufferDesc *desc = GetBufferDescriptor(buf - 1);
+ uint64 buf_state;
+ bool buffer_flushed;
+
+ buf_state = pg_atomic_read_u64(&desc->state);
+
+ /*
+ * Nobody is expected to allocate new buffers while resizing is going
+ * on hence unlocked precheck should be safe and saves some cycles.
+ */
+ if (!(buf_state & BM_VALID))
+ continue;
+
+ ResourceOwnerEnlarge(CurrentResourceOwner);
+ ReservePrivateRefCountEntry();
+
+ LockBufHdr(desc);
+
+ /*
+ * Now that we have locked buffer descriptor, make sure that the
+ * buffer without valid data has been skipped above.
+ */
+ Assert(buf_state & BM_VALID);
+
+ if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
+ {
+ elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ result = false;
+ break;
+ }
+ }
+
+ return result;
+}
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 4d5ee52ddc0..e397c5ce47a 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -37,8 +37,9 @@ typedef struct
/*
* clock-sweep hand: index of next buffer to consider grabbing. Note that
* this isn't a concrete buffer - we only ever increase the value. So, to
- * get an actual buffer, it needs to be used modulo size of the buffer
- * pool.
+ * get an actual buffer, it needs to be used modulo the size of the buffer
+ *
+ * allocation area.
*/
pg_atomic_uint32 nextVictimBuffer;
@@ -101,11 +102,52 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
static void AddBufferToRing(BufferAccessStrategy strategy,
BufferDesc *buf);
+/*
+ * StrategyWrapAround - Wrap around the clock-sweep hand.
+ *
+ * `new_pos` is the new_pos position of the clock hand after wrap around.
+ * `num_passes` is the number of completed passes when wrapping around.
+ */
+static void
+StrategyWrapAround(uint32 new_pos, uint32 num_passes)
+{
+ bool success = false;
+
+ while (!success)
+ {
+ uint32 wrapped;
+
+ /*
+ * Acquire the spinlock while increasing completePasses. That allows
+ * other readers to read nextVictimBuffer and completePasses in a
+ * consistent manner which is required for StrategySyncStart(). In
+ * theory delaying the increment could lead to an overflow of
+ * nextVictimBuffers, but that's highly unlikely and wouldn't be
+ * particularly harmful.
+ */
+ SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
+
+ wrapped = new_pos % activeNBuffers;
+
+ success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
+ &new_pos, wrapped);
+ if (success)
+ StrategyControl->completePasses += num_passes;
+ SpinLockRelease(&StrategyControl->buffer_strategy_lock);
+ }
+}
+
/*
* ClockSweepTick - Helper routine for StrategyGetBuffer()
*
- * Move the clock hand one buffer ahead of its current position and return the
- * id of the buffer now under the hand.
+ * Move the clock hand one buffer ahead of its current position and return the id
+ * of the buffer now under the hand.
+ *
+ * We use the same value of activeNBuffers through out the function. Hence, even
+ * if the multiple backends end up wrapping around nextVictimBuffer using
+ * different activeNBuffers, they end up increasing completePasses incrementally
+ * and consistent to the respective victims. Use the latest activeNBuffers so as
+ * to be as consistent with the other allocators as possible.
*/
static inline uint32
ClockSweepTick(void)
@@ -119,13 +161,14 @@ ClockSweepTick(void)
*/
victim =
pg_atomic_fetch_add_u32(&StrategyControl->nextVictimBuffer, 1);
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
- if (victim >= NBuffers)
+ if (victim >= activeNBuffers)
{
uint32 originalVictim = victim;
/* always wrap what we look up in BufferDescriptors */
- victim = victim % NBuffers;
+ victim = victim % activeNBuffers;
/*
* If we're the one that just caused a wraparound, force
@@ -134,35 +177,9 @@ ClockSweepTick(void)
* value consisting of nextVictimBuffer and completePasses.
*/
if (victim == 0)
- {
- uint32 expected;
- uint32 wrapped;
- bool success = false;
-
- expected = originalVictim + 1;
-
- while (!success)
- {
- /*
- * Acquire the spinlock while increasing completePasses. That
- * allows other readers to read nextVictimBuffer and
- * completePasses in a consistent manner which is required for
- * StrategySyncStart(). In theory delaying the increment
- * could lead to an overflow of nextVictimBuffers, but that's
- * highly unlikely and wouldn't be particularly harmful.
- */
- SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
-
- wrapped = expected % NBuffers;
-
- success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer,
- &expected, wrapped);
- if (success)
- StrategyControl->completePasses++;
- SpinLockRelease(&StrategyControl->buffer_strategy_lock);
- }
- }
+ StrategyWrapAround(originalVictim + 1, 1);
}
+
return victim;
}
@@ -180,6 +197,11 @@ ClockSweepTick(void)
*
* The buffer is pinned and marked as owned, using TrackNewBufferPin(),
* before returning.
+ *
+ * We do not process a ProcSignalBarrier between choosing a victim and pinning
+ * it, so the buffer will remain valid even if the buffer pool is shrunk. For
+ * better safety we may want to disable interrupt handling explicitly during
+ * this time.
*/
BufferDesc *
StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
@@ -238,7 +260,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
pg_atomic_fetch_add_u32(&StrategyControl->numBufferAllocs, 1);
/* Use the "clock sweep" algorithm to find a free buffer */
- trycounter = NBuffers;
+ trycounter = activeNBuffers;
for (;;)
{
uint64 old_buf_state;
@@ -291,7 +313,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r
if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
local_buf_state))
{
- trycounter = NBuffers;
+ trycounter = activeNBuffers;
break;
}
}
@@ -334,9 +356,15 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
uint32 nextVictimBuffer;
int result;
+ /*
+ * Update backend local activeNBuffers for the same reason as in
+ * StrategyGetBuffer.
+ */
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
SpinLockAcquire(&StrategyControl->buffer_strategy_lock);
nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
- result = nextVictimBuffer % NBuffers;
+ result = nextVictimBuffer % activeNBuffers;
if (complete_passes)
{
@@ -346,7 +374,7 @@ StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc)
* Additionally add the number of wraparounds that happened before
* completePasses could be incremented. C.f. ClockSweepTick().
*/
- *complete_passes += nextVictimBuffer / NBuffers;
+ *complete_passes += nextVictimBuffer / activeNBuffers;
}
if (num_buf_alloc)
@@ -392,6 +420,29 @@ StrategyCtlShmemRequest(void *arg)
);
}
+/*
+ * StrategyAdjustNewBufAllocSize
+ *
+ * Adjust clock hand when resizing the buffer pool.
+ */
+void
+StrategyAdjustNewBufAllocSize(void)
+{
+ int num_passes;
+ uint32 nextVictimBuffer;
+
+ /* Should be called only when resizing is in progress. */
+ Assert(pg_atomic_read_u32(&BufferControl->resizer_pid) != 0);
+
+ activeNBuffers = pg_atomic_read_u32(&BufferControl->activeNBuffers);
+
+ /* Consistently wrap around the clock sweep hand, if necessary. */
+ nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer);
+ num_passes = nextVictimBuffer / activeNBuffers;
+ if (num_passes > 0)
+ StrategyWrapAround(nextVictimBuffer, num_passes);
+}
+
/*
* StrategyCtlShmemInit -- initialize the buffer cache replacement strategy.
*/
@@ -634,12 +685,27 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
strategy->current = 0;
/*
- * If the slot hasn't been filled yet, tell the caller to allocate a new
- * buffer with the normal allocation strategy. He will then fill this
- * slot by calling AddBufferToRing with the new buffer.
+ * If the slot hasn't been filled yet or the buffer in the slot is outside
+ * the buffer allocation area tell the caller to allocate a new buffer
+ * with the normal allocation strategy. He will then fill this slot by
+ * calling AddBufferToRing with the new buffer. Usually the buffers in the
+ * ring will be within the buffer allocation area, but if the buffer pool
+ * has been shrunk since the last time the ring was filled, some of the
+ * buffers in the ring may be outside the new buffer allocation area.
+ *
+ * TODO: buffer ids in the ring will never be greater than the size of
+ * buffer pool, except maybe the first time ring is accessed after
+ * shrinking the buffer pool. Checking the upper bound on buffer id always
+ * may mask a bug bugs that introduces buffer ids higher than the size of
+ * buffer pool in the ring. But performing that check only once after
+ * shrinking seems impossible. The BufferAccessStrategy objects are not
+ * accessible outside the ScanState. Hence we can not purge the buffers
+ * while evicting the buffers. After the resizing is finished, it's not
+ * possible to notice when we touch the first of those objects and the
+ * last of objects. See if this can fixed.
*/
bufnum = strategy->buffers[strategy->current];
- if (bufnum == InvalidBuffer)
+ if (bufnum == InvalidBuffer || bufnum > activeNBuffers)
return NULL;
buf = GetBufferDescriptor(bufnum - 1);
diff --git a/src/backend/storage/buffer/meson.build b/src/backend/storage/buffer/meson.build
index ed84bf08971..f219e29d5ef 100644
--- a/src/backend/storage/buffer/meson.build
+++ b/src/backend/storage/buffer/meson.build
@@ -6,4 +6,5 @@ backend_sources += files(
'bufmgr.c',
'freelist.c',
'localbuf.c',
+ 'buf_resize.c',
)
diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c
index 21a77f98c1d..f97215dae9c 100644
--- a/src/backend/storage/ipc/procsignal.c
+++ b/src/backend/storage/ipc/procsignal.c
@@ -28,6 +28,7 @@
#include "replication/logicalworker.h"
#include "replication/slotsync.h"
#include "replication/walsender.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/latch.h"
@@ -598,6 +599,15 @@ ProcessProcSignalBarrier(void)
case PROCSIGNAL_BARRIER_CHECKSUM_OFF:
processed = AbsorbDataChecksumsBarrier(type);
break;
+ case PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC:
+ processed = ProcessBarrierNewBufferAlloc();
+ break;
+ case PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE:
+ processed = ProcessBarrierBufferPoolResize();
+ break;
+ case PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE:
+ processed = ProcessBarrierBufferPoolSize();
+ break;
}
/*
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 780cdb5e9ff..8332bc0b252 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -43,6 +43,7 @@
#include "postmaster/autovacuum.h"
#include "replication/slotsync.h"
#include "replication/syncrep.h"
+#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
@@ -583,6 +584,21 @@ InitProcess(void)
* the reasons mentioned there.
*/
ShmemReprotectResizableStructs();
+
+#ifndef EXEC_BACKEND
+
+ /*
+ * Pick up the current buffer pool size from shared memory. Fork'ed
+ * backends would otherwise inherit the postmaster's potentially stale
+ * values. EXEC_BACKEND children do this via the buffer manager attach
+ * callback above.
+ *
+ * We also call this function after ProcSignalInit() for the reasons
+ * specified there, but we need it here so that InitBufferManagerAccess()
+ * can use the current buffer pool size.
+ */
+ BufferManagerInitProc();
+#endif
}
/*
@@ -772,6 +788,21 @@ InitAuxiliaryProcess(void)
* the reasons mentioned there.
*/
ShmemReprotectResizableStructs();
+
+#ifndef EXEC_BACKEND
+
+ /*
+ * Pick up the current buffer pool size from shared memory. Fork'ed
+ * backends would otherwise inherit the postmaster's potentially stale
+ * values. EXEC_BACKEND children do this via the buffer manager attach
+ * callback above.
+ *
+ * We also call this function after ProcSignalInit() for the reasons
+ * specifid there, but we need it here so that InitBufferManagerAccess()
+ * can use the current buffer pool size.
+ */
+ BufferManagerInitProc();
+#endif
}
/*
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index b6bdfe213fe..2586b7087ac 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -4288,6 +4288,9 @@ PostgresSingleUserMain(int argc, char *argv[],
/* Initialize size of fast-path lock cache. */
InitializeFastPathLocks();
+ /* Initialize MaxNBuffers for buffer pool resizing. */
+ InitializeMaxNBuffers();
+
/*
* Also call any legacy shmem request hooks that might'be been installed
* by preloaded libraries.
diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c
index ccf845e87b9..4c5cb3c3908 100644
--- a/src/backend/utils/init/globals.c
+++ b/src/backend/utils/init/globals.c
@@ -142,6 +142,8 @@ int max_parallel_maintenance_workers = 2;
* register background workers.
*/
int NBuffersGUC = 16384;
+bool finalMaxNBuffers = false;
+int MaxNBuffers = 0;
int MaxConnections = 100;
int max_worker_processes = 8;
int max_parallel_workers = 8;
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 815b865aa38..c129d3a8a70 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -610,6 +610,55 @@ InitializeFastPathLocks(void)
pg_nextpower2_32(FastPathLockGroupsPerBackend));
}
+/*
+ * Initialize MaxNBuffers variable with validation.
+ *
+ * This must be called after GUCs have been loaded but before shared memory size
+ * is determined.
+ *
+ * Since MaxNBuffers limits the size of the buffer pool, it must be at least as
+ * much as NBuffersGUC. If MaxNBuffers is 0 (default), set it to
+ * NBuffersGUC. Otherwise, validate that MaxNBuffers is not less than
+ * NBuffersGUC.
+ */
+void
+InitializeMaxNBuffers(void)
+{
+ if (MaxNBuffers == 0) /* default/boot value */
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", NBuffersGUC);
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+
+ /*
+ * We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
+ * However, if the DBA explicitly set max_shared_buffers = 0 in the
+ * config file, then PGC_S_DYNAMIC_DEFAULT will fail to override that
+ * and we must force the matter with PGC_S_OVERRIDE.
+ */
+ if (MaxNBuffers == 0) /* failed to apply it? */
+ SetConfigOption("max_shared_buffers", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+ else
+ {
+ if (MaxNBuffers < NBuffersGUC)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("max_shared_buffers (%d) cannot be less than current shared_buffers (%d)",
+ MaxNBuffers, NBuffersGUC),
+ errhint("Increase max_shared_buffers or decrease shared_buffers.")));
+ }
+ }
+
+ Assert(MaxNBuffers > 0);
+ Assert(!finalMaxNBuffers);
+ finalMaxNBuffers = true;
+}
+
/*
* Early initialization of a backend (either standalone or under postmaster).
* This happens even before InitPostgres.
@@ -760,32 +809,38 @@ InitPostgres(const char *in_dbname, Oid dboid,
SharedInvalBackendInit(false);
/*
- * Prevent consuming interrupts between setting ProcSignalInit and setting
- * the initial local data checksum value. If a barrier is emitted, and
- * absorbed, before local cached state is initialized the state transition
- * can be invalid.
+ * Prevent consuming interrupts between ProcSignalInit() and the
+ * initialization of state that is kept in sync with shared memory via
+ * procsignal-based barriers (currently the data_checksum_version cache
+ * and the local NBuffers/activeNBuffers cache). If a barrier is emitted,
+ * and absorbed, before that local cached state is initialized the state
+ * transition can be invalid.
*/
HOLD_INTERRUPTS();
ProcSignalInit(MyCancelKey, MyCancelKeyLength);
/*
- * Initialize a local cache of the data_checksum_version, to be updated by
- * the procsignal-based barriers.
+ * Initialize the per-backend caches that are kept in sync via
+ * procsignal-based barriers: currently the local data_checksum_version
+ * and the local copy of the buffer pool size (NBuffers/activeNBuffers).
*
- * This intentionally happens after initializing the procsignal, otherwise
- * we might miss a state change. This means we can get a barrier for the
- * state we've just initialized.
+ * These initializations intentionally happen after ProcSignalInit(),
+ * otherwise we might miss a state change. This means we may also receive
+ * a barrier for the state we've just initialized.
*
* The postmaster (which is what gets forked into the new child process)
* does not handle barriers, therefore it may not have the current value
- * of LocalDataChecksumState value (it'll have the value read from the
- * control file, which may be arbitrarily old).
+ * of LocalDataChecksumState (it'll have the value read from the control
+ * file, which may be arbitrarily old) or NBuffers/activeNBuffers (which
+ * may have been changed by an online resize after the postmaster
+ * started).
*
* NB: Even if the postmaster handled barriers, the value might still be
* stale, as it might have changed after this process forked.
*/
InitLocalDataChecksumState();
+ BufferManagerInitProc();
/*
* Refresh per-backend protections for resizable shmem structures. Usually
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 1a5a168bc1a..5f1dd338f56 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -2631,7 +2631,7 @@ convert_to_base_unit(double value, const char *unit,
* the value without loss. For example, if the base unit is GUC_UNIT_KB, 1024
* is converted to 1 MB, but 1025 is represented as 1025 kB.
*/
-static void
+void
convert_int_from_base_unit(int64 base_value, int base_unit,
int64 *value, const char **unit)
{
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index 7a9bee3ed3f..f5aa96c544d 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2146,6 +2146,15 @@
max => 'MAX_BACKENDS /* XXX? */',
},
+{ name => "max_shared_buffers", type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+ short_desc => 'Sets the upper limit for the shared_buffers value.',
+ flags => 'GUC_UNIT_BLOCKS',
+ variable => 'MaxNBuffers',
+ boot_val => '0',
+ min => '0',
+ max => 'INT_MAX / 2',
+},
+
{ name => 'max_slot_wal_keep_size', type => 'int', context => 'PGC_SIGHUP', group => 'REPLICATION_SENDING',
short_desc => 'Sets the maximum WAL size that can be reserved by replication slots.',
long_desc => 'Replication slots will be marked as failed, and segments released for deletion or recycling, if this much space is occupied by WAL on disk. -1 means no maximum.',
@@ -2720,13 +2729,15 @@
# We sometimes multiply the number of shared buffers by two without
# checking for overflow, so we mustn't allow more than INT_MAX / 2.
-{ name => 'shared_buffers', type => 'int', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM',
+{ name => 'shared_buffers', type => 'int', context => 'PGC_SIGHUP', group => 'RESOURCES_MEM',
short_desc => 'Sets the number of shared memory buffers used by the server.',
flags => 'GUC_UNIT_BLOCKS',
variable => 'NBuffersGUC',
boot_val => '16384',
- min => '16',
+ min => 'MIN_NUM_BUFFERS',
max => 'INT_MAX / 2',
+ check_hook => 'check_shared_buffers',
+ show_hook => 'show_shared_buffers',
},
{ name => 'shared_memory_initial_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS',
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 712172760b3..0487576ffd7 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12715,4 +12715,19 @@
proname => 'hashoid8extended', prorettype => 'int8',
proargtypes => 'oid8 int8', prosrc => 'hashoid8extended' },
+{ oid => '9999', descr => 'resize shared buffers according to the value of GUC `shared_buffers`',
+ proname => 'pg_resize_shared_buffers',
+ provolatile => 'v',
+ prorettype => 'bool',
+ proargtypes => '',
+ prosrc => 'pg_resize_shared_buffers'},
+{ oid => '9998', descr => 'report shared buffer pool resize status',
+ proname => 'pg_get_buffer_resize_status',
+ provolatile => 'v',
+ prorettype => 'record',
+ proargtypes => '',
+ proallargtypes => '{int4,int4,int4,int4}',
+ proargmodes => '{o,o,o,o}',
+ proargnames => '{active_nbuffers,current_nbuffers,target_nbuffers,resizer_pid}',
+ prosrc => 'pg_get_buffer_resize_status'},
]
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index e2df51d275a..772f92057e8 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -176,6 +176,8 @@ extern PGDLLIMPORT char *DataDir;
extern PGDLLIMPORT int data_directory_mode;
extern PGDLLIMPORT int NBuffersGUC;
+extern PGDLLIMPORT bool finalMaxNBuffers;
+extern PGDLLIMPORT int MaxNBuffers;
extern PGDLLIMPORT int MaxBackends;
extern PGDLLIMPORT int MaxConnections;
extern PGDLLIMPORT int max_worker_processes;
@@ -515,6 +517,7 @@ extern PGDLLIMPORT ProcessingMode Mode;
extern void pg_split_opts(char **argv, int *argcp, const char *optstr);
extern void InitializeMaxBackends(void);
extern void InitializeFastPathLocks(void);
+extern void InitializeMaxNBuffers(void);
extern void InitPostgres(const char *in_dbname, Oid dboid,
const char *username, Oid useroid,
uint32 flags,
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 678065b3de2..a4971cebde4 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -264,6 +264,33 @@ BufMappingPartitionLockByIndex(uint32 index)
return &MainLWLockArray[BUFFER_MAPPING_LWLOCK_OFFSET + index].lock;
}
+/*
+ * BufferControl -- shared area controlling buffer pool
+ *
+ * This structure stores information about the size of the buffer pool and
+ * whether it is being resized.
+ */
+typedef struct BufferControlBlock
+{
+ /*
+ * size of the part of the buffer pool from where buffers are being
+ * allocated to new requests.
+ */
+ pg_atomic_uint32 activeNBuffers;
+
+ /* current size of the buffer pool, in number of buffers */
+ pg_atomic_uint32 currentNBuffers;
+
+ /* target size of the buffer pool, in number of buffers */
+ pg_atomic_uint32 targetNBuffers;
+
+ /*
+ * PID of the backend currently performing a resize, or 0 when no resize
+ * is in progress. Also acts as a lock prohibiting concurrent resizes.
+ */
+ pg_atomic_uint32 resizer_pid;
+} BufferControlBlock;
+
/*
* BufferDesc -- shared descriptor/state data for a single shared buffer.
*
@@ -411,6 +438,7 @@ typedef struct WritebackContext
} WritebackContext;
/* in buf_init.c */
+extern PGDLLIMPORT BufferControlBlock *BufferControl;
extern PGDLLIMPORT BufferDescPadded *BufferDescriptors;
extern PGDLLIMPORT ConditionVariableMinimallyPadded *BufferIOCVArray;
extern PGDLLIMPORT WritebackContext BackendWritebackContext;
@@ -422,9 +450,26 @@ extern PGDLLIMPORT BufferDesc *LocalBufferDescriptors;
static inline BufferDesc *
GetBufferDescriptor(int id)
{
+ BufferDesc *bdesc;
+
Assert(id >= 0 && id < NBuffers);
- return &(BufferDescriptors[id]).bufferdesc;
+ bdesc = &(BufferDescriptors[id]).bufferdesc;
+
+ /*
+ * TODO: This assertion was proposed in
+ * https://www.postgresql.org/message-id/CAExHW5uzRMYVZsXXS3HXXT0fG_sNrpUhUqwP4NorhaCqH9JDhA@mail.gmail.com,
+ * but was ultimately removed since there was no adequate reason to keep
+ * it in the code without shared buffer resizing. With resizing we may
+ * write and rewrite parts of the buffer descriptor array. So it's better
+ * to make sure that the buffer descriptor is initialized correctly. For
+ * now just make sure that the id in the buffer descriptor is the same as
+ * the id used to access it. Later we may want to expand the assertion to
+ * check the BufferDesc invariants or remove this assertion.
+ */
+ Assert(bdesc->buf_id == id);
+
+ return bdesc;
}
static inline BufferDesc *
@@ -594,6 +639,7 @@ extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
extern int StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc);
extern void StrategyNotifyBgWriter(int bgwprocno);
+extern void StrategyAdjustNewBufAllocSize(void);
/* buf_table.c */
extern uint32 BufTableHashCode(BufferTag *tagPtr);
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index f1f6e601f51..187e9b29a9e 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -14,6 +14,7 @@
#ifndef BUFMGR_H
#define BUFMGR_H
+#include "fmgr.h"
#include "port/pg_iovec.h"
#include "storage/aio_types.h"
#include "storage/block.h"
@@ -159,6 +160,7 @@ typedef struct ReadBuffersOperation ReadBuffersOperation;
typedef struct WritebackContext WritebackContext;
/* in globals.c ... this duplicates miscadmin.h */
+#define MIN_NUM_BUFFERS 16
extern PGDLLIMPORT int NBuffersGUC;
/* in bufmgr.c */
@@ -167,6 +169,7 @@ extern PGDLLIMPORT int bgwriter_lru_maxpages;
extern PGDLLIMPORT double bgwriter_lru_multiplier;
extern PGDLLIMPORT bool track_io_timing;
extern PGDLLIMPORT int NBuffers;
+extern PGDLLIMPORT int activeNBuffers;
#define DEFAULT_EFFECTIVE_IO_CONCURRENCY 16
#define DEFAULT_MAINTENANCE_IO_CONCURRENCY 16
@@ -371,6 +374,12 @@ extern void MarkDirtyRelUnpinnedBuffers(Relation rel,
extern void MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
int32 *buffers_already_dirty,
int32 *buffers_skipped);
+extern bool EvictExtraBuffers(int targetNBuffers, int currentNBuffers);
+
+/* in buf_init.c */
+extern bool BufferManagerShmemResize(int currentNBuffers, int targetNBuffers);
+extern void BufferManagerShmemProtect(void);
+extern void BufferManagerInitProc(void);
/* in localbuf.c */
extern void AtProcExit_LocalBuffers(void);
@@ -473,4 +482,10 @@ BufferGetPage(Buffer buffer)
#endif /* FRONTEND */
+/* buf_resize.c */
+extern Datum pg_resize_shared_buffers(PG_FUNCTION_ARGS);
+extern bool ProcessBarrierNewBufferAlloc(void);
+extern bool ProcessBarrierBufferPoolResize(void);
+extern bool ProcessBarrierBufferPoolSize(void);
+
#endif /* BUFMGR_H */
diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h
index aaa158bfd66..78dda9c8d21 100644
--- a/src/include/storage/procsignal.h
+++ b/src/include/storage/procsignal.h
@@ -54,6 +54,11 @@ typedef enum
PROCSIGNAL_BARRIER_CHECKSUM_INPROGRESS_ON,
PROCSIGNAL_BARRIER_CHECKSUM_INPROGRESS_OFF,
PROCSIGNAL_BARRIER_CHECKSUM_ON,
+ PROCSIGNAL_BARRIER_NEW_BUFFER_ALLOC, /* New buffer allocation pool size
+ * changed */
+ PROCSIGNAL_BARRIER_BUFFER_POOL_RESIZE, /* Buffer pool shared structures
+ * resized */
+ PROCSIGNAL_BARRIER_BUFFER_POOL_SIZE, /* Buffer pool size updated */
} ProcSignalBarrierType;
/*
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 2a6e2ed18b3..2e3dc919018 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -462,6 +462,8 @@ extern config_handle *get_config_handle(const char *name);
extern void AlterSystemSetConfigFile(AlterSystemStmt *altersysstmt);
extern char *GetConfigOptionByName(const char *name, const char **varname,
bool missing_ok);
+extern void convert_int_from_base_unit(int64 base_value, int base_unit,
+ int64 *value, const char **unit);
extern void TransformGUCArray(ArrayType *array, List **names,
List **values);
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index df048517a0f..7e549a36530 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -181,4 +181,6 @@ extern void assign_synchronized_standby_slots(const char *newval, void *extra);
extern bool check_log_min_messages(char **newval, void **extra, GucSource source);
extern void assign_log_min_messages(const char *newval, void *extra);
+extern const char *show_shared_buffers(bool use_units);
+extern bool check_shared_buffers(int *newval, void **extra, GucSource source);
#endif /* GUC_HOOKS_H */
diff --git a/src/test/Makefile b/src/test/Makefile
index 3eb0a06abb4..7a0d74086c1 100644
--- a/src/test/Makefile
+++ b/src/test/Makefile
@@ -20,7 +20,8 @@ SUBDIRS = \
postmaster \
recovery \
regress \
- subscription
+ subscription \
+ buffermgr
ifeq ($(with_icu),yes)
SUBDIRS += icu
diff --git a/src/test/README b/src/test/README
index afdc7676519..77f11607ff7 100644
--- a/src/test/README
+++ b/src/test/README
@@ -15,6 +15,9 @@ examples/
Demonstration programs for libpq that double as regression tests via
"make check"
+buffermgr/
+ Tests for resizing buffer pool without restarting the server
+
isolation/
Tests for concurrent behavior at the SQL level
diff --git a/src/test/buffermgr/Makefile b/src/test/buffermgr/Makefile
new file mode 100644
index 00000000000..24c245c900a
--- /dev/null
+++ b/src/test/buffermgr/Makefile
@@ -0,0 +1,35 @@
+#-------------------------------------------------------------------------
+#
+# Makefile for src/test/buffermgr
+#
+# Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+# Portions Copyright (c) 1994, Regents of the University of California
+#
+# src/test/buffermgr/Makefile
+#
+#-------------------------------------------------------------------------
+
+EXTRA_INSTALL = contrib/pg_buffercache \
+ src/test/modules/injection_points \
+ src/test/modules/test_shmem
+
+REGRESS = buffer_resize
+
+# Custom configuration for buffer manager tests
+TEMP_CONFIG = $(srcdir)/buffermgr_test.conf
+
+export enable_injection_points
+
+subdir = src/test/buffermgr
+top_builddir = ../../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+
+check:
+ $(prove_check)
+
+installcheck:
+ $(prove_installcheck)
+
+clean distclean:
+ rm -rf tmp_check
diff --git a/src/test/buffermgr/README b/src/test/buffermgr/README
new file mode 100644
index 00000000000..c375ad80989
--- /dev/null
+++ b/src/test/buffermgr/README
@@ -0,0 +1,26 @@
+src/test/buffermgr/README
+
+Regression tests for buffer manager
+===================================
+
+This directory contains a test suite for resizing buffer manager without restarting the server.
+
+
+Running the tests
+=================
+
+NOTE: You must have given the --enable-tap-tests argument to configure.
+
+Run
+ make check
+or
+ make installcheck
+You can use "make installcheck" if you previously did "make install".
+In that case, the code in the installation tree is tested. With
+"make check", a temporary installation tree is built from the current
+sources and then tested.
+
+Either way, this test initializes, starts, and stops a test Postgres
+cluster.
+
+See src/test/perl/README for more info about running these tests.
diff --git a/src/test/buffermgr/buffermgr_test.conf b/src/test/buffermgr/buffermgr_test.conf
new file mode 100644
index 00000000000..a15f3e442a5
--- /dev/null
+++ b/src/test/buffermgr/buffermgr_test.conf
@@ -0,0 +1,11 @@
+# Configuration for buffer manager regression tests
+
+# Even if max_shared_buffers is set multiple times only the last one is used to
+# as the limit on shared_buffers.
+max_shared_buffers = 128kB
+# Set initial shared_buffers as expected by test
+shared_buffers = 128MB
+# Set a larger value for max_shared_buffers to allow testing resize operations
+max_shared_buffers = 300MB
+# Turn huge pages off, since that affects the size of memory segments
+huge_pages = off
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
new file mode 100644
index 00000000000..1dc3efd830a
--- /dev/null
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -0,0 +1,290 @@
+-- Test buffer pool resizing and shared memory allocation tracking This test
+-- resizes the buffer pool multiple times and monitors shared memory allocations
+-- related to buffer management
+-- TODOs
+--
+-- 1. The test sets shared_buffers values in MBs. Instead it could use values in
+-- kBs so that the test runs on very small machines.
+--
+-- 2. The size, minimum_size and maximum_size columns in pg_shmem_allocations
+-- for "Buffer Blocks" should be same as the value of GUC shared_buffers. We
+-- should test that.
+--
+-- 3. We should make sure that when the shared_buffers value is increased, the
+-- size and allocated_size for all buffer related shared memory allocations
+-- increases and when the shared_buffers value is decreased, the size and
+-- allocated_size for all buffer related shared memory allocations decreases
+-- proportionately.
+--
+-- 4. allocated_size for allocations should be greater than or equal to size for
+-- all buffer related shared memory allocations. Similarly reserved_space should
+-- be greater than or equal to maximum_size for all buffer related shared memory
+-- allocations. We should test these conditions as well.
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+-- Load test_shmem for test_shmem_pagesize().
+CREATE EXTENSION IF NOT EXISTS test_shmem;
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, size,
+ allocated_size >= size AS alloc_size_cmp,
+ allocated_size - size < 2 * test_shmem_pagesize() AS alloc_size_diff,
+ minimum_size, maximum_size, reserved_space
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 134217728 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1048576 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 262144 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 134217728 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1048576 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 262144 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16384
+(1 row)
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 128MB (pending: 64MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 64MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 67108864 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 524288 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 131072 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 8192
+(1 row)
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+-----------------------
+ 64MB (pending: 256MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 256MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 268435456 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 2097152 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 524288 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 32768
+(1 row)
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 256MB (pending: 100MB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 100MB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+-----------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 104857600 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 819200 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 204800 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 12800
+(1 row)
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+ shared_buffers
+------------------------
+ 100MB (pending: 128kB)
+(1 row)
+
+SELECT pg_resize_shared_buffers();
+ pg_resize_shared_buffers
+--------------------------
+ t
+(1 row)
+
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SELECT * FROM buffer_allocations;
+ name | size | alloc_size_cmp | alloc_size_diff | minimum_size | maximum_size | reserved_space
+-------------------------------+--------+----------------+-----------------+--------------+--------------+----------------
+ Buffer Blocks | 131072 | t | t | 131072 | 314572800 | 314574336
+ Buffer Descriptors | 1024 | t | t | 1024 | 2457600 | 2457600
+ Buffer IO Condition Variables | 256 | t | t | 256 | 614400 | 614400
+ Checkpoint BufferIds | 768000 | t | t | 768000 | 768000 | 768120
+(4 rows)
+
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+ buffer_count
+--------------
+ 16
+(1 row)
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+ERROR: invalid value for parameter "shared_buffers": 51200
+DETAIL: "shared_buffers" must be less than "max_shared_buffers".
+SELECT pg_reload_conf();
+ pg_reload_conf
+----------------
+ t
+(1 row)
+
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+ shared_buffers
+----------------
+ 128kB
+(1 row)
+
+SHOW max_shared_buffers;
+ max_shared_buffers
+--------------------
+ 300MB
+(1 row)
+
+-- TODO: Test that a non-superuser can not invoke pg_resize_shared_buffers()
+-- function.
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
new file mode 100644
index 00000000000..7a6d5e29f8d
--- /dev/null
+++ b/src/test/buffermgr/meson.build
@@ -0,0 +1,25 @@
+# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+
+tests += {
+ 'name': 'buffermgr',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'buffer_resize',
+ ],
+ 'regress_args': ['--temp-config', files('buffermgr_test.conf')],
+ },
+ 'tap': {
+ 'env': {
+ 'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
+ },
+ 'tests': [
+ 't/001_resize_buffer.pl',
+ 't/003_resize_fault_tolerance.pl',
+ 't/004_client_join_buffer_resize.pl',
+ 't/005_resize_failures.pl',
+ 't/006_resize_with_syslogger.pl',
+ ],
+ },
+}
diff --git a/src/test/buffermgr/sql/buffer_resize.sql b/src/test/buffermgr/sql/buffer_resize.sql
new file mode 100644
index 00000000000..4219e33a805
--- /dev/null
+++ b/src/test/buffermgr/sql/buffer_resize.sql
@@ -0,0 +1,106 @@
+-- Test buffer pool resizing and shared memory allocation tracking This test
+-- resizes the buffer pool multiple times and monitors shared memory allocations
+-- related to buffer management
+
+-- TODOs
+--
+-- 1. The test sets shared_buffers values in MBs. Instead it could use values in
+-- kBs so that the test runs on very small machines.
+--
+-- 2. The size, minimum_size and maximum_size columns in pg_shmem_allocations
+-- for "Buffer Blocks" should be same as the value of GUC shared_buffers. We
+-- should test that.
+--
+-- 3. We should make sure that when the shared_buffers value is increased, the
+-- size and allocated_size for all buffer related shared memory allocations
+-- increases and when the shared_buffers value is decreased, the size and
+-- allocated_size for all buffer related shared memory allocations decreases
+-- proportionately.
+--
+-- 4. allocated_size for allocations should be greater than or equal to size for
+-- all buffer related shared memory allocations. Similarly reserved_space should
+-- be greater than or equal to maximum_size for all buffer related shared memory
+-- allocations. We should test these conditions as well.
+
+CREATE EXTENSION IF NOT EXISTS pg_buffercache;
+
+-- Load test_shmem for test_shmem_pagesize().
+CREATE EXTENSION IF NOT EXISTS test_shmem;
+
+-- Create a view for buffer-related shared memory allocations
+CREATE VIEW buffer_allocations AS
+SELECT name, size,
+ allocated_size >= size AS alloc_size_cmp,
+ allocated_size - size < 2 * test_shmem_pagesize() AS alloc_size_diff,
+ minimum_size, maximum_size, reserved_space
+FROM pg_shmem_allocations
+WHERE name IN ('Buffer Blocks', 'Buffer Descriptors', 'Buffer IO Condition Variables',
+ 'Checkpoint BufferIds')
+ORDER BY name;
+
+-- Test 1: Default shared_buffers
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+-- Calling pg_resize_shared_buffers() without changing shared_buffers should be a no-op.
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 2: Set to 64MB
+ALTER SYSTEM SET shared_buffers = '64MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 3: Set to 256MB
+ALTER SYSTEM SET shared_buffers = '256MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 4: Set to 100MB (non-power-of-two)
+ALTER SYSTEM SET shared_buffers = '100MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 5: Set to minimum 128kB
+ALTER SYSTEM SET shared_buffers = '128kB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+SHOW shared_buffers;
+SELECT pg_resize_shared_buffers();
+SHOW shared_buffers;
+SELECT * FROM buffer_allocations;
+SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
+
+-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
+ALTER SYSTEM SET shared_buffers = '400MB';
+SELECT pg_reload_conf();
+-- reconnect to ensure new setting is loaded
+\c
+-- This should show the old value since the configuration was rejected
+SHOW shared_buffers;
+SHOW max_shared_buffers;
+
+-- TODO: Test that a non-superuser can not invoke pg_resize_shared_buffers()
+-- function.
diff --git a/src/test/buffermgr/t/001_resize_buffer.pl b/src/test/buffermgr/t/001_resize_buffer.pl
new file mode 100644
index 00000000000..fb5a42be26a
--- /dev/null
+++ b/src/test/buffermgr/t/001_resize_buffer.pl
@@ -0,0 +1,193 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Minimal test testing shared_buffer resizing under load
+
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Function to check if pgbench is still running.
+#
+# Relying on IPC::Run's pumpable status to check if pgbench is still running has
+# been proven unreliable. Instead we rely on existence of pgbench processes in
+# pg_stat_activity. Since we use -C with pgbench, there can be a non-zero
+# chance that no pgbench process is running even thought pgbench is running. But
+# that's a very rare possibility that can be ignored.
+sub pgbench_processes_active
+{
+ my ($node, $application_name) = @_;
+
+ my $result = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_stat_activity WHERE application_name = '$application_name';");
+ return int($result) > 0;
+}
+
+my $resize_sql_func_def = q{
+create or replace function pg_resize_shared_buffers_sql(new_size int, out num_tries int) returns int as $$
+declare
+ success boolean := false;
+ tries int := 0;
+ cur_setting text;
+ pending_pattern text;
+ target text := new_size::text;
+begin
+ -- Wait until pg_settings reports the new value as pending,
+ -- i.e. "<old value> (pending: <new value>)".
+ pending_pattern := '%(pending: ' || target || ')%';
+ loop
+ select setting into cur_setting
+ from pg_settings where name = 'shared_buffers';
+ exit when cur_setting like pending_pattern or cur_setting = target;
+ perform pg_sleep(0.1);
+ raise notice 'Current setting: %', cur_setting;
+ end loop;
+
+ -- pg_resize_shared_buffers() returns true on success; retry until it succeeds.
+ while not success loop
+ tries := tries + 1;
+ select pg_resize_shared_buffers() into success;
+ if not success then
+ perform pg_sleep(0.1);
+ end if;
+ raise notice 'pg_resize_shared_buffers() attempt %: success = %', tries, success;
+ end loop;
+
+ -- Confirm the new value is in effect (no longer pending).
+ select setting into cur_setting
+ from pg_settings where name = 'shared_buffers';
+ if cur_setting <> target then
+ raise exception 'shared_buffers resize did not take effect: expected %, got %',
+ target, cur_setting;
+ end if;
+
+ num_tries := tries;
+ return;
+end;
+$$ language plpgsql;
+};
+
+# Function to resize buffer pool and verify the change.
+sub apply_and_verify_buffer_change
+{
+ my ($node, $new_size) = @_;
+
+ # Use the new pg_resize_shared_buffers() interface which handles everything synchronously
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$new_size'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ $node->safe_psql('postgres', "SELECT pg_resize_shared_buffers_sql($new_size)");
+
+ # Any failure in resizing the buffer pool will cause the test to timeout. So
+ # if we reach here, the resize was successful. Just declare it as a
+ # successful test so that we can see progress in the test output.
+ ok(1, "Buffer pool resized to $new_size");
+}
+
+my @buffer_sizes = (128, 28, 16 * 1024, 32 * 1024, 1024, 512, 16, 24, 256, 128 * 1024, 16 * 1204);
+
+# Initialize a cluster and start pgbench in the background for concurrent load.
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Permit resizing up to 1GB for this test and let the server start with 128MB.
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = } . (sort { $b <=> $a } @buffer_sizes)[0] . qq{
+shared_buffers = 16
+log_statement = none
+restart_after_crash = off
+});
+
+$node->start;
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+$node->safe_psql('postgres', $resize_sql_func_def);
+
+my $pgb_scale = 10;
+my $pgb_duration = 120;
+my $pgb_num_clients = 10;
+# make it easy to identify pgbench processes in pg_stat_activity
+my $application_name = 'pgbench_buffer_resize_test';
+$node->pgbench(
+ "--initialize --init-steps=dtpvg --scale=$pgb_scale --quiet",
+ 0,
+ [qr{^$}],
+ [ # stderr patterns to verify initialization stages
+ qr{dropping old tables},
+ qr{creating tables},
+ qr{done in \d+\.\d\d s }
+ ],
+ "pgbench initialization (scale=$pgb_scale)"
+);
+my ($pgbench_stdin, $pgbench_stdout, $pgbench_stderr) = ('', '', '');
+# Use --exit-on-abort so that the test stops on the first server crash or error,
+# thus making it easy to debug the failure. Use -C to increase the chances of a
+# new backend being created while resizing the buffer pool.
+my $pgbench_process = IPC::Run::start(
+ [
+ 'pgbench',
+ '-p', $node->port,
+ '-h', $node->host,
+ '-T', $pgb_duration,
+ '-c', $pgb_num_clients,
+ '-C',
+ '--exit-on-abort',
+ '--continue-on-error',
+ "dbname=postgres application_name=$application_name"
+ ],
+ '<' => \$pgbench_stdin,
+ '>' => \$pgbench_stdout,
+ '2>' => \$pgbench_stderr
+);
+
+ok($pgbench_process, "pgbench started successfully");
+
+# Resize buffer pool to various sizes while pgbench is running in the
+# background. We use smaller sizes to induce frequent buffer eviction and
+# allocation. Also smaller buffer pool means frequent wraparound in background
+# writer, default buffer allocation strategy and checkpointer.
+#
+# TODO: These are pseudo-randomly picked sizes, but we can do better.
+my $tests_completed = 0;
+
+# Reset background writer stats before starting the resize cycle
+$node->safe_psql('postgres', "SELECT pg_stat_reset_shared('bgwriter')");
+
+# Resize as many times as possible while pgbench is running.
+while (pgbench_processes_active($node, $application_name))
+{
+ for my $target_size (@buffer_sizes)
+ {
+ # Stop if pgbench finished
+ if (!pgbench_processes_active($node, $application_name))
+ {
+ last;
+ }
+
+ apply_and_verify_buffer_change($node, $target_size);
+ $tests_completed++;
+
+ # Wait for the resized buffer pool to stabilize.
+ sleep(1);
+ }
+}
+
+ok($tests_completed > scalar(@buffer_sizes), "All buffer size transitions were tested");
+note "Completed $tests_completed buffer resize operations while pgbench was running";
+
+# Check that the background writer did some work during the resize cycle
+is($node->safe_psql('postgres', "SELECT buffers_clean > 0 FROM pg_stat_bgwriter"), 't', "Background writer ran during resize cycle");
+
+# Make sure that pgbench finishes
+$pgbench_process->signal('TERM');
+ok((IPC::Run::finish $pgbench_process), "pgbench finished successfully");
+
+# Log any error output from pgbench for debugging
+diag("pgbench stderr:\n$pgbench_stderr");
+diag("pgbench stdout:\n$pgbench_stdout");
+
+# Ensure database is still functional after all the buffer changes
+$node->connect_ok("dbname=postgres",
+ "Database remains accessible after $tests_completed buffer resize operations");
+
+done_testing();
diff --git a/src/test/buffermgr/t/003_resize_fault_tolerance.pl b/src/test/buffermgr/t/003_resize_fault_tolerance.pl
new file mode 100644
index 00000000000..366929f4c45
--- /dev/null
+++ b/src/test/buffermgr/t/003_resize_fault_tolerance.pl
@@ -0,0 +1,839 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test that only one pg_resize_shared_buffers() call succeeds when multiple
+# sessions attempt to resize buffers concurrently
+
+use strict;
+use warnings;
+use Config;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# =============================================================================
+# Initialization
+# =============================================================================
+my $initial_nbuffers = 16;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
+$node->append_conf('postgresql.conf', 'max_shared_buffers = 32');
+$node->append_conf('postgresql.conf', 'restart_after_crash = on');
+$node->start;
+
+# Load injection points extension for test coordination
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# =============================================================================
+# Helper functions
+# =============================================================================
+
+# Setup resize operation to be interrupted.
+#
+# Prepare to resize the buffer pool to a target size. Start a resize session
+# through a background psql session. Adjust GUCs for the mode of interruption.
+# If injection point is provided, setup it up with the injection point and wait
+# for the resize session to reach the injection point. The resize session is
+# returned to the caller.
+sub start_resize_session
+{
+ my ($target_nbuffers, $mode, $injection_point) = @_;
+
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$target_nbuffers'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ my $session = $node->background_psql('postgres', on_error_stop => 0);
+
+ my $injection_action;
+ if (defined $injection_point)
+ {
+ $injection_action = ($mode eq 'error') ? 'error' : 'wait';
+ $session->query_safe('SELECT injection_points_set_local()', verbose => 0);
+ $session->query_safe(
+ "SELECT injection_points_attach('$injection_point', '$injection_action')",
+ verbose => 0);
+ }
+
+ apply_session_gucs_for_mode($session, $mode);
+
+ $session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ ));
+
+ # Wait for the pg_resize_shared_buffers to start waiting at the injection
+ # point.
+ if (defined $injection_point && $injection_action eq 'wait')
+ {
+ my $resize_pid = $session->{backend_pid};
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = '$injection_point' FROM pg_stat_activity WHERE pid = $resize_pid")
+ or die "timed out waiting for resize backend $resize_pid at $injection_point";
+ }
+
+ return $session;
+}
+
+# Start a backend which can be used to test the barrier handler fault tolerance.
+# We use a long pg_sleep() to simulate a load that checks for interrupts
+# regularly. The given injection point is attached to the peer backend locally
+# to induce a fault in barrier handler.
+sub start_peer_session_with_injection_point
+{
+ my ($injection_point, $action) = @_;
+
+ my $session = $node->background_psql('postgres', on_error_stop => 0);
+
+ $session->query_safe('SELECT injection_points_set_local()', verbose => 0);
+ $session->query_safe("SELECT injection_points_attach('$injection_point', '$action')",
+ verbose => 0);
+ $session->query_until(
+ qr/starting_sleep/,
+ q(
+ \echo starting_sleep
+ SELECT pg_sleep(60);
+ ));
+
+ my $peer_pid = $session->{backend_pid};
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = 'PgSleep' FROM pg_stat_activity WHERE pid = $peer_pid")
+ or die "timed out waiting for peer $peer_pid to enter pg_sleep";
+
+ return $session;
+}
+
+# Apply per-mode session GUCs locally in the given session if required.
+sub apply_session_gucs_for_mode
+{
+ my ($session, $mode) = @_;
+
+ # In timeout mode, set a statement timeout long enough for the resizing
+ # session to reach the injection point and stay there but short enough that
+ # the test doesn't take too long to fail if something goes wrong.
+ if ($mode eq 'timeout')
+ {
+ $session->query_safe("SET statement_timeout = '500ms'", verbose => 0);
+ }
+
+ # Let the resize session detect a client disconnection when testing client
+ # disconnections.
+ if ($mode eq 'disconnect')
+ {
+ $session->query_safe("SET client_connection_check_interval = '100ms'",
+ verbose => 0);
+ }
+}
+
+# Administer the interrupt corresponding to $mode against a resize session
+# that is waiting to be interrupted while resizing the buffer pool.
+sub interrupt_resize_session
+{
+ my ($mode, $session) = @_;
+
+ if ($mode eq 'terminate')
+ {
+ $node->safe_psql('postgres', "SELECT pg_terminate_backend(" . $session->{backend_pid} . ")");
+ }
+ elsif ($mode eq 'cancel')
+ {
+ $node->safe_psql('postgres', "SELECT pg_cancel_backend(" . $session->{backend_pid} . ")");
+ }
+ elsif ($mode eq 'disconnect')
+ {
+ $session->{run}->kill_kill;
+ }
+ elsif ($mode eq 'timeout')
+ {
+ # Nothing to do; statement_timeout will fire from within the resize
+ # session itself.
+ }
+ elsif ($mode eq 'error')
+ {
+ # Nothing to do; the injection point raised ERROR from within the
+ # resize backend itself.
+ }
+ else
+ {
+ die "interrupt_resize_session: unknown mode '$mode'";
+ }
+}
+
+# Function to perform checks after the resize operation has been interrupted. As
+# a result of the interruption, the resize function may finish rolling back the
+# resize or the backend executing that function may exit rolling back the resize
+# or the postmaster may restart all the backends. Perform appropriate checks by
+# detecting the post-interrupt state.
+#
+# - sentinel_session: a background psql session that is used to detect whether the
+# postmaster restarted all backends or not.
+# - resize_session: the background psql session that was executing the resize
+# operation and was interrupted.
+# - log_offset: the offset in the server log file before the resize operation was
+# initiated.
+# - injection_point and mode: the injection point and mode of interruption that
+# was used to interrupt the resize operation.
+# - orig_nbuffers and target_nbuffers: the original and target buffer sizes for
+# the resize operation.
+# - test_label: a label to create unique test names for different tests
+sub check_interrupted_resize
+{
+ my ($sentinel_session, $resize_session, $log_offset, $mode,
+ $injection_point, $orig_nbuffers, $target_nbuffers, $test_label) = @_;
+
+ my $resize_pid = $resize_session->{backend_pid};
+ my $sentinel_pid = $sentinel_session->{backend_pid};
+
+ # Wait until the resize backend is no longer running the resize query.
+ $node->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity "
+ . "WHERE pid = $resize_pid AND state = 'active' "
+ . "AND query LIKE '%pg_resize_shared_buffers%'")
+ or die
+ "timed out waiting for resize backend $resize_pid to finish";
+
+ # Wait for the postmaster to be ready in case it restarted the backends.
+ $node->poll_query_until('postgres', 'SELECT true')
+ or die "timed out waiting for postmaster liveliness check";
+
+ # Confirm the resize backend was interrupted by the intended signal.
+ # Match against the log line emitted by the resize PID so we don't
+ # accidentally pick up an unrelated message.
+ my %expected_msg = (
+ terminate => 'terminating connection due to administrator command',
+ cancel => 'canceling statement due to user request',
+ timeout => 'canceling statement due to statement timeout',
+ disconnect => 'connection to client lost',
+ error => "error triggered for injection point $injection_point",
+ );
+ my $log_pattern = qr/\[$resize_pid\][^\n]*\Q$expected_msg{$mode}\E/;
+
+ $node->wait_for_log($log_pattern, $log_offset);
+ ok($node->log_contains($log_pattern, $log_offset),
+ "$test_label: server log shows expected $mode message from pid $resize_pid"
+ );
+
+ my $server_restarted = $node->safe_psql('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE pid = $sentinel_pid"
+ ) eq 't';
+
+ if ($server_restarted)
+ {
+ # Postmaster restarted all backends; sentinel and resize sessions
+ # are dead, just reap their IPC::Run handles.
+ $sentinel_session->finish;
+ $resize_session->finish;
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after crash recovery");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after crash recovery");
+ }
+ else
+ {
+ $sentinel_session->quit;
+ # The resize session may have exited (e.g. on FATAL or disconnect).
+ if ($node->safe_psql('postgres',
+ "SELECT count(*) = 1 FROM pg_stat_activity WHERE pid = $resize_pid") eq 't')
+ {
+ $resize_session->quit;
+ }
+ else
+ {
+ $resize_session->finish;
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$orig_nbuffers|$orig_nbuffers|$orig_nbuffers|0",
+ "$test_label: buffer resize rolled back after $mode");
+
+ # TODO: Also check that the pg_shmem_allocations values are not changed
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$orig_nbuffers (pending: $target_nbuffers)",
+ "$test_label: pg_settings reports pending new value after $mode");
+
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't',
+ "$test_label: resize succeeds after interrupted resize is cleaned up");
+ }
+}
+
+# =============================================================================
+# Concurrent resize test functions
+#
+# Verify that only one pg_resize_shared_buffers() call can succeed at a time
+# using injection points.
+# =============================================================================
+
+# Workhorse function:
+#
+# Make the resize session wait at the given injection point and start another
+# concurrent resize session. The concurrent resize should fail.
+sub test_concurrent_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $session = start_resize_session($target_nbuffers, 'concurrent_resize',
+ $injection_point);
+ my $resize_pid = $session->{backend_pid};
+
+ is($node->safe_psql('postgres',
+ "SELECT resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$resize_pid", "$test_label: resizer_pid reports resize backend");
+
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 'f', "$test_label: concurrent resize fails");
+
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('$injection_point')");
+
+ $session->quit;
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool resized to target after wakeup");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after wakeup");
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points
+sub test_concurrent_resize
+{
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'pgrsb-buffer-pool-resize-barrier-sent',
+ );
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = (['expand', 24], ['shrink', $initial_nbuffers]);
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_concurrent_resize_at_injection_point($point, $target,
+ "$name: $point");
+ }
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test resize operation interruption
+#
+# Verify that an interruption in resize operation does not leave the buffer pool
+# in an inconsistent state.
+# =============================================================================
+
+# Workhorse function:
+#
+# Interrupt pg_resize_shared_buffers() when it is waiting on an injection point.
+# Check that the buffer pool is left in a consistent state as an aftermath.
+#
+# $mode selects how the resize session is interrupted.
+# - 'terminate' - SIGTERM via pg_terminate_backend() from another session.
+# - 'cancel' - SIGINT via pg_cancel_backend() from another session.
+# - 'timeout' - statement_timeout fires inside the resize session itself.
+# - 'disconnect' - the resize session's client connection is closed abruptly.
+# - 'error' - the injection point itself raises ERROR from within the
+# resize backend.
+sub test_interrupt_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $mode, $test_label) = @_;
+
+ my $orig_nbuffers = $node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+ my $log_offset = -s $node->logfile;
+
+ # Start a sentinel session that will be used to detect whether the
+ # postmaster restarted all backends or not after the resize session is
+ # interrupted.
+ my $sentinel_session = $node->background_psql('postgres', on_error_stop => 0);
+
+ my $resize_session = start_resize_session($target_nbuffers, $mode,
+ $injection_point);
+
+ interrupt_resize_session($mode, $resize_session);
+
+ check_interrupted_resize($sentinel_session, $resize_session, $log_offset,
+ $mode, $injection_point, $orig_nbuffers, $target_nbuffers,
+ $test_label);
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points passing it the
+# given mode of interruption.
+sub test_interrupt_resize_session
+{
+ my ($mode) = @_;
+
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'buffer-mgr-resize-struct',
+ 'pgrsb-buffer-pool-resize-barrier-sent',
+ );
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = (['expand', 24], ['shrink', $initial_nbuffers]);
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "$mode: buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_interrupt_resize_at_injection_point($point, $target, $mode,
+ "$mode $name: $point");
+ }
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "$mode: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test error handling in barrier handler
+#
+# Verify that an error in barrier handler does not cause a resize session to
+# fail. The barrier handler may run in a peer backend or the backend which is
+# performing the resize itself.
+# =============================================================================
+
+# Workhorse function:
+#
+# Make a peer session wait at the given injection point in the barrier handler
+# and simulate an error in the handler.
+sub test_error_in_barrier_handler_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $peer_session = start_peer_session_with_injection_point($injection_point, 'error');
+ my $peer_pid = $peer_session->{backend_pid};
+
+ my $log_offset = -s $node->logfile;
+
+ # Resize the buffer pool which will send a barrier to the peer backend
+ # simulating an error in the barrier handler.
+ $node->safe_psql('postgres',"ALTER SYSTEM SET shared_buffers = '$target_nbuffers'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"), 't');
+
+ # Confirm the peer raised the expected error from inside the handler.
+ my $log_pattern = qr/\[$peer_pid\][^\n]*\Qerror triggered for injection point $injection_point\E/;
+ $node->wait_for_log($log_pattern, $log_offset);
+ ok($node->log_contains($log_pattern, $log_offset),
+ "$test_label: server log shows error from peer pid $peer_pid at $injection_point"
+ );
+
+ # Check that the resize was completed as expected
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size");
+
+ # pg_resize_shared_buffers() returns only after every peer has
+ # acknowledged the barrier, so by this point the erroring peer has
+ # already left procArray. Assert that and then reap its IPC::Run handle.
+ is($node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_stat_activity WHERE pid = $peer_pid"),
+ '0',
+ "$test_label: peer pid $peer_pid exited after handler error");
+
+ $peer_session->finish;
+}
+
+# Driver function:
+#
+# Simulate a failure to change the protection on the shared memory. This should
+# cause the barrier handler to raise an error. The barrier handler may run in a
+# peer backend or the backend which is performing the resize itself. The resize
+# session should still complete successfully.
+sub test_error_in_barrier_handler
+{
+ my $injection_point = 'buffer-mgr-protect-struct';
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "error-in-handler: buffer pool size is $initial_nbuffers at start");
+
+ # Test error in barrier handler in a peer backend. Expand then shrink so
+ # the pool returns to its starting size.
+ for my $dir (['expand', 24], ['shrink', $initial_nbuffers])
+ {
+ my ($name, $target) = @$dir;
+ test_error_in_barrier_handler_at_injection_point($injection_point,
+ $target, "error-in-handler peer $name");
+ }
+
+ # Test the same error in the barrier handler in the resize backend itself.
+ # Expand then shrink so the pool returns to its starting size.
+ for my $dir (['expand', 24], ['shrink', $initial_nbuffers])
+ {
+ my ($name, $target) = @$dir;
+ test_interrupt_resize_at_injection_point($injection_point,
+ $target, 'error', "error-in-handler resize-backend $name");
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "error-in-handler: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test fault tolerance of resize operation waiting for barrier
+#
+# Test that, when interrupted, a resizing operation waiting for a barrier to be
+# acknowledged doesn't leave the buffer pool in an inconsistent state.
+# =============================================================================
+
+# Workhorse function:
+#
+# We start a peer session with the given injection point in the barrier handler
+# code attached locally. Once the resize operation starts, the peer session will
+# hit the injection point and wait there. Interrupt the resize session and check
+# that the buffer pool is left in a consistent state as an aftermath.
+#
+# - injection_point: the injection point to attach to the peer session.
+# - target_nbuffers: the target buffer size for the resize operation.
+# - mode: the mode of interruption to apply to the resize session.
+# - test_label: a label to create unique test names for different tests
+sub test_fault_resize_waiting_barrier
+{
+ my ($injection_point, $target_nbuffers, $mode, $test_label) = @_;
+
+ my $orig_nbuffers = $node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+ my $log_offset = -s $node->logfile;
+
+ my $peer_session = start_peer_session_with_injection_point($injection_point, 'wait');
+ my $peer_pid = $peer_session->{backend_pid};
+
+ # Sentinel session to detect a postmaster restart.
+ my $sentinel_session = $node->background_psql('postgres', on_error_stop => 0);
+
+ my $resize_session = start_resize_session($target_nbuffers, $mode);
+ my $resize_pid = $resize_session->{backend_pid};
+
+ # Wait for the peer to reach the injection point. At this point the resize
+ # backend should be blocked in WaitForProcSignalBarrier.
+ $node->poll_query_until('postgres',
+ "SELECT wait_event = '$injection_point' FROM pg_stat_activity WHERE pid = $peer_pid")
+ or die "$test_label: timed out waiting for peer $peer_pid at $injection_point";
+ is($node->safe_psql('postgres',
+ "SELECT wait_event FROM pg_stat_activity WHERE pid = $resize_pid"),
+ 'ProcSignalBarrier',
+ "$test_label: resize $resize_pid is waiting at ProcSignalBarrier");
+
+ interrupt_resize_session($mode, $resize_session);
+
+ check_interrupted_resize($sentinel_session, $resize_session, $log_offset,
+ $mode, $injection_point, $orig_nbuffers, $target_nbuffers, $test_label);
+
+ # Cleanup peer session. If the postmaster restarted all backends, the peer
+ # backend is already gone.
+ if ($node->safe_psql('postgres',
+ "SELECT count(*) = 1 FROM pg_stat_activity WHERE pid = $peer_pid") eq 't')
+ {
+ $peer_session->quit;
+ }
+ else
+ {
+ $peer_session->finish;
+ }
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for different injection points passing it the
+# given mode of interruption.
+sub test_fault_resize_waiting_barrier_for_mode
+{
+ my ($mode) = @_;
+
+ my @injection_points = (
+ 'pgrsb-handle-new-buffer-alloc-barrier',
+ 'pgrsb-handle-buffer-pool-size-barrier',
+ 'pgrsb-handle-buffer-pool-resize-barrier',
+ );
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = (['expand', 24], ['shrink', $initial_nbuffers]);
+
+ is($node->safe_psql('postgres', "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "fault-resize-on-peer $mode: buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+
+ test_fault_resize_waiting_barrier($point, $target, $mode,
+ "fault-resize-on-peer $mode $name: $point");
+ }
+ }
+
+ is($node->safe_psql('postgres', "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "fault-resize-on-peer $mode: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Functions to test server restart during a resize
+#
+# Verify that a server can be stopped and started while a resize operation is in
+# progress and the server is started with buffer pool in a consistent state that
+# reflects the target size.
+# =============================================================================
+
+# Workhorse function for fast/immediate shutdown:
+#
+# Make the resize session wait at the given injection point and restart the
+# server in the given mode.
+sub test_server_restart_during_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $stop_mode, $test_label) = @_;
+
+ my $resize_session = start_resize_session($target_nbuffers,
+ 'server_restart', $injection_point);
+
+ $node->stop($stop_mode);
+
+ # Cleanup resize session, the backend must have gone now.
+ $resize_session->finish;
+
+ $node->start;
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after $stop_mode restart");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after $stop_mode restart");
+}
+
+# Workhorse function for smart shutdown:
+#
+# Let the resize operation wait at the given injection point, send a smart
+# shutdown asynchronously. Once the postmaster enters smart shutdown, wakeup the
+# resize backend and let it complete. Verify that the pool reflects the target
+# size immediately and also after the restart.
+#
+# - injection_point: the injection point to park the resize backend at.
+# - target_nbuffers: the target buffer size for the resize operation.
+# - test_label: a label to create unique test names for different tests
+sub test_server_restart_smart_during_resize_at_injection_point
+{
+ my ($injection_point, $target_nbuffers, $test_label) = @_;
+
+ my $log_offset = -s $node->logfile;
+
+ my $resize_session = start_resize_session($target_nbuffers, 'server_restart', $injection_point);
+ my $resize_pid = $resize_session->{backend_pid};
+
+ # Open another session which can be used to wakeup the resize backend.
+ my $control_session = $node->background_psql('postgres', on_error_stop => 0);
+
+ # Start the process to stop the server in smart mode.
+ local %ENV = $node->_get_env();
+ my @stop_cmd = ('pg_ctl', '--pgdata' => $node->data_dir, '--mode' => 'smart', 'stop');
+ my ($stop_in, $stop_out, $stop_err) = ('', '', '');
+ my $stop_session = IPC::Run::start(\@stop_cmd,
+ \$stop_in, \$stop_out, \$stop_err);
+
+ # Confirm the postmaster entered smart shutdown.
+ $node->wait_for_log(qr/received smart shutdown request/, $log_offset);
+
+ # Make sure that the resize backend is still alive
+ is($control_session->query("SELECT wait_event FROM pg_stat_activity WHERE pid = $resize_pid"),
+ $injection_point,
+ "$test_label: resize backend alive during smart shutdown");
+
+ # Wake up the resize backnd and let it finish.
+ $control_session->query_safe("SELECT injection_points_wakeup('$injection_point')",
+ verbose => 0);
+
+ # Check that the resize finished successfully by querying from the same
+ # session. The queries won't return if the resize didn't finish. Accomodate
+ # the output 't' from pg_resize_shared_buffers() in the expected output of
+ # the first query.
+ is($resize_session->query("SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()",
+ verbose => 0),
+ "t\n$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: pg_resize_shared_buffers() succeeded and pool at target during smart shutdown");
+ is($resize_session->query("SELECT setting FROM pg_settings WHERE name = 'shared_buffers'",
+ verbose => 0),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size during smart shutdown");
+
+ $resize_session->quit;
+ $control_session->quit;
+
+ # Wait for server to stop
+ IPC::Run::finish($stop_session)
+ or die "$test_label: pg_ctl smart stop failed: $stop_err";
+
+ # Sync Cluster.pm internal state and start the cluster back.
+ $node->{_pid} = undef;
+ $node->start;
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$target_nbuffers|$target_nbuffers|$target_nbuffers|0",
+ "$test_label: buffer pool reflects target size after smart restart");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$target_nbuffers",
+ "$test_label: pg_settings reports target size after smart restart");
+}
+
+# Driver function:
+#
+# Invoke the workhorse function for every injection point in resize operation in
+# both directions for the given stop mode.
+sub test_server_restart_during_resize
+{
+ my ($stop_mode) = @_;
+
+ my @injection_points = (
+ 'pg-resize-shared-buffers-flag-set',
+ 'pgrsb-new-buffer-alloc-barrier-sent',
+ 'pgrsb-buffer-pool-size-barrier-sent',
+ 'pgrsb-buffer-pool-resize-barrier-sent',
+ );
+
+ # Expand then shrink so the pool returns to its starting size.
+ my @directions = (['expand', 24], ['shrink', $initial_nbuffers]);
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "server-restart $stop_mode: buffer pool size is $initial_nbuffers at start");
+
+ for my $point (@injection_points)
+ {
+ for my $dir (@directions)
+ {
+ my ($name, $target) = @$dir;
+ my $label = "server-restart $stop_mode $name: $point";
+
+ if ($stop_mode eq 'smart')
+ {
+ test_server_restart_smart_during_resize_at_injection_point(
+ $point, $target, $label);
+ }
+ else
+ {
+ test_server_restart_during_resize_at_injection_point($point,
+ $target, $stop_mode, $label);
+ }
+ }
+ }
+
+ is($node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers",
+ "server-restart $stop_mode: buffer pool size is $initial_nbuffers at end");
+}
+
+# =============================================================================
+# Run tests
+# =============================================================================
+test_concurrent_resize();
+test_error_in_barrier_handler();
+
+test_interrupt_resize_session('terminate');
+test_interrupt_resize_session('cancel');
+test_interrupt_resize_session('timeout');
+test_interrupt_resize_session('error');
+
+# A resize session waiting for a barrier to be acknowledged can not be
+# interrupted by an error. Hence don't test that mode.
+test_fault_resize_waiting_barrier_for_mode('terminate');
+test_fault_resize_waiting_barrier_for_mode('cancel');
+test_fault_resize_waiting_barrier_for_mode('timeout');
+
+test_server_restart_during_resize('immediate');
+test_server_restart_during_resize('fast');
+test_server_restart_during_resize('smart');
+
+# client_connection_check_interval is only effective on systems that expose
+# POLLRDHUP/EPOLLRDHUP (Linux, and a few other Unix variants). On other
+# platforms the GUC is silently a no-op, so the disconnect test would hang.
+if ($Config::Config{osname} eq 'linux')
+{
+ test_interrupt_resize_session('disconnect');
+ test_fault_resize_waiting_barrier_for_mode('disconnect');
+}
+else
+{
+ diag("skipping disconnect interrupt test on $Config::Config{osname} "
+ . "(requires POLLRDHUP support)");
+}
+
+done_testing();
+
+# Few more tests to add but may be somewhere else
+# TODO: test when there are backends that have not attached to the shared memory
+# TODO: test that a non-superuser cannot run pg_resize_shared_buffers()
+# TODO: the resize_sql_func_def in 001_resize_buffer may be useful in other
+# tests (not necessarily this one). Maybe we can use it in other tests where we
+# are looping in TAP test code.
diff --git a/src/test/buffermgr/t/004_client_join_buffer_resize.pl b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
new file mode 100644
index 00000000000..fda0f01bb27
--- /dev/null
+++ b/src/test/buffermgr/t/004_client_join_buffer_resize.pl
@@ -0,0 +1,221 @@
+# Copyright (c) 2025-2025, PostgreSQL Global Development Group
+#
+# Test shared_buffer resizing coordination with client connections joining using injection points
+use strict;
+use warnings;
+use IPC::Run;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+use Time::HiRes qw(sleep);
+
+# Skip this test if injection points are not supported
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Function to calculate the size of test table required to fill up maximum
+# buffer pool when populating it.
+sub calculate_test_sizes
+{
+ my ($node, $block_size) = @_;
+
+ # Get the maximum buffer pool size from configuration
+ my $max_shared_buffers = $node->safe_psql('postgres', "SHOW max_shared_buffers");
+ my ($max_val, $max_unit) = ($max_shared_buffers =~ /(\d+)(\w+)/);
+ my $max_size_bytes;
+ if (lc($max_unit) eq 'kb') {
+ $max_size_bytes = $max_val * 1024;
+ } elsif (lc($max_unit) eq 'mb') {
+ $max_size_bytes = $max_val * 1024 * 1024;
+ } elsif (lc($max_unit) eq 'gb') {
+ $max_size_bytes = $max_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $max_size_bytes = $max_val * 1024;
+ }
+
+ # Fill more pages than minimally required to increase the chances of pages
+ # from the test table filling the buffer cache.
+ $max_size_bytes = $max_size_bytes;
+ my $pages_needed = int($max_size_bytes / $block_size) + 10; # Add some extra to ensure buffers are filled
+ my $rows_to_insert = $pages_needed * 100; # Assuming roughly 100 rows per page for our table structure
+ return ($max_size_bytes, $pages_needed, $rows_to_insert);
+}
+
+# Function to calculate expected buffer count from size string
+sub calculate_buffer_count
+{
+ my ($size_string, $block_size) = @_;
+ # Parse size and convert to bytes
+ my ($size_val, $unit) = ($size_string =~ /(\d+)(\w+)/);
+ my $size_bytes;
+ if (lc($unit) eq 'kb') {
+ $size_bytes = $size_val * 1024;
+ } elsif (lc($unit) eq 'mb') {
+ $size_bytes = $size_val * 1024 * 1024;
+ } elsif (lc($unit) eq 'gb') {
+ $size_bytes = $size_val * 1024 * 1024 * 1024;
+ } else {
+ # Default to kB if unit is not recognized
+ $size_bytes = $size_val * 1024;
+ }
+ return int($size_bytes / $block_size);
+}
+
+# Initialize cluster with very small buffer sizes for testing
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# Configure for buffer resizing with very small buffer pool sizes for faster tests.
+# TODO: for some reason parallel workers try to load default number of shared_buffers which doesn't work with lower max_shared_buffers. We need to fix that - somewhere it's picking default value of shared buffers. For now disable parallelism
+$node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', qq{
+max_shared_buffers = 512kB
+shared_buffers = 320kB
+max_parallel_workers_per_gather = 0
+});
+$node->start;
+
+# Enable injection points
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+# Get the block size (this is fixed for the binary)
+my $block_size = $node->safe_psql('postgres', "SHOW block_size");
+
+# Try to create pg_buffercache extension for buffer analysis
+eval {
+ $node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+};
+if ($@) {
+ $node->stop;
+ plan skip_all => 'pg_buffercache extension not available - cannot verify buffer usage';
+}
+
+# Create a small test table, and fetch its properties for later reference if required.
+$node->safe_psql('postgres', qq{
+ CREATE TABLE client_test (c1 int, data char(50));
+});
+my $table_oid = $node->safe_psql('postgres', "SELECT oid FROM pg_class WHERE relname = 'client_test'");
+my $table_relfilenode = $node->safe_psql('postgres', "SELECT relfilenode FROM pg_class WHERE relname = 'client_test'");
+note("Test table client_test: OID = $table_oid, relfilenode = $table_relfilenode");
+my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $block_size);
+
+# Create dedicated sessions for injection point handling and test queries,
+# so that we don't create new backends for test operations after starting
+# resize operation. Only one backend, which tests new backend synchronization
+# with resizing operation, should start after resizing has commenced.
+my $injection_session = $node->background_psql('postgres');
+my $query_session = $node->background_psql('postgres');
+my $resize_session = $node->background_psql('postgres');
+
+# Function to run a single injection point test
+sub run_injection_point_test
+{
+ my ($test_name, $injection_point, $target_size, $operation_type) = @_;
+
+ # Silence the logging of the statements we run to avoid
+ # unnecessarily bloating the test logs. This runs before the
+ # upgrade we're testing, so the details should not be very
+ # interesting for debugging. But if needed, you can make it more
+ # verbose by setting this.
+ my $verbose = 0;
+
+ note("Test with $test_name ($operation_type)");
+
+ # Calculate test parameters before starting resize
+ my ($max_size_bytes, $pages_needed, $rows_to_insert) = calculate_test_sizes($node, $target_size, $block_size);
+
+ # Update buffer pool size and wait for it to reflect pending state
+ $resize_session->query_safe("ALTER SYSTEM SET shared_buffers = '$target_size'", verbose => $verbose);
+ $resize_session->query_safe("SELECT pg_reload_conf()", verbose => $verbose);
+ my $pending_size_str = "pending: $target_size";
+ $resize_session->poll_query_until("SELECT substring(current_setting('shared_buffers'), '$pending_size_str')", $pending_size_str, verbose => $verbose);
+
+ # Set up injection point in injection session
+ $injection_session->query_safe("SELECT injection_points_attach('$injection_point', 'wait')", verbose => $verbose);
+
+ # Trigger resize
+ $resize_session->query_until(
+ qr/starting_resize/,
+ q(
+ \echo starting_resize
+ SELECT pg_resize_shared_buffers();
+ )
+ );
+
+ # Wait until resize actually reaches the injection point using the query session
+ $query_session->wait_for_event('client backend', $injection_point, verbose => $verbose);
+
+ # Start a client while resize is paused
+ my $client = $node->background_psql('postgres');
+ note("Background client backend PID: " . $client->query_safe("SELECT pg_backend_pid()", verbose => $verbose));
+
+ # Wake up the injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_wakeup('$injection_point')", verbose => $verbose);
+
+ # Test buffer functionality immediately after waking up injection point
+ # Insert data to test buffer pool functionality during/after resize
+ $client->query_safe("INSERT INTO client_test SELECT i, 'test_data_' || i FROM generate_series(1, $rows_to_insert) i", verbose => $verbose);
+ # Verify the data was inserted correctly and can be read back
+ is($client->query_safe("SELECT COUNT(*) FROM client_test", verbose => $verbose), $rows_to_insert, "inserted $rows_to_insert during $test_name ($operation_type) successful");
+
+ # Verify table size is reasonable (should be substantial for testing)
+ ok($query_session->query_safe("SELECT pg_total_relation_size('client_test')", verbose => $verbose) >= $max_size_bytes,"table size is large enough to overflow buffer pool in test $test_name ($operation_type)");
+
+ # Wait for the resize operation to complete. There is no direct way to do so
+ # in background_psql. Hence fire a psql command and wait for it to finish
+ $resize_session->query(q(\echo 'done'), verbose => $verbose);
+
+ # Detach injection point from injection session
+ $injection_session->query_safe("SELECT injection_points_detach('$injection_point')", verbose => $verbose);
+
+ # Verify resize completed successfully
+ is($query_session->query_safe("SELECT current_setting('shared_buffers')", verbose => $verbose), $target_size,
+ "resize completed successfully to $target_size");
+
+ # Check buffer pool size using pg_buffercache after resize completion
+ is($query_session->query_safe("SELECT COUNT(*) FROM pg_buffercache", verbose => $verbose), calculate_buffer_count($target_size, $block_size), "all buffers in the buffer pool used in $test_name ($operation_type)");
+
+ # Wait for client to complete
+ ok($client->quit, "client succeeded during $test_name ($operation_type)");
+
+ # Clean up for next test
+ $query_session->query_safe("DELETE FROM client_test", verbose => $verbose);
+}
+
+# Test new client joining during various phases of buffer resizing operation using injection points
+my @injection_tests = (
+ {
+ name => 'flag setting phase',
+ injection_point => 'pg-resize-shared-buffers-flag-set',
+ },
+ {
+ name => 'new buffer alloc barrier complete',
+ injection_point => 'pgrsb-new-buffer-alloc-barrier-sent',
+ },
+ {
+ name => 'buffer pool size barrier complete',
+ injection_point => 'pgrsb-buffer-pool-size-barrier-sent',
+ },
+ {
+ name => 'buffer pool resize barrier complete',
+ injection_point => 'pgrsb-buffer-pool-resize-barrier-sent',
+ },
+);
+
+foreach my $test (@injection_tests)
+{
+ # Test shrinking scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '272kB', 'shrinking');
+
+ # Test expanding scenario
+ run_injection_point_test($test->{name}, $test->{injection_point}, '400kB', 'expanding');
+}
+
+$injection_session->quit;
+$query_session->quit;
+$resize_session->quit;
+
+done_testing();
diff --git a/src/test/buffermgr/t/005_resize_failures.pl b/src/test/buffermgr/t/005_resize_failures.pl
new file mode 100644
index 00000000000..820aa07f006
--- /dev/null
+++ b/src/test/buffermgr/t/005_resize_failures.pl
@@ -0,0 +1,171 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Test that pg_resize_shared_buffers() rolls back cleanly when resize fails.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $have_injection_points = ($ENV{enable_injection_points} eq 'yes');
+
+# Start the pool large enough that there is room to shrink below a pinned
+# buffer while still satisfying the shared_buffers GUC minimum.
+my $initial_nbuffers = 24;
+my $max_nbuffers = 32;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+if ($have_injection_points)
+{
+ $node->append_conf('postgresql.conf', 'shared_preload_libraries = injection_points');
+}
+$node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
+$node->append_conf('postgresql.conf', "max_shared_buffers = $max_nbuffers");
+$node->start;
+
+# pg_buffercache lets us locate the bufferid holding a given page.
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+if ($have_injection_points)
+{
+ $node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+}
+
+# ---------------------------------------------------------------------------
+# Test the case when shrinking is aborted by a pinned buffer
+# ---------------------------------------------------------------------------
+
+my $min_nbuffers = $node->safe_psql('postgres',
+ "SELECT min_val::int FROM pg_settings WHERE name = 'shared_buffers'");
+
+# In order to reliably pin a buffer above $min_nbuffers, we create as many
+# tables $min_nbuffers + 1, open a cursor on the tables and fetch one row from
+# each cursor one at a time. This will pin one buffer per table, guaranteeing
+# that at least one of the pinned buffers will be above $min_nbuffers.
+my $ntables = $min_nbuffers + 1;
+for my $i (1 .. $ntables)
+{
+ $node->safe_psql('postgres', "CREATE TABLE evict_target_$i AS SELECT generate_series(1, 2) AS i");
+}
+my $pinner = $node->background_psql('postgres', on_error_stop => 0);
+$pinner->query_safe("BEGIN", verbose => 0);
+my $pinned_buf = 0;
+for my $i (1 .. $ntables)
+{
+ $pinner->query_safe("DECLARE c_$i CURSOR FOR SELECT * FROM evict_target_$i", verbose => 0);
+ $pinner->query_safe("FETCH 1 FROM c_$i", verbose => 0);
+
+ my $buf = $node->safe_psql('postgres',
+ "SELECT min(bufferid) FROM pg_buffercache WHERE pinning_backends > 0 AND bufferid > $min_nbuffers");
+
+ if ($buf =~ /^\d+$/)
+ {
+ $pinned_buf = $buf;
+ last;
+ }
+}
+cmp_ok($pinned_buf, '>', $min_nbuffers, "pinned a buffer above $min_nbuffers");
+
+# Set the target so that the pinned buffer is in the range of buffers to be evicted.
+my $shrink_target = $pinned_buf - 1;
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$shrink_target'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+my $log_offset = -s $node->logfile;
+
+is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 'f',
+ "shrink returns false when a buffer to be evicted is pinned");
+ok($node->log_contains(qr/could not remove buffer $pinned_buf, it is pinned/, $log_offset),
+ "log reports the pinned buffer that blocked eviction");
+ok($node->log_contains(qr/failed to evict extra buffers during shrinking/, $log_offset),
+ "log reports the eviction failure");
+
+is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$initial_nbuffers|$initial_nbuffers|$initial_nbuffers|0",
+ "pool unchanged after eviction failure");
+
+is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$initial_nbuffers (pending: $shrink_target)",
+ "pg_settings reports pending shrink target");
+
+# Releasing all pins lets the retry succeed.
+$pinner->quit;
+
+is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't',
+ "shrink succeeds after pins released");
+
+is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$shrink_target|$shrink_target|$shrink_target|0",
+ "pool shrunk to $shrink_target after pins released");
+
+# ---------------------------------------------------------------------------
+# Test the case when memory allocation fails when expanding the buffer pool.
+# Uses an injection point to simulate the failure without exhausting real
+# memory.
+# ---------------------------------------------------------------------------
+
+SKIP:
+{
+ skip "injection points not supported by this build"
+ unless $have_injection_points;
+
+ # The buffer manager's resizable structures whose sizes must be rolled
+ # back if any one of them fails to grow.
+ my $resizable_structs =
+ q{('Buffer Descriptors', 'Buffer Blocks', 'Buffer IO Condition Variables', 'Checkpoint BufferIds')};
+ my $sizes_query = "SELECT name, size FROM pg_shmem_allocations WHERE name IN $resizable_structs ORDER BY name";
+ my $sizes_before = $node->safe_psql('postgres', $sizes_query);
+
+ my $resizer = $node->background_psql('postgres');
+ $resizer->query_safe("SELECT injection_points_set_local()", verbose => 0);
+ $resizer->query_safe("SELECT injection_points_attach('buffer-mgr-resize-struct-fail', 'notice')",
+ verbose => 0);
+
+ my $expand_target = $max_nbuffers;
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$expand_target'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+ my $expand_log_offset = -s $node->logfile;
+
+ is($resizer->query("SELECT pg_resize_shared_buffers()"), 'f',
+ "expansion fails when a structure can not be expanded");
+
+ # Discard the expected WARNINGs so later query_safe calls do not die.
+ $resizer->{stderr} = '';
+
+ ok($node->log_contains(qr/failed to expand buffer pool structures/, $expand_log_offset),
+ "log reports the expansion failure");
+
+ is($node->safe_psql('postgres',
+ "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$shrink_target|$shrink_target|$shrink_target|0",
+ "buffer pool status after expansion failure");
+
+ is($node->safe_psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
+ "$shrink_target (pending: $expand_target)",
+ "pg_settings reports pending expand target");
+
+ is($node->safe_psql('postgres', $sizes_query), $sizes_before,
+ "resizable buffer manager structures rolled back to previous sizes");
+
+ # Detach the injection point, to retry again. The retry should succeed.
+ $resizer->query_safe(
+ "SELECT injection_points_detach('buffer-mgr-resize-struct-fail')",
+ verbose => 0);
+ is($resizer->query("SELECT pg_resize_shared_buffers()"), 't',
+ "expand succeeds after the injection point is detached");
+
+ $resizer->quit;
+
+ is($node->safe_psql('postgres', "SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid FROM pg_get_buffer_resize_status()"),
+ "$expand_target|$expand_target|$expand_target|0",
+ "pool expanded to $expand_target after detach");
+}
+
+done_testing();
diff --git a/src/test/buffermgr/t/006_resize_with_syslogger.pl b/src/test/buffermgr/t/006_resize_with_syslogger.pl
new file mode 100644
index 00000000000..75b047ad984
--- /dev/null
+++ b/src/test/buffermgr/t/006_resize_with_syslogger.pl
@@ -0,0 +1,58 @@
+# Copyright (c) 2026-2026, PostgreSQL Global Development Group
+#
+# Test that pg_resize_shared_buffers() works when a backend that never
+# attaches to shared memory is running.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $initial_nbuffers = 16;
+my $expanded_nbuffers = 24;
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+# When logging_collector is on, the server starts a syslogger process that never
+# attaches to the shared memory. We use that as a proxy for a backend that never
+# attaches to shared memory.
+$node->append_conf(
+ 'postgresql.conf', qq{
+shared_buffers = $initial_nbuffers
+max_shared_buffers = $expanded_nbuffers
+logging_collector = on
+});
+$node->start;
+
+# Check that the syslogger is running by writing a log marker and waiting for it
+# to appear in the log file.
+sub check_syslogger_running
+{
+ my ($marker) = @_;
+
+ $node->safe_psql('postgres', "DO \$\$ BEGIN RAISE LOG '$marker'; END \$\$");
+ return $node->poll_query_until('postgres', "SELECT pg_read_file(pg_current_logfile()) ~ '$marker'");
+}
+
+check_syslogger_running('syslogger_marker_before_resize')
+ or die "syslogger is not running";
+
+# Resize the buffer pool, and check that the syslogger continues to run while
+# the resize is in progress.
+# TODO: Instead of custom markers we could use the log line that is emitted when
+# the resize is complete, when we have frozen those.
+for my $dir (['expand', $expanded_nbuffers], ['shrink', $initial_nbuffers])
+{
+ my ($name, $target) = @$dir;
+
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$target'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+ is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
+ 't',
+ "$name to $target succeeds with syslogger running");
+ ok(check_syslogger_running("syslogger_marker_after_$name"),
+ "syslogger drains logs after $name");
+}
+
+done_testing();
diff --git a/src/test/meson.build b/src/test/meson.build
index cd45cbf57fb..e9550933063 100644
--- a/src/test/meson.build
+++ b/src/test/meson.build
@@ -4,6 +4,7 @@ subdir('regress')
subdir('isolation')
subdir('authentication')
+subdir('buffermgr')
subdir('postmaster')
subdir('recovery')
subdir('subscription')
diff --git a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
index 699334320d9..e02c5314f4b 100644
--- a/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
+++ b/src/test/perl/PostgreSQL/Test/BackgroundPsql.pm
@@ -61,6 +61,7 @@ use Config;
use IPC::Run;
use PostgreSQL::Test::Utils qw(pump_until);
use Test::More;
+use Time::HiRes qw(usleep);
=pod
@@ -403,4 +404,79 @@ sub set_query_timer_restart
return $self->{query_timer_restart};
}
+=pod
+
+=item $session->poll_query_until($query [, $expected ])
+
+Run B<$query> repeatedly in this background session, until it returns the
+B<$expected> result ('t', or SQL boolean true, by default).
+Continues polling if the query returns an error result.
+Times out after a reasonable number of attempts.
+Returns 1 if successful, 0 if timed out.
+
+=cut
+
+sub poll_query_until
+{
+ my ($self, $query, $expected, %params) = @_;
+
+ $expected = 't' unless defined($expected); # default value
+
+ my $max_attempts = 10 * $PostgreSQL::Test::Utils::timeout_default;
+ my $attempts = 0;
+ my ($stdout, $stderr_flag);
+
+ while ($attempts < $max_attempts)
+ {
+ ($stdout, $stderr_flag) = $self->query($query, %params);
+
+ chomp($stdout);
+
+ # If query succeeded and returned expected result
+ if (!$stderr_flag && $stdout eq $expected)
+ {
+ return 1;
+ }
+
+ # Wait 0.1 second before retrying.
+ usleep(100_000);
+
+ $attempts++;
+ }
+
+ # Give up. Print the output from the last attempt, hopefully that's useful
+ # for debugging.
+ my $stderr_output = $stderr_flag ? $self->{stderr} : '';
+ diag qq(poll_query_until timed out executing this query:
+$query
+expecting this output:
+$expected
+last actual query output:
+$stdout
+with stderr:
+$stderr_output);
+ return 0;
+}
+
+=item $session->wait_for_event(backend_type, wait_event_name)
+
+Poll pg_stat_activity until backend_type reaches wait_event_name using this
+background session.
+
+=cut
+
+sub wait_for_event
+{
+ my ($self, $backend_type, $wait_event_name, %params) = @_;
+
+ $self->poll_query_until(qq[
+ SELECT count(*) > 0 FROM pg_stat_activity
+ WHERE backend_type = '$backend_type' AND wait_event = '$wait_event_name'
+ ], undef, %params)
+ or die
+ qq(timed out when waiting for $backend_type to reach wait event '$wait_event_name');
+
+ return;
+}
+
1;
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index a823271262c..01a7b957b3a 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -359,6 +359,7 @@ BufferAccessStrategy
BufferAccessStrategyType
BufferCacheOsPagesContext
BufferCacheOsPagesRec
+BufferControlBlock
BufferDesc
BufferDescPadded
BufferHeapTupleTableSlot
[application/octet-stream] v20260917-0009-fixup-Decouple-GUC-shared_buffers-and-size-of-the-buffer-pool.patch (927B, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/11-v20260917-0009-fixup-Decouple-GUC-shared_buffers-and-size-of-the-buffer-pool.patch)
download | inline diff:
From 513b9342dff0b6f7bb75126b7cdc8e89feaf4cb7 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 5 Aug 2026 17:38:00 +0530
Subject: [PATCH] fixup! Decouple GUC shared_buffers and size of the buffer
pool
---
contrib/pg_prewarm/autoprewarm.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index e12323fb7c6..dd202894b4e 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -855,7 +855,7 @@ autoprewarm_start_worker(PG_FUNCTION_ARGS)
* SQL-callable function to perform an immediate block dump.
*
* Note: this is declared to return int8, as insurance against some
- * very distant day when we might make NBuffers wider than int.
+ * very distant day when we might make the size of buffer pool wider than int.
*/
Datum
autoprewarm_dump_now(PG_FUNCTION_ARGS)
[application/octet-stream] v20260917-0010-squash-Follow-up-changes-since-last-email-on-hackers.patch (3.1K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/12-v20260917-0010-squash-Follow-up-changes-since-last-email-on-hackers.patch)
download | inline diff:
From 44debab4e13fe84bec3cc8ce7cb4ea8d971cfd43 Mon Sep 17 00:00:00 2001
From: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
Date: Wed, 5 Aug 2026 18:16:45 +0530
Subject: [PATCH] squash! Follow-up changes since last email on hackers
Fix apw_dump_now() to not assume that the buffer pool is fixed sized.
---
contrib/pg_prewarm/autoprewarm.c | 32 ++++++++++++++-----
src/test/buffermgr/t/016_stress_pg_prewarm.pl | 3 --
2 files changed, 24 insertions(+), 11 deletions(-)
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index dd202894b4e..2150d82f19c 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -675,6 +675,7 @@ static int
apw_dump_now(bool is_bgworker, bool dump_unlogged)
{
int num_blocks;
+ int max_blocks;
int i;
int ret;
BlockInfoRecord *block_info_array;
@@ -702,26 +703,35 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
return 0;
}
- /*
- * TODO: we need to modify this function to not rely on NBuffers being
- * constant.
- */
-
/*
* With sufficiently large shared_buffers, allocation will exceed 1GB, so
- * allow for a huge allocation to prevent outright failure.
+ * allow for a huge allocation to prevent outright failure. Use the current
+ * size of the buffer pool as the estimate of the number of blocks to dump,
+ * and grow the array if necessary.
*
* (In the future, it might be a good idea to redesign this to use a more
* memory-efficient data structure.)
*/
+ max_blocks = NBuffers;
block_info_array = (BlockInfoRecord *)
- palloc_extended((sizeof(BlockInfoRecord) * NBuffers), MCXT_ALLOC_HUGE);
+ palloc_extended((sizeof(BlockInfoRecord) * max_blocks), MCXT_ALLOC_HUGE);
for (num_blocks = 0, i = 0; i < NBuffers; i++)
{
uint64 buf_state;
- CHECK_FOR_INTERRUPTS();
+ /*
+ * Expand the array if necessary using the latest size of the buffer
+ * pool as the estimate of the number of blocks to dump.
+ */
+ if (num_blocks >= max_blocks)
+ {
+ max_blocks = NBuffers;
+ block_info_array = (BlockInfoRecord *)
+ repalloc_extended(block_info_array,
+ sizeof(BlockInfoRecord) * max_blocks,
+ MCXT_ALLOC_HUGE);
+ }
bufHdr = GetBufferDescriptor(i);
@@ -747,6 +757,12 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
}
UnlockBufHdr(bufHdr);
+
+ /*
+ * Check for interrupts here, at the end of the loop, so that the buffer
+ * index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
snprintf(transient_dump_file_path, MAXPGPATH, "%s.tmp", AUTOPREWARM_FILE);
diff --git a/src/test/buffermgr/t/016_stress_pg_prewarm.pl b/src/test/buffermgr/t/016_stress_pg_prewarm.pl
index 23161f8865a..39b44191595 100644
--- a/src/test/buffermgr/t/016_stress_pg_prewarm.pl
+++ b/src/test/buffermgr/t/016_stress_pg_prewarm.pl
@@ -1,9 +1,6 @@
# Copyright (c) 2025-2026, PostgreSQL Global Development Group
#
# Stress test the autoprewarm concurrently with shared_buffers resizing.
-#
-# TODO: This test fails because apw_dump_now() assumes NBuffers is
-# constant. Fix is on the way.
use strict;
use warnings;
[application/octet-stream] v20260917-0011-pg_buffercache-process-one-buffer-at-a-time.patch (15.0K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/13-v20260917-0011-pg_buffercache-process-one-buffer-at-a-time.patch)
download | inline diff:
From 6aa7ffd67406eba7365fdae41305532741157aa4 Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palak.chaturvedi@example.com>
Date: Fri, 14 Aug 2026 04:30:51 +0000
Subject: [PATCH] pg_buffercache: process one buffer at a time in
pg_buffercache_os_pages()
Following Ashutosh's suggestion, walk the buffer pool one buffer at a
time. For each buffer we compute the OS pages it overlaps, query the
NUMA state of just those pages via pg_numa_query_pages() with a small
stack-allocated scratch array, and stream one row per OS page directly
into the tuplestore.
This eliminates the fixed-size upfront allocation of fctx->record[] and
os_page_status[] that the old code sized from NBuffers, which is the
value pg_resize_shared_buffers() may change while this function runs.
No snapshot, no upfront allocation tied to NBuffers, no min-cap loop
bound, no assertions guarding invariants of a fixed-size array.
The loop is bounded by the live shadow NBuffers, matching the shape of
pg_buffercache_pages(): a concurrent shrink exits early, a concurrent
expand may or may not include the newly added tail depending on
timing, both cases return a partial-but-consistent snapshot.
Trade-offs vs the old batch approach:
- One pg_numa_query_pages syscall per buffer instead of one big call
covering the whole pool. Measured overhead at default 128MB shared
buffers is ~3 ms out of ~36 ms for pg_buffercache_numa (+7-10%).
For include_numa=false there is no syscall path at all.
- Per-invocation memory drops from O(NBuffers) records (megabytes at
typical sizes, ~750 MB at 128 GB shared_buffers) to O(1) stack.
- Non-NUMA path becomes ~30 % faster because streaming rows into the
tuplestore is cheaper than building fctx->record[] and returning
per-call.
Switches the SRF from multi-call to materialized, matching the sibling
functions pg_buffercache_pages() and pg_buffercache_usage_counts() in
the same file. Deletes BufferCacheOsPagesRec and
BufferCacheOsPagesContext typedefs and all SRF_IS_FIRSTCALL /
SRF_RETURN_NEXT scaffolding.
Verified: 015_stress_pg_buffercache 8/8 pass, zero retryable errors,
no crashes.
---
contrib/pg_buffercache/pg_buffercache_pages.c | 341 ++++--------------
1 file changed, 77 insertions(+), 264 deletions(-)
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 7335ccae150..d5166df32fb 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -32,33 +32,18 @@
#define NUM_BUFFERCACHE_OS_PAGES_ELEM 3
+/*
+ * Upper bound on OS pages a single BLCKSZ buffer can overlap. With BLCKSZ
+ * up to 32 KB and os_page_size at least 4 KB that's at most 9; 16 is a safe
+ * cap so the per-iteration scratch arrays fit on the stack.
+ */
+#define MAX_PAGES_PER_BUFFER 16
+
PG_MODULE_MAGIC_EXT(
.name = "pg_buffercache",
.version = PG_VERSION
);
-/*
- * Record structure holding the to be exposed cache data for OS pages. This
- * structure is used by pg_buffercache_os_pages(), where NUMA information may
- * or may not be included.
- */
-typedef struct
-{
- uint32 bufferid;
- int64 page_num;
- int32 numa_node;
-} BufferCacheOsPagesRec;
-
-/*
- * Function context for data persisting over repeated calls.
- */
-typedef struct
-{
- TupleDesc tupdesc;
- bool include_numa;
- BufferCacheOsPagesRec *record;
-} BufferCacheOsPagesContext;
-
static TupleDesc build_buffercache_pages_tupledesc(int natts);
@@ -285,279 +270,107 @@ build_buffercache_pages_tupledesc(int natts)
static Datum
pg_buffercache_os_pages_internal(FunctionCallInfo fcinfo, bool include_numa)
{
- FuncCallContext *funcctx;
- MemoryContext oldcontext;
- BufferCacheOsPagesContext *fctx; /* User function context. */
- TupleDesc tupledesc;
- TupleDesc expected_tupledesc;
- HeapTuple tuple;
- Datum result;
-
- /*
- * TODO: This allocates memory using NBuffers which may change while this
- * function is executed. We need to change this function so that it
- * doesn't rely on NBuffers being static throughout the execution of this
- * function.
- */
-
- if (SRF_IS_FIRSTCALL())
- {
- int i,
- idx;
- Size os_page_size;
- int pages_per_buffer;
- int *os_page_status = NULL;
- uint64 os_page_count = 0;
- int max_entries;
- char *startptr,
- *endptr;
-
- /* If NUMA information is requested, initialize NUMA support. */
- if (include_numa && pg_numa_init() == -1)
- elog(ERROR, "libnuma initialization failed or NUMA is not supported on this platform");
-
- /*
- * The database block size and OS memory page size are unlikely to be
- * the same. The block size is 1-32KB, the memory page size depends on
- * platform. On x86 it's usually 4KB, on ARM it's 4KB or 64KB, but
- * there are also features like THP etc. Moreover, we don't quite know
- * how the pages and buffers "align" in memory - the buffers may be
- * shifted in some way, using more memory pages than necessary.
- *
- * So we need to be careful about mapping buffers to memory pages. We
- * calculate the maximum number of pages a buffer might use, so that
- * we allocate enough space for the entries. And then we count the
- * actual number of entries as we scan the buffers.
- *
- * This information is needed before calling move_pages() for NUMA
- * node id inquiry.
- */
- os_page_size = pg_get_shmem_pagesize();
-
- /*
- * The pages and block size is expected to be 2^k, so one divides the
- * other (we don't know in which direction). This does not say
- * anything about relative alignment of pages/buffers.
- */
- Assert((os_page_size % BLCKSZ == 0) || (BLCKSZ % os_page_size == 0));
-
- if (include_numa)
- {
- void **os_page_ptrs = NULL;
-
- /*
- * How many addresses we are going to query? Simply get the page
- * for the first buffer, and first page after the last buffer, and
- * count the pages from that.
- */
- startptr = (char *) TYPEALIGN_DOWN(os_page_size,
- BufferGetBlock(1));
- endptr = (char *) TYPEALIGN(os_page_size,
- (char *) BufferGetBlock(NBuffers) + BLCKSZ);
- os_page_count = (endptr - startptr) / os_page_size;
-
- /* Used to determine the NUMA node for all OS pages at once */
- os_page_ptrs = palloc0_array(void *, os_page_count);
- os_page_status = palloc_array(int, os_page_count);
-
- /*
- * Fill pointers for all the memory pages. This loop stores and
- * touches (if needed) addresses into os_page_ptrs[] as input to
- * one big move_pages(2) inquiry system call, as done in
- * pg_numa_query_pages().
- */
- idx = 0;
- for (char *ptr = startptr; ptr < endptr; ptr += os_page_size)
- {
- os_page_ptrs[idx++] = ptr;
-
- /* Only need to touch memory once per backend process lifetime */
- if (firstNumaTouch)
- pg_numa_touch_mem_if_required(ptr);
- }
-
- Assert(idx == os_page_count);
-
- elog(DEBUG1, "NUMA: NBuffers=%d os_page_count=" UINT64_FORMAT " "
- "os_page_size=%zu", NBuffers, os_page_count, os_page_size);
-
- /*
- * If we ever get 0xff back from kernel inquiry, then we probably
- * have bug in our buffers to OS page mapping code here.
- */
- memset(os_page_status, 0xff, sizeof(int) * os_page_count);
-
- /* Query NUMA status for all the pointers */
- if (pg_numa_query_pages(0, os_page_count, os_page_ptrs, os_page_status) == -1)
- elog(ERROR, "failed NUMA pages inquiry: %m");
- }
-
- /* Initialize the multi-call context, load entries about buffers */
-
- funcctx = SRF_FIRSTCALL_INIT();
-
- /* Switch context when allocating stuff to be used in later calls */
- oldcontext = MemoryContextSwitchTo(funcctx->multi_call_memory_ctx);
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+ Size os_page_size;
+ void *os_page_ptrs[MAX_PAGES_PER_BUFFER];
+ int os_page_status[MAX_PAGES_PER_BUFFER];
+ Datum values[NUM_BUFFERCACHE_OS_PAGES_ELEM];
+ bool nulls[NUM_BUFFERCACHE_OS_PAGES_ELEM];
+ char *startptr;
+ int i;
- /* Create a user function context for cross-call persistence */
- fctx = palloc_object(BufferCacheOsPagesContext);
+ InitMaterializedSRF(fcinfo, 0);
- if (get_call_result_type(fcinfo, NULL, &expected_tupledesc) != TYPEFUNC_COMPOSITE)
- elog(ERROR, "return type must be a row type");
+ if (include_numa && pg_numa_init() == -1)
+ elog(ERROR, "libnuma initialization failed or NUMA is not supported on this platform");
- if (expected_tupledesc->natts != NUM_BUFFERCACHE_OS_PAGES_ELEM)
- elog(ERROR, "incorrect number of output arguments");
+ os_page_size = pg_get_shmem_pagesize();
- /* Construct a tuple descriptor for the result rows. */
- tupledesc = CreateTemplateTupleDesc(expected_tupledesc->natts);
- TupleDescInitEntry(tupledesc, (AttrNumber) 1, "bufferid",
- INT4OID, -1, 0);
- TupleDescInitEntry(tupledesc, (AttrNumber) 2, "os_page_num",
- INT8OID, -1, 0);
- TupleDescInitEntry(tupledesc, (AttrNumber) 3, "numa_node",
- INT4OID, -1, 0);
+ Assert((os_page_size % BLCKSZ == 0) || (BLCKSZ % os_page_size == 0));
+ Assert(Max(1, BLCKSZ / os_page_size) + 1 <= MAX_PAGES_PER_BUFFER);
- TupleDescFinalize(tupledesc);
- fctx->tupdesc = BlessTupleDesc(tupledesc);
- fctx->include_numa = include_numa;
+ if (include_numa && firstNumaTouch)
+ elog(DEBUG1, "NUMA: page-faulting the buffercache for proper NUMA readouts");
- /*
- * Each buffer needs at least one entry, but it might be offset in
- * some way, and use one extra entry. So we allocate space for the
- * maximum number of entries we might need, and then count the exact
- * number as we're walking buffers. That way we can do it in one pass,
- * without reallocating memory.
- */
- pages_per_buffer = Max(1, BLCKSZ / os_page_size) + 1;
- max_entries = NBuffers * pages_per_buffer;
+ startptr = (char *) TYPEALIGN_DOWN(os_page_size,
+ (char *) BufferGetBlock(1));
- /* Allocate entries for BufferCacheOsPagesRec records. */
- fctx->record = (BufferCacheOsPagesRec *)
- MemoryContextAllocHuge(CurrentMemoryContext,
- sizeof(BufferCacheOsPagesRec) * max_entries);
+ /*
+ * We don't hold the partition locks, so we don't get a consistent
+ * snapshot across all buffers, but we do grab the buffer header locks,
+ * so the information of each buffer is self-consistent.
+ */
+ for (i = 0; i < NBuffers; i++)
+ {
+ char *buffptr = (char *) BufferGetBlock(i + 1);
+ char *startptr_buff = (char *) TYPEALIGN_DOWN(os_page_size,
+ buffptr);
+ char *endptr_buff = buffptr + BLCKSZ;
+ BufferDesc *bufHdr;
+ uint32 bufferid;
+ int32 page_num;
+ int num_pages_this_buffer = 0;
+ int j;
+ char *ptr;
- /* Return to original context when allocating transient memory */
- MemoryContextSwitchTo(oldcontext);
+ bufHdr = GetBufferDescriptor(i);
+ LockBufHdr(bufHdr);
+ bufferid = BufferDescriptorGetBuffer(bufHdr);
+ UnlockBufHdr(bufHdr);
- if (include_numa && firstNumaTouch)
- elog(DEBUG1, "NUMA: page-faulting the buffercache for proper NUMA readouts");
+ page_num = (startptr_buff - startptr) / os_page_size;
- /*
- * Scan through all the buffers, saving the relevant fields in the
- * fctx->record structure.
- *
- * We don't hold the partition locks, so we don't get a consistent
- * snapshot across all buffers, but we do grab the buffer header
- * locks, so the information of each buffer is self-consistent.
- */
- startptr = (char *) TYPEALIGN_DOWN(os_page_size, (char *) BufferGetBlock(1));
- idx = 0;
- for (i = 0; i < NBuffers; i++)
+ for (ptr = startptr_buff; ptr < endptr_buff; ptr += os_page_size)
{
- char *buffptr = (char *) BufferGetBlock(i + 1);
- BufferDesc *bufHdr;
- uint32 bufferid;
- int32 page_num;
- char *startptr_buff,
- *endptr_buff;
-
- bufHdr = GetBufferDescriptor(i);
+ os_page_ptrs[num_pages_this_buffer++] = ptr;
- /* Lock each buffer header before inspecting. */
- LockBufHdr(bufHdr);
- bufferid = BufferDescriptorGetBuffer(bufHdr);
- UnlockBufHdr(bufHdr);
-
- /* start of the first page of this buffer */
- startptr_buff = (char *) TYPEALIGN_DOWN(os_page_size, buffptr);
-
- /* end of the buffer (no need to align to memory page) */
- endptr_buff = buffptr + BLCKSZ;
-
- Assert(startptr_buff < endptr_buff);
-
- /* calculate ID of the first page for this buffer */
- page_num = (startptr_buff - startptr) / os_page_size;
-
- /* Add an entry for each OS page overlapping with this buffer. */
- for (char *ptr = startptr_buff; ptr < endptr_buff; ptr += os_page_size)
- {
- fctx->record[idx].bufferid = bufferid;
- fctx->record[idx].page_num = page_num;
- fctx->record[idx].numa_node = include_numa ? os_page_status[page_num] : -1;
-
- /* advance to the next entry/page */
- ++idx;
- ++page_num;
- }
-
- /*
- * Check for interrupts here, at the end of the loop, so that the
- * buffer index i remains valid till the next iteration.
- */
- CHECK_FOR_INTERRUPTS();
+ /* Only need to touch memory once per backend process lifetime */
+ if (include_numa && firstNumaTouch)
+ pg_numa_touch_mem_if_required(ptr);
}
- Assert(idx <= max_entries);
-
- if (include_numa)
- Assert(idx >= os_page_count);
-
- /* Set max calls and remember the user function context. */
- funcctx->max_calls = idx;
- funcctx->user_fctx = fctx;
-
- /* Remember this backend touched the pages (only relevant for NUMA) */
if (include_numa)
- firstNumaTouch = false;
- }
-
- funcctx = SRF_PERCALL_SETUP();
-
- /* Get the saved state */
- fctx = funcctx->user_fctx;
-
- if (funcctx->call_cntr < funcctx->max_calls)
- {
- uint32 i = funcctx->call_cntr;
- Datum values[NUM_BUFFERCACHE_OS_PAGES_ELEM];
- bool nulls[NUM_BUFFERCACHE_OS_PAGES_ELEM];
+ {
+ memset(os_page_status, 0xff, sizeof(int) * num_pages_this_buffer);
+ if (pg_numa_query_pages(0, num_pages_this_buffer,
+ os_page_ptrs, os_page_status) == -1)
+ elog(ERROR, "failed NUMA pages inquiry: %m");
+ }
- values[0] = Int32GetDatum(fctx->record[i].bufferid);
+ values[0] = Int32GetDatum(bufferid);
nulls[0] = false;
-
- values[1] = Int64GetDatum(fctx->record[i].page_num);
nulls[1] = false;
- if (fctx->include_numa)
+ for (j = 0; j < num_pages_this_buffer; j++)
{
- /* status is valid node number */
- if (fctx->record[i].numa_node >= 0)
+ values[1] = Int64GetDatum(page_num++);
+
+ if (include_numa && os_page_status[j] >= 0)
{
- values[2] = Int32GetDatum(fctx->record[i].numa_node);
+ values[2] = Int32GetDatum(os_page_status[j]);
nulls[2] = false;
}
else
{
- /* some kind of error (e.g. pages moved to swap) */
values[2] = (Datum) 0;
nulls[2] = true;
}
- }
- else
- {
- values[2] = (Datum) 0;
- nulls[2] = true;
- }
- /* Build and return the tuple. */
- tuple = heap_form_tuple(fctx->tupdesc, values, nulls);
- result = HeapTupleGetDatum(tuple);
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc,
+ values, nulls);
+ }
- SRF_RETURN_NEXT(funcctx, result);
+ /*
+ * Check for interrupts here, at the end of the loop, so that the
+ * buffer index i remains valid till the next iteration.
+ */
+ CHECK_FOR_INTERRUPTS();
}
- else
- SRF_RETURN_DONE(funcctx);
+
+ if (include_numa)
+ firstNumaTouch = false;
+
+ return (Datum) 0;
}
/*
[application/octet-stream] v20260917-0012-buffermgr-recompute-pin-limit.patch (3.0K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/14-v20260917-0012-buffermgr-recompute-pin-limit.patch)
download | inline diff:
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Date: Mon, 14 Sep 2026 15:33:01 +0000
Subject: [PATCH] buffermgr: recompute pin limit after resize
MaxProportionalPins is a per-backend pin limit derived from NBuffers at
InitBufferManagerAccess() time and never touched again. After a resize
changes NBuffers, the stale value under- or over-estimates the limit
for the rest of the backend's lifetime. Factor the computation out into
RecomputeMaxProportionalPins() and call it both at init and from
ProcessBarrierBufferPoolSize(), which is where every backend learns the
buffer pool size actually changed.
---
src/backend/storage/buffer/buf_resize.c | 1 +
src/backend/storage/buffer/bufmgr.c | 11 ++++++++++-
src/include/storage/bufmgr.h | 1 +
3 files changed, 12 insertions(+), 1 deletion(-)
diff --git a/src/backend/storage/buffer/buf_resize.c b/src/backend/storage/buffer/buf_resize.c
index 90f5fb1c71d..d8b5a0ad845 100644
--- a/src/backend/storage/buffer/buf_resize.c
+++ b/src/backend/storage/buffer/buf_resize.c
@@ -503,6 +503,7 @@ ProcessBarrierBufferPoolSize(void)
Assert(activeNBuffers == pg_atomic_read_u32(&BufferControl->activeNBuffers));
NBuffers = pg_atomic_read_u32(&BufferControl->currentNBuffers);
+ RecomputeMaxProportionalPins();
return true;
}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index bb436734585..92b245c6c43 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -275,6 +275,15 @@ static int PrivateRefCountEntryLast = -1;
static uint32 MaxProportionalPins;
+/*
+ * Recompute the per-backend pin limit after a buffer pool resize.
+ */
+void
+RecomputeMaxProportionalPins(void)
+{
+ MaxProportionalPins = NBuffers / (MaxBackends + NUM_AUXILIARY_PROCS);
+}
+
static void ReservePrivateRefCountEntry(void);
static PrivateRefCountEntry *NewPrivateRefCountEntry(Buffer buffer);
static PrivateRefCountEntry *GetPrivateRefCountEntry(Buffer buffer, bool do_move);
@@ -4328,7 +4337,7 @@ InitBufferManagerAccess(void)
* allow plenty of pins. LimitAdditionalPins() and
* GetAdditionalPinLimit() can be used to check the remaining balance.
*/
- MaxProportionalPins = NBuffers / (MaxBackends + NUM_AUXILIARY_PROCS);
+ RecomputeMaxProportionalPins();
memset(&PrivateRefCountArray, 0, sizeof(PrivateRefCountArray));
memset(&PrivateRefCountArrayKeys, 0, sizeof(PrivateRefCountArrayKeys));
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 187e9b29a9e..db04a260870 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -272,6 +272,7 @@ extern Buffer ExtendBufferedRel(BufferManagerRelation bmr,
ForkNumber forkNum,
BufferAccessStrategy strategy,
uint32 flags);
+extern void RecomputeMaxProportionalPins(void);
extern BlockNumber ExtendBufferedRelBy(BufferManagerRelation bmr,
ForkNumber fork,
BufferAccessStrategy strategy,
--
2.43.0
[application/octet-stream] v20260917-0014-buffermgr-fix-tagged-buffer-eviction.patch (9.4K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/15-v20260917-0014-buffermgr-fix-tagged-buffer-eviction.patch)
download | inline diff:
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Subject: [PATCH] buffermgr: fix tagged-buffer eviction during shrink
---
src/backend/storage/buffer/bufmgr.c | 28 +++--
src/test/buffermgr/meson.build | 1
src/test/buffermgr/t/003_resize_failures.pl | 2
src/test/buffermgr/t/006_evict_mid_io_race.pl | 153 +++++++++++++++++++++++++
4 files changed, 175 insertions(+), 9 deletions(-)
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2163,6 +2163,13 @@ AsyncReadBuffers(ReadBuffersOperation *operation, int *nblocks_progress)
pgaio_io_set_flag(ioh, ioh_flags);
+ /*
+ * Test hook: stall after BM_TAG_VALID is set but before the read
+ * issues. With IOMETHOD_SYNC this is the only window where the
+ * buffer is visible (BM_TAG_VALID) but not yet valid (BM_VALID).
+ */
+ INJECTION_POINT("start-read-buffers-before-readv", NULL);
+
/* ---
* Even though we're trying to issue IO asynchronously, track the time
* in smgrstartreadv():
@@ -9152,26 +9159,31 @@ EvictExtraBuffers(int targetNBuffers, int currentNBuffers)
buf_state = pg_atomic_read_u64(&desc->state);
/*
- * Nobody is expected to allocate new buffers while resizing is going
- * on hence unlocked precheck should be safe and saves some cycles.
+ * A tagged buffer must be evicted even if its data is not yet
+ * valid. Skipping it could leave a mapping to a removed buffer.
*/
- if (!(buf_state & BM_VALID))
+ if (!(buf_state & BM_TAG_VALID))
continue;
ResourceOwnerEnlarge(CurrentResourceOwner);
ReservePrivateRefCountEntry();
- LockBufHdr(desc);
+ buf_state = LockBufHdr(desc);
/*
- * Now that we have locked buffer descriptor, make sure that the
- * buffer without valid data has been skipped above.
+ * Concurrent invalidation may have cleared the tag since the
+ * unlocked precheck.
*/
- Assert(buf_state & BM_VALID);
+ if (!(buf_state & BM_TAG_VALID))
+ {
+ UnlockBufHdr(desc);
+ continue;
+ }
if (!EvictUnpinnedBufferInternal(desc, &buffer_flushed))
{
- elog(WARNING, "could not remove buffer %u, it is pinned", buf);
+ ereport(WARNING,
+ (errmsg("could not evict buffer %u", buf)));
result = false;
break;
}
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
--- a/src/test/buffermgr/meson.build
+++ b/src/test/buffermgr/meson.build
@@ -25,6 +25,7 @@ tests += {
't/003_resize_failures.pl',
't/004_resize_with_syslogger.pl',
't/005_resize_unsupported.pl',
+ 't/006_evict_mid_io_race.pl',
't/010_stress_resize_buffer.pl',
't/011_stress_drop_relation_buffers.pl',
't/012_stress_drop_database_buffers.pl',
diff --git a/src/test/buffermgr/t/003_resize_failures.pl b/src/test/buffermgr/t/003_resize_failures.pl
--- a/src/test/buffermgr/t/003_resize_failures.pl
+++ b/src/test/buffermgr/t/003_resize_failures.pl
@@ -83,7 +83,7 @@ my $log_offset = -s $node->logfile;
is($node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()"),
'f',
"shrink returns false when a buffer to be evicted is pinned");
-ok($node->log_contains(qr/could not remove buffer $pinned_buf, it is pinned/, $log_offset),
+ok($node->log_contains(qr/could not evict buffer $pinned_buf\b/, $log_offset),
"log reports the pinned buffer that blocked eviction");
ok($node->log_contains(qr/failed to evict extra buffers during shrinking/, $log_offset),
"log reports the eviction failure");
diff --git a/src/test/buffermgr/t/006_evict_mid_io_race.pl b/src/test/buffermgr/t/006_evict_mid_io_race.pl
new file mode 100644
--- /dev/null
+++ b/src/test/buffermgr/t/006_evict_mid_io_race.pl
@@ -0,0 +1,153 @@
+# Copyright (c) 2025-2026, PostgreSQL Global Development Group
+#
+# Check that shrink rolls back while a tagged buffer above the target has
+# an unfinished read, then succeeds after the reader completes.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+if (!$ENV{enable_injection_points} || $ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => "test requires injection points";
+}
+
+my $initial_nbuffers = 64;
+my $grown_nbuffers = 128;
+my $max_nbuffers = 256;
+
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf',
+ 'shared_preload_libraries = injection_points');
+$node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
+$node->append_conf('postgresql.conf', "max_shared_buffers = $max_nbuffers");
+$node->append_conf('postgresql.conf', 'io_method = sync');
+$node->append_conf('postgresql.conf', 'huge_pages = off');
+$node->start;
+
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => "resizable shared memory not supported by this build";
+}
+
+$node->safe_psql('postgres', "CREATE EXTENSION pg_buffercache");
+$node->safe_psql('postgres', "CREATE EXTENSION injection_points");
+
+$node->safe_psql('postgres',
+ "CREATE TABLE evict_race AS SELECT generate_series(1,2) AS i");
+$node->safe_psql('postgres', "CHECKPOINT");
+
+# Grow the pool so we have buffers above the initial size.
+$node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$grown_nbuffers'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+$node->safe_psql('postgres', "SELECT pg_resize_shared_buffers()");
+
+my $current = $node->safe_psql('postgres',
+ "SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+is($current, "$grown_nbuffers", "pool grown to $grown_nbuffers");
+
+my $reader = $node->background_psql('postgres');
+my $reader_pid = $reader->query_safe("SELECT pg_backend_pid()", verbose => 0);
+chomp $reader_pid;
+
+# Warm catalogs before attaching the injection point so that the reader waits
+# while reading the table, not a catalog.
+$reader->query_safe("SELECT * FROM evict_race", verbose => 0);
+
+$reader->query_safe("SELECT injection_points_set_local()", verbose => 0);
+$reader->query_safe(
+ "SELECT injection_points_attach('start-read-buffers-before-readv', 'wait')",
+ verbose => 0);
+
+my $injector = $node->background_psql('postgres');
+
+# Evict the victim relation's buffers.
+$node->safe_psql('postgres', qq{
+ SELECT count(*) FROM pg_buffercache b,
+ LATERAL pg_buffercache_evict(b.bufferid) e
+ WHERE b.relfilenode = (SELECT relfilenode FROM pg_class
+ WHERE relname = 'evict_race')
+});
+
+# Evict low-numbered buffers; check the actual victim slot below.
+$node->safe_psql('postgres', qq{
+ SELECT count(*) FROM pg_buffercache b,
+ LATERAL pg_buffercache_evict(b.bufferid) e
+ WHERE b.bufferid <= $initial_nbuffers
+ AND b.relfilenode IS NOT NULL
+});
+
+# Wait before issuing the table read, with BM_TAG_VALID set and BM_VALID clear.
+$reader->query_until(qr/READER_STARTED/,
+ "\\echo READER_STARTED\nSELECT * FROM evict_race;\n");
+
+$node->poll_query_until('postgres', qq{
+ SELECT EXISTS (
+ SELECT 1
+ FROM pg_stat_activity
+ WHERE pid = $reader_pid
+ AND wait_event = 'start-read-buffers-before-readv')
+})
+ or die "reader never reached start-read-buffers-before-readv";
+
+my $victim_bufid = $node->safe_psql('postgres', q{
+ SELECT b.bufferid + 1
+ FROM pg_buffercache_lookup_table b JOIN pg_database d
+ ON b.database = d.oid AND b.tablespace = d.dattablespace
+ WHERE d.datname = current_database()
+ AND b.relfilenode = pg_relation_filenode('evict_race')
+ AND b.forknum = 0 AND b.blocknum = 0
+});
+like($victim_bufid, qr/^[0-9]+$/, 'one mapping for the victim page');
+BAIL_OUT('victim mapping missing or ambiguous')
+ unless $victim_bufid =~ /^[0-9]+$/;
+cmp_ok($victim_bufid, '>', $initial_nbuffers,
+ 'victim is above the shrink target');
+BAIL_OUT('victim is outside the range being removed')
+ unless $victim_bufid > $initial_nbuffers;
+
+# Shrink back to original size.
+$node->safe_psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '$initial_nbuffers'");
+$node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
+my $log_offset = -s $node->logfile;
+my $shrink_result = $node->safe_psql('postgres',
+ 'SELECT pg_resize_shared_buffers()');
+
+is($shrink_result, 'f', 'shrink fails during the read');
+BAIL_OUT('shrink did not roll back') unless $shrink_result eq 'f';
+ok($node->log_contains(
+ qr/could not evict buffer \Q$victim_bufid\E\b/, $log_offset),
+ 'shrink reports the victim buffer');
+is($node->safe_psql('postgres', q{
+ SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid
+ FROM pg_get_buffer_resize_status()
+}), "$grown_nbuffers|$grown_nbuffers|$grown_nbuffers|0",
+ 'failed shrink restores the resize state');
+
+$injector->query_safe(
+ "SELECT injection_points_detach('start-read-buffers-before-readv')",
+ verbose => 0);
+$injector->query_safe(
+ "SELECT injection_points_wakeup('start-read-buffers-before-readv')",
+ verbose => 0);
+is($reader->query_safe(''), "1\n2", 'reader completes successfully');
+$reader->quit;
+$injector->quit;
+
+is($node->safe_psql('postgres', 'SELECT pg_resize_shared_buffers()'),
+ 't', 'shrink succeeds after the read completes');
+is($node->safe_psql('postgres', q{
+ SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid
+ FROM pg_get_buffer_resize_status()
+}), "$initial_nbuffers|$initial_nbuffers|$initial_nbuffers|0",
+ 'pool shrunk to target and resizer released');
+is($node->safe_psql('postgres', 'SELECT count(*) FROM evict_race'),
+ '2', 'table remains readable after shrink');
+
+done_testing();
[application/octet-stream] v20260917-0015-buffermgr-fix-worker-guc-reset.patch (6.5K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/16-v20260917-0015-buffermgr-fix-worker-guc-reset.patch)
download | inline diff:
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Subject: [PATCH] buffermgr: allow worker GUC initialization resets
Parallel workers reset GUCs to their boot values before restoring the
leader's settings. When max_shared_buffers is below the shared_buffers
boot value, the check hook rejects this temporary reset and worker startup
fails. Allow it while InitializingParallelWorker is set.
Do not exempt other PGC_S_DEFAULT assignments: configuration-file removal
also uses that source and must not install a target above MaxNBuffers.
Correct the check-hook detail to reflect that equality with the maximum
is allowed.
Add a TAP test that launches parallel workers with a 32MB pool, checks an
explicit oversized setting, and removes the last configured value. The
reset must be rejected and a subsequent resize must leave the pool intact.
Base: 0ad260c6365b21708037db9b863bb488a0a521c0
Replaces consolidated patch 0005; does not require that patch first.
---
src/backend/storage/buffer/buf_init.c | 9 +++
src/test/buffermgr/expected/buffer_resize.out | 2 -
src/test/buffermgr/meson.build | 1
src/test/buffermgr/t/008_shared_buffers_guc.pl | 71 ++++++++++++++++++++++++
4 files changed, 81 insertions(+), 2 deletions(-)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 91671bce888..58d8f9029b6 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -14,6 +14,7 @@
*/
#include "postgres.h"
+#include "access/parallel.h"
#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
@@ -328,13 +329,19 @@ BufferManagerShmemResize(int currentNBuffers, int targetNBuffers)
*
* When reloading the configuration, shared_buffers should not be set to a value
* higher than max_shared_buffers fixed at the boot time.
+ *
+ * Parallel workers temporarily restore the boot value before applying the
+ * leader's settings. Other default resets must still respect MaxNBuffers.
*/
bool
check_shared_buffers(int *newval, void **extra, GucSource source)
{
+ if (source == PGC_S_DEFAULT && InitializingParallelWorker)
+ return true;
+
if (finalMaxNBuffers && *newval > MaxNBuffers)
{
- GUC_check_errdetail("\"shared_buffers\" must be less than \"max_shared_buffers\".");
+ GUC_check_errdetail("\"shared_buffers\" must not exceed \"max_shared_buffers\".");
return false;
}
return true;
diff --git a/src/test/buffermgr/expected/buffer_resize.out b/src/test/buffermgr/expected/buffer_resize.out
index 1dc3efd830a..2443673e706 100644
--- a/src/test/buffermgr/expected/buffer_resize.out
+++ b/src/test/buffermgr/expected/buffer_resize.out
@@ -264,7 +264,7 @@ SELECT COUNT(*) AS buffer_count FROM pg_buffercache;
-- Test 6: Try to set shared_buffers higher than max_shared_buffers (should fail)
ALTER SYSTEM SET shared_buffers = '400MB';
ERROR: invalid value for parameter "shared_buffers": 51200
-DETAIL: "shared_buffers" must be less than "max_shared_buffers".
+DETAIL: "shared_buffers" must not exceed "max_shared_buffers".
SELECT pg_reload_conf();
pg_reload_conf
----------------
diff --git a/src/test/buffermgr/meson.build b/src/test/buffermgr/meson.build
index 826e1094ab3..4db62d185ce 100644
--- a/src/test/buffermgr/meson.build
+++ b/src/test/buffermgr/meson.build
@@ -26,6 +26,7 @@ tests += {
't/004_resize_with_syslogger.pl',
't/005_resize_unsupported.pl',
't/006_evict_mid_io_race.pl',
+ 't/008_shared_buffers_guc.pl',
't/010_stress_resize_buffer.pl',
't/011_stress_drop_relation_buffers.pl',
't/012_stress_drop_database_buffers.pl',
diff --git a/src/test/buffermgr/t/008_shared_buffers_guc.pl b/src/test/buffermgr/t/008_shared_buffers_guc.pl
new file mode 100644
index 00000000000..22ced7ba41e
--- /dev/null
+++ b/src/test/buffermgr/t/008_shared_buffers_guc.pl
@@ -0,0 +1,71 @@
+# Test shared_buffers validation during worker startup and configuration reset.
+
+use strict;
+use warnings;
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf('postgresql.conf', qq{
+shared_buffers = '32MB'
+huge_pages = off
+max_worker_processes = 4
+max_parallel_workers = 2
+restart_after_crash = off
+});
+$node->start;
+
+if ($node->safe_psql('postgres', 'SHOW have_resizable_shmem') ne 'on')
+{
+ plan skip_all => 'resizable shared memory not supported by this build';
+}
+
+is($node->safe_psql('postgres', 'SHOW max_shared_buffers'), '32MB',
+ 'maximum follows the initial shared_buffers setting');
+my $initial_nbuffers = $node->safe_psql('postgres',
+ 'SELECT current_nbuffers FROM pg_get_buffer_resize_status()');
+
+$node->safe_psql('postgres', q{
+ CREATE TABLE parallel_scan AS SELECT generate_series(1, 100000) AS value;
+ ALTER TABLE parallel_scan SET (parallel_workers = 2);
+ ANALYZE parallel_scan;
+});
+my $plan = $node->safe_psql('postgres', q{
+ SET max_parallel_workers_per_gather = 2;
+ SET min_parallel_table_scan_size = 0;
+ SET parallel_setup_cost = 0;
+ SET parallel_tuple_cost = 0;
+ EXPLAIN (ANALYZE, COSTS OFF, TIMING OFF)
+ SELECT count(*) FROM parallel_scan;
+});
+like($plan, qr/Workers Launched: 2/, 'parallel workers restore GUC state');
+
+$node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '32MB'");
+my ($ret, $stdout, $stderr) = $node->psql('postgres',
+ "ALTER SYSTEM SET shared_buffers = '64MB'");
+isnt($ret, 0, 'explicit value above the maximum is rejected');
+like($stderr, qr/must not exceed/, 'error allows equality with the maximum');
+
+$node->adjust_conf('postgresql.conf', 'shared_buffers', undef);
+$node->safe_psql('postgres', 'ALTER SYSTEM RESET shared_buffers');
+my $log_offset = -s $node->logfile;
+$node->reload;
+$node->wait_for_log(qr/"shared_buffers" must not exceed/, $log_offset);
+
+is($node->safe_psql('postgres', 'SHOW shared_buffers'), '32MB',
+ 'configuration removal does not install an oversized default');
+is($node->safe_psql('postgres', 'SELECT pg_resize_shared_buffers()'), 't',
+ 'resize after rejected default reset is a successful no-op');
+is($node->safe_psql('postgres', q{
+ SELECT active_nbuffers, current_nbuffers, target_nbuffers, resizer_pid
+ FROM pg_get_buffer_resize_status();
+}), "$initial_nbuffers|$initial_nbuffers|$initial_nbuffers|0",
+ 'buffer pool remains unchanged');
+is($node->safe_psql('postgres', 'SELECT count(*) FROM parallel_scan'),
+ '100000', 'table remains readable');
+$node->log_check('no crash during configuration reset', $log_offset,
+ log_unlike => [qr/PANIC|terminated by signal/]);
+
+done_testing();
\ No newline at end of file
[application/octet-stream] v20260917-0013-shmem-limit-madvise-ranges.patch (4.4K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/17-v20260917-0013-shmem-limit-madvise-ranges.patch)
download | inline diff:
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Date: Thu, 17 Sep 2026 00:00:00 +0000
Subject: [PATCH] shmem: limit madvise calls to resized ranges
When shrinking, stop MADV_REMOVE at the current page-aligned end rather
than extending into the protected reserved tail. Kernels that require
the mapping to be currently writable reject the latter with EACCES,
which causes buffer pool resizing to PANIC. Retain the maximum-size
bound to preserve any page shared with the next structure.
When growing, pass only the newly added page-aligned range to
MADV_POPULATE_WRITE. This avoids prefaulting the existing range again
for small increases in a large structure.
Update configure with the MADV_REMOVE and MADV_POPULATE_WRITE declaration
checks already present in configure.ac.
Reported-by: Yuhang Qiu <iamqyh@gmail.com>
Discussion: https://postgr.es/m/78DD860A-DD0E-4B70-A0CA-EE5CCA3E0E60@gmail.com
Discussion: https://postgr.es/m/8E7D0939-ADE1-405A-91D2-0139A6E322B3@gmail.com
---
configure | 26 ++++++++++++++++++++++++++
src/backend/storage/ipc/shmem.c | 18 ++++++++++--------
2 files changed, 36 insertions(+), 8 deletions(-)
diff --git a/configure b/configure
index d42a7a794ff..6e61394a81d 100755
--- a/configure
+++ b/configure
@@ -16414,6 +16414,32 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux-specific madvise constants needed for resizable shared memory. See similar checks in meson.build for explanation of why these checks are here.
+ac_fn_c_check_decl "$LINENO" "MADV_POPULATE_WRITE" "ac_cv_have_decl_MADV_POPULATE_WRITE" "#include <sys/mman.h>
+"
+if test "x$ac_cv_have_decl_MADV_POPULATE_WRITE" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_MADV_POPULATE_WRITE $ac_have_decl
+_ACEOF
+
+ac_fn_c_check_decl "$LINENO" "MADV_REMOVE" "ac_cv_have_decl_MADV_REMOVE" "#include <sys/mman.h>
+"
+if test "x$ac_cv_have_decl_MADV_REMOVE" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_MADV_REMOVE $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c
index fbc05f0dc13..c4aa1660fca 100644
--- a/src/backend/storage/ipc/shmem.c
+++ b/src/backend/storage/ipc/shmem.c
@@ -917,15 +917,17 @@ ShmemResizeStruct(const char *name, Size new_size)
* When shrinking, release memory pages beyond the new end, but not the
* page containing maximal end of the structure, as it may be used by the
* next structure.
- *
- * We do not consider the current end of the structure as it simplifies
- * the calculations. Instead we rely on the underlying APIs not to touch
- * the memory pages that will not be affected by the change in size.
*/
new_end = (char *) TYPEALIGN(page_size, (char *) result->location + new_size);
if (new_size < result->size)
{
- char *max_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+ /*
+ * Stop at the current page-aligned end to avoid the protected tail,
+ * and preserve any page shared with the next structure.
+ */
+ char *current_end = (char *) TYPEALIGN(page_size, (char *) result->location + result->size);
+ char *reserved_end = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location + result->maximum_size);
+ char *max_end = Min(current_end, reserved_end);
if (max_end > new_end)
{
@@ -938,9 +940,9 @@ ShmemResizeStruct(const char *name, Size new_size)
}
else if (new_size > result->size)
{
- char *struct_start = (char *) TYPEALIGN_DOWN(page_size, (char *) result->location);
+ char *old_end = (char *) TYPEALIGN(page_size, (char *) result->location + result->size);
- if (new_end > struct_start)
+ if (new_end > old_end)
{
ShmemIndexEnt entry_copy = *result;
@@ -951,7 +953,7 @@ ShmemResizeStruct(const char *name, Size new_size)
entry_copy.size = new_size;
ShmemProtectStructInternal(&entry_copy);
- if (!PGSharedMemoryEnsureAllocated(struct_start, new_end - struct_start))
+ if (!PGSharedMemoryEnsureAllocated(old_end, new_end - old_end))
{
ShmemProtectStructInternal(result);
ereport(WARNING,
[application/octet-stream] v20260917-0016-doc-update-buffer-resize-todos.patch (3.4K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/18-v20260917-0016-doc-update-buffer-resize-todos.patch)
download | inline diff:
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Date: Mon, 14 Sep 2026 15:33:02 +0000
Subject: [PATCH] doc: update buffer resize documentation
- InitProcess()/InitProcessPhase2(): 'specifid' -> 'specified'
- func-admin.sgml: garbled paragraph in pg_resize_shared_buffers() docs
- config.sgml: quantify the memory max_shared_buffers reserves upfront
---
doc/src/sgml/config.sgml | 9 +++++----
doc/src/sgml/func/func-admin.sgml | 11 +++++------
src/backend/storage/lmgr/proc.c | 2 +-
3 files changed, 11 insertions(+), 11 deletions(-)
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 8418af072f9..031a82c04da 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -1844,9 +1844,6 @@ include_dir 'conf.d'
other GUCS which use this GUCs value to set their defaults will not be
changed. They may still require a server restart to consider new value.
</para>
-
- <para>
- </para>
</listitem>
</varlistentry>
@@ -1875,7 +1872,11 @@ include_dir 'conf.d'
memory required for the buffer lookup table and the array used to sort
buffers during a checkpoint is allocated at the server start
considering the largest buffer pool size allowed by this parameter.
- <!-- TODO: Provide a numeric example of how much extra memory say max_shared_buffers = 1GB consume. -->
+ For example, with the default block size of 8kB,
+ <literal>max_shared_buffers = 1GB</literal> allows for 131072 buffers,
+ whose lookup table and sort array together occupy about 9.5MB,
+ regardless of the current value of
+ <varname>shared_buffers</varname>.
</para>
<para>
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 2a621fff060..9d345c7ecaa 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -160,12 +160,11 @@ postgres=# SHOW shared_buffers; -- Step 5 (verification)
pending to be applied.
</para>
<para>
- <!-- TODO: Document behaviour when the function is called on platforms that do not support resizable shared memory -->
- linkend="functions-admin-signal-table"/> send control signals to
- other server processes. Use of these functions is restricted to
- superusers by default but access may be granted to others using
- <command>GRANT</command>, with noted exceptions.
- </para>
+ If <xref linkend="guc-have-resizable-shmem"/> is <literal>off</literal>,
+ this function reports an error instead of resizing the buffer pool, and
+ the pending value of <varname>shared_buffers</varname> takes effect only
+ after the server is restarted.
+ </para>
</entry>
</row>
</tbody>
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 8332bc0b252..92cfb37f660 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -798,7 +798,7 @@ InitAuxiliaryProcess(void)
* callback above.
*
* We also call this function after ProcSignalInit() for the reasons
- * specifid there, but we need it here so that InitBufferManagerAccess()
+ * specified there, but we need it here so that InitBufferManagerAccess()
* can use the current buffer pool size.
*/
BufferManagerInitProc();
--
2.43.0
[application/octet-stream] v20260917-0019-shmem-distinguish-mprotect-errors.patch (1.1K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/19-v20260917-0019-shmem-distinguish-mprotect-errors.patch)
download | inline diff:
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Subject: [PATCH] shmem: distinguish mprotect failure messages
Refresh the earlier follow-up for the complete series.
---
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index c052776e94c..179ceab57ff 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -1200,7 +1200,8 @@ PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
if (mprotect(rw_start, (char *) rw_end - (char *) rw_start,
PROT_READ | PROT_WRITE) != 0)
{
- ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ ereport(WARNING,
+ errmsg("could not make shared memory read-write: %m"));
return false;
}
}
@@ -1210,7 +1211,8 @@ PGSharedMemoryProtect(void *rw_start, void *rw_end, void *prot_end)
if (mprotect(rw_end, (char *) prot_end - (char *) rw_end,
PROT_NONE) != 0)
{
- ereport(WARNING, errmsg("could not protect shared memory: %m"));
+ ereport(WARNING,
+ errmsg("could not make reserved shared memory inaccessible: %m"));
return false;
}
}
[application/octet-stream] v20260917-0018-buffermgr-fix-resizer-reload-order.patch (1.3K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/20-v20260917-0018-buffermgr-fix-resizer-reload-order.patch)
download | inline diff:
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Subject: [PATCH] buffermgr: start the resizer after configuration reload
Refresh the earlier follow-up for the complete series.
---
diff --git a/src/test/buffermgr/t/003_resize_failures.pl b/src/test/buffermgr/t/003_resize_failures.pl
index aea172d323c..13a3f3af633 100644
--- a/src/test/buffermgr/t/003_resize_failures.pl
+++ b/src/test/buffermgr/t/003_resize_failures.pl
@@ -128,15 +128,15 @@ SKIP:
my $sizes_query = "SELECT name, size FROM pg_shmem_allocations WHERE name IN $resizable_structs ORDER BY name";
my $sizes_before = $node->safe_psql('postgres', $sizes_query);
+ my $expand_target = $max_nbuffers;
+ $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$expand_target'");
+ $node->safe_psql('postgres', "SELECT pg_reload_conf()");
+
my $resizer = $node->background_psql('postgres');
$resizer->query_safe("SELECT injection_points_set_local()", verbose => 0);
$resizer->query_safe("SELECT injection_points_attach('buffer-mgr-resize-struct-fail', 'notice')",
verbose => 0);
- my $expand_target = $max_nbuffers;
- $node->safe_psql('postgres', "ALTER SYSTEM SET shared_buffers = '$expand_target'");
- $node->safe_psql('postgres', "SELECT pg_reload_conf()");
-
my $expand_log_offset = -s $node->logfile;
is($resizer->query("SELECT pg_resize_shared_buffers()"), 'f',
[application/octet-stream] v20260917-0017-buffermgr-validate-hints-and-rollback-metadata.patch (7.0K, ../../CALfch18r6Fx2TLfNq9vPMLuuZpVO-GbxVRGNX=A+R8cW6-hhmw@mail.gmail.com/21-v20260917-0017-buffermgr-validate-hints-and-rollback-metadata.patch)
download | inline diff:
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Palak Chaturvedi <palakchaturvedi2843@gmail.com>
Date: Mon, 14 Sep 2026 15:33:02 +0000
Subject: [PATCH] buffermgr: validate buffer hints and rollback metadata
- ReadRecentBuffer(): range-check the recent_buffer hint, which can go
stale after a shrink
- 001_resize_fault_tolerance.pl: verify shmem structure sizes are
actually restored after a rolled-back resize, not just the NBuffers
counter
---
src/backend/access/transam/xlogutils.c | 2 +
src/backend/storage/buffer/bufmgr.c | 16 +++++------
src/test/buffermgr/t/001_resize_fault_tolerance.pl | 30 +++++++++++++++++---
3 files changed, 34 insertions(+), 14 deletions(-)
diff --git a/src/backend/access/transam/xlogutils.c b/src/backend/access/transam/xlogutils.c
index 58b9dab6a90..ed5559b6624 100644
--- a/src/backend/access/transam/xlogutils.c
+++ b/src/backend/access/transam/xlogutils.c
@@ -491,8 +491,8 @@ XLogReadBufferExtended(RelFileLocator rlocator, ForkNumber forknum,
Assert(blkno != P_NEW);
/* Do we have a clue where the buffer might be already? */
- if (BufferIsValid(recent_buffer) &&
+ if (recent_buffer != InvalidBuffer &&
mode == RBM_NORMAL &&
ReadRecentBuffer(rlocator, forknum, blkno, recent_buffer))
{
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index d51d1f72392..8cc5a6316a7 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -828,14 +828,9 @@ PrefetchBuffer(Relation reln, ForkNumber forkNum, BlockNumber blockNum)
* successful. Return true if the buffer is valid and still has the expected
* tag. In that case, the buffer is pinned and the usage count is bumped.
*
- * The callers of this function should make sure that the buffer is valid even
- * if the shared buffer pool has undergone a resize. Buffer pool resizing waits
- * for all backends to acknowledge the barrier before changing the buffer pool
- * size. Hence the caller should call this function after validation without an
- * intervening ProcSignalBarrier processing.
- *
- * TODO: This function could perform the validation in this function itself
- * instead of relying on the two callers who do it currently.
+ * A shrink can leave recent_buffer outside the current pool. Reject such
+ * hints before accessing the descriptor. Do not process resize barriers
+ * between this check and acquiring the pin.
*/
bool
ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockNum,
@@ -845,7 +840,10 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
BufferTag tag;
uint64 buf_state;
- Assert(BufferIsValid(recent_buffer));
+ if (recent_buffer == InvalidBuffer ||
+ recent_buffer > NBuffers ||
+ recent_buffer < -NLocBuffer)
+ return false;
ResourceOwnerEnlarge(CurrentResourceOwner);
ReservePrivateRefCountEntry();
diff --git a/src/test/buffermgr/t/001_resize_fault_tolerance.pl b/src/test/buffermgr/t/001_resize_fault_tolerance.pl
index 87c95dbedf1..086fdc1b702 100644
--- a/src/test/buffermgr/t/001_resize_fault_tolerance.pl
+++ b/src/test/buffermgr/t/001_resize_fault_tolerance.pl
@@ -27,6 +27,7 @@
$node->append_conf('postgresql.conf', "shared_buffers = $initial_nbuffers");
$node->append_conf('postgresql.conf', 'max_shared_buffers = 32');
+$node->append_conf('postgresql.conf', 'huge_pages = off');
$node->append_conf('postgresql.conf', 'restart_after_crash = on');
$node->start;
# Bail out if this build does not support resizable shared memory, which
@@ -43,6 +44,21 @@ $node->safe_psql('postgres', "CREATE EXTENSION injection_points");
# Helper functions
# =============================================================================
+# Snapshot the buffer manager's allocation metadata.
+sub buffer_shmem_allocations
+{
+ my $allocations = $node->safe_psql('postgres', q{
+ SELECT string_agg(name || '=' || allocated_size, ', ' ORDER BY name)
+ FROM pg_shmem_allocations
+ WHERE name IN ('Buffer Blocks', 'Buffer Descriptors',
+ 'Buffer IO Condition Variables', 'Checkpoint BufferIds',
+ 'Shared Buffer Lookup Table')
+ HAVING count(*) = 5
+ });
+ BAIL_OUT('missing buffer manager allocation metadata') if $allocations eq '';
+ return $allocations;
+}
+
# Setup resize operation to be interrupted.
#
# Prepare to resize the buffer pool to a target size. Start a resize session
@@ -191,11 +204,13 @@ sub interrupt_resize_session
# was used to interrupt the resize operation.
# - orig_nbuffers and target_nbuffers: the original and target buffer sizes for
# the resize operation.
+# - orig_allocations: buffer_shmem_allocations() taken before the resize began.
# - test_label: a label to create unique test names for different tests
sub check_interrupted_resize
{
my ($sentinel_session, $resize_session, $log_offset, $mode,
- $injection_point, $orig_nbuffers, $target_nbuffers, $test_label) = @_;
+ $injection_point, $orig_nbuffers, $target_nbuffers, $orig_allocations,
+ $test_label) = @_;
my $resize_pid = $resize_session->{backend_pid};
my $sentinel_pid = $sentinel_session->{backend_pid};
@@ -269,7 +284,8 @@ sub check_interrupted_resize
"$orig_nbuffers|$orig_nbuffers|$orig_nbuffers|0",
"$test_label: buffer resize rolled back after $mode");
- # TODO: Also check that the pg_shmem_allocations values are not changed
+ is(buffer_shmem_allocations(), $orig_allocations,
+ "$test_label: shared memory allocations unchanged after $mode");
is($node->safe_psql('postgres',
"SELECT setting FROM pg_settings WHERE name = 'shared_buffers'"),
@@ -386,6 +402,7 @@ sub test_interrupt_resize_at_injection_point
my $orig_nbuffers = $node->safe_psql('postgres',
"SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+ my $orig_allocations = buffer_shmem_allocations();
my $log_offset = -s $node->logfile;
# Start a sentinel session that will be used to detect whether the
@@ -400,7 +417,7 @@ sub test_interrupt_resize_at_injection_point
check_interrupted_resize($sentinel_session, $resize_session, $log_offset,
$mode, $injection_point, $orig_nbuffers, $target_nbuffers,
- $test_label);
+ $orig_allocations, $test_label);
}
# Driver function:
@@ -563,6 +580,7 @@ sub test_fault_resize_waiting_barrier
my $orig_nbuffers = $node->safe_psql('postgres',
"SELECT current_nbuffers FROM pg_get_buffer_resize_status()");
+ my $orig_allocations = buffer_shmem_allocations();
my $log_offset = -s $node->logfile;
my $peer_session = start_peer_session_with_injection_point($injection_point, 'wait');
@@ -587,7 +605,8 @@ sub test_fault_resize_waiting_barrier
interrupt_resize_session($mode, $resize_session);
check_interrupted_resize($sentinel_session, $resize_session, $log_offset,
- $mode, $injection_point, $orig_nbuffers, $target_nbuffers, $test_label);
+ $mode, $injection_point, $orig_nbuffers, $target_nbuffers,
+ $orig_allocations, $test_label);
# Cleanup peer session. If the postmaster restarted all backends, the peer
# backend is already gone.
--
2.43.0
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-07 13:00 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-08 03:09 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-08 11:24 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-11 10:56 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-17 16:02 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
@ 2026-09-22 03:45 ` Yuhang Qiu <iamqyh@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Yuhang Qiu @ 2026-09-22 03:45 UTC (permalink / raw)
To: Palak Chaturvedi <chaturvedipalak1911@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi Palak,
Thanks for the update. 0013 and 0015 LGTM.
Two more points:
- pg_resize_shared_buffers() lacks a permission check, although the
documentation says it is restricted to superusers.
- In test 006, evicting the low-numbered buffers leaves those slots
available for reuse, so the victim may still be below the shrink target.
The test can therefore bail out before attempting the shrink. What would
be a reliable way to arrange an in-progress read above the target?
Best regards,
Yuhang Qiu
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Re: Changing shared_buffers without restart Andres Freund <andres@anarazel.de>
2025-10-13 15:58 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-11-14 11:53 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Re: Changing shared_buffers without restart Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-20 11:13 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-07 13:00 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-08 03:09 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-08 11:24 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-11 10:56 ` Re: Changing shared_buffers without restart Yuhang Qiu <iamqyh@gmail.com>
2026-09-17 16:02 ` Re: Changing shared_buffers without restart Palak Chaturvedi <chaturvedipalak1911@gmail.com>
@ 2026-09-22 07:02 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 0 replies; 167+ messages in thread
From: Ashutosh Bapat @ 2026-09-22 07:02 UTC (permalink / raw)
To: Palak Chaturvedi <chaturvedipalak1911@gmail.com>; +Cc: Yuhang Qiu <iamqyh@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <heikki.linnakangas@databricks.com>; Haoyu Huang <haoyu.huang@databricks.com>; Tomas Vondra <tomas@vondra.me>; Peter Eisentraut <peter@eisentraut.org>; Thomas Munro <thomas.munro@gmail.com>; Dmitry Dolgov <9erthalion6@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Andres Freund <andres@anarazel.de>; Jakub Wartak <jakub.wartak@enterprisedb.com>
On Thu, Sep 17, 2026 at 9:32 PM Palak Chaturvedi
<chaturvedipalak1911@gmail.com> wrote:
> > I also tested the patches and did some more review.
> >
> > With shared_buffers=32MB and max_shared_buffers left at its default, parallel
> > workers fail to start, even without any resize:
> > FATAL: failed to initialize shared_buffers to 16384
> > CONTEXT: parallel worker
> >
> > RestoreGUCState() resets the GUC to its 128MB boot value, which exceeds
> > MaxNBuffers and fails the check hook.
>
> Reproduced this with shared_buffers=32MB and max_shared_buffers left at
> its default. Fixed in 0015 by allowing the temporary PGC_S_DEFAULT
> assignment only while InitializingParallelWorker is set.
>
> Also checked removing the last configured value. Allowing PGC_S_DEFAULT
> unconditionally accepts the 128MB default above the reserved 32MB
> maximum, followed by a PANIC on resize. The revised check rejects this
> reset.
>
> Added test 008 for both cases. All nine assertions pass, including
> launching two workers. Without the fix, it fails with the reported
> initialization error. It also fails with the unconditional
> default-source exception.
>
In the latest patchset this is 0015. I am wondering why does the
finalMaxNBuffers guard is not working OR why max_shared_buffers isn't
being reset. If max_shared_buffers also gets reset, it will default to
shared_buffers and the check will pass. This needs more investigation.
Please note that this is noted at
src/test/buffermgr/t/002_client_join_buffer_resize.pl:89, so a
separate test may not be needed.
> > There is also a performance issue in the grow path:
> > PGSharedMemoryEnsureAllocated() gets the whole structure's range, rather than
> > just the added range. Even a small increase makes MADV_POPULATE_WRITE walk the
> > existing range again, adding overhead for large buffer pools. Could we limit
> > this to the page-aligned [current_end, new_end), as in the shrink path?
>
> Changed this in 0013 to populate only the newly added pages. Also did
> the testing. For a 512MB to 528MB grow, seven untraced runs gave a median
> resize time of 332.7ms with the old range and 11.4ms with the new range.
> Repeated the optimized run after the control and got 11.4ms again.
>
> This was on Linux 6.17 with huge_pages=off, no swap and no concurrent
> workload. These are total resize times for this case. Existing pages
> that have been swapped out will be brought back when accessed. I have
> not measured the effect under memory pressure with swap enabled.
>
We mmap the shared memory with MAP_NORESERVE so swap is not reserved
for this segment. From your description I can't figure out whether the
performance problem is real or not. The reason it's coded like that is
purely to simplify code - it avoids tracking the current allocation
boundary and relies only on the current size of the structure. We
should fix it if there's visible performance problem. I think there's
also some hazard if we try to remove pages from protected range. That
may require this patch irrespective of the performance problem. But I
don't see that being mentioned in the commit message.
>
> Done:
>
> - 0011: Included the earlier pg_buffercache fix to process one buffer
> at a time in pg_buffercache_os_pages.
I will review this in detail, but the idea to process one buffer at a
time looks promising. If we could avoid performance degradation in the
NUMA path that would be great.
>
> - 0012: Recompute MaxProportionalPins when a backend processes a resize
> barrier, so the pin limit follows the new pool size.
This looks ok too. Recompute in the function name does look misleading
since the same function is used to compute MaxProportionalPins the
first time and later. Did you check whether changing the value of
MaxProportionalPins in a barrier is safe?
>
> - 0013: Retained the corrected MADV_REMOVE shrink range and included
> the grow-range change discussed above. Also included the configure
> declaration checks for the madvise constants.
configure.ac has these changes. Do we need to include the
corresponding changes to configure in the patch?
>
> - 0014: Check BM_TAG_VALID during eviction. A buffer can already have
> a lookup-table entry while its read is still in progress, so checking
> only whether its contents are valid can miss it. Recheck under the
> header lock and roll back if eviction fails. Test 006 checks rollback,
> reader completion and successful retry.
> 003 already covers the fully-valid pinned-buffer case.
EvictExtraBuffers is inspired from other Evict*Unpinned* functions in
that file. Those functions do not check for this flag? Do we need to
handle it there as well?
>
> - 0016: Covered the TODO to document unsupported platforms. The
> documentation now says that the resize function reports an error
> when resizable shared memory is unavailable. Also repaired the
> garbled paragraph and the typo.
> See doc/src/sgml/func/func-admin.sgml:163.
- <!-- TODO: Document behaviour when the function is called on
platforms that do not support resizable shared memory -->
- linkend="functions-admin-signal-table"/> send control signals to
- other server processes. Use of these functions is restricted to
- superusers by default but access may be granted to others using
- <command>GRANT</command>, with noted exceptions.
- </para>
+ If <xref linkend="guc-have-resizable-shmem"/> is
<literal>off</literal>,
+ this function reports an error instead of resizing the buffer pool, and
+ the pending value of <varname>shared_buffers</varname> takes
effect only
+ after the server is restarted.
+ </para>
Have you accidentally removed some other function's description?
>
> - 0016: Covered the numeric memory-overhead example TODO. Verified with
> max_shared_buffers=1GB and 8kB blocks on the 64-bit build. The lookup
> table and checkpoint sort array together allocate 9,976,780 bytes,
> about 9.5MB. Got the same result with shared_buffers=32MB and 128MB.
> See doc/src/sgml/config.sgml:1875.
Thanks.
>
> - 0017: Covered the TODO to validate hints inside ReadRecentBuffer().
> A buffer number saved before shrink may be outside the new pool.
> The function now rejects it before accessing the descriptor,
> allowing the caller to do a normal lookup.
> See src/backend/storage/buffer/bufmgr.c:843.
XLogReadBufferExtended() still checks for InvalidBuffer and the other
caller invalidate_one_block remains unchanged. Am I missing something.
>
> - 0017: Covered the TODO to check pg_shmem_allocations after rollback.
> Test 001 now compares allocation metadata for all five buffer-manager
> structures instead of checking only buffer counts. Also disabled
> huge pages so rounding does not hide the small allocation changes.
> All 414 assertions pass. This checks allocation metadata, not whether
> physical memory was released.
> See src/test/buffermgr/t/001_resize_fault_tolerance.pl:290.
> The snapshot helper is at line 48; huge_pages=off is at line 29.
>
Even with huge_pages = off, the previous allocations should match
after resize rolls back, right?
> - 0018: Included the earlier SIGHUP test-ordering fix from
> v20260817-0009. Start the resizer after configuration reload so the
> expansion-failure test does not attempt resize with the old target.
Thanks a lot for this fix. Included in my patch set now.
>
> - 0019: Included the earlier mprotect diagnostic fix from
> v20260908-0011. The messages distinguish failure to make the active
> range writable from failure to protect the reserved tail.
>
I think we will come to these fine tuning changes as we are nearer to
the finalization patch. Is this something causing inconvenience during
testing or review?
> Not done:
>
> - The variable-naming changes raised earlier are still pending.
These too should probably wait till we have a wider agreement on the
high level design.
>
> - Direct stale-hint coverage for 0017. The code check is included, but
> there is no test that keeps a buffer-number hint across a shrink
> which removes that buffer slot. Such a test should check that the
> old hint is rejected and the caller finds the page by normal lookup.
> The rollback tests do not exercise this case.
>
I couldn't easily find the discussion that leads to this. Can you
please elaborate?
> - SHOW shared_buffers still includes the pending target, so its output
> cannot be passed to pg_size_bytes(). Should SHOW return a parseable
> size and leave the pending target to pg_get_buffer_resize_status()?
>
pg_size_bytes() would fail only if there's "pending" size in the
output, otherwise it should succeed. But if a size change is pending
then SHOWing just one of either the current NBuffers or the pending
size is going to be misleading. SHOW is not supposed to provide the
value of the GUC, not the size of the buffer pool. Maybe we will just
report the size of buffer pool or just the pending size ultimately,
but I would wait for a wider opinion before actually making that
change.
> - A failed grow may leave some pages allocated in the unused range.
> MADV_POPULATE_WRITE can fail after populating part of the range, and
> making it inaccessible again does not free those pages. I have not
> tested this failure case or added cleanup for it.
Is that possible? The madvise() documentation does not mention that possibility.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-09-28 09:24 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-29 06:51 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-09-28 09:24 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Thu, Sep 18, 2025 at 10:25:29AM +0530, Ashutosh Bapat wrote:
> Given these things, I think we should set up the buffer lookup table
> to hold maximum entries required to expand the buffer pool to its
> maximum, right at the beginning.
Thanks for investigating. I think another option would be to rebuild the
buffer lookup table (create a new table based on the new size and copy
the data over from the original one) as part of the resize procedure,
alongsize with buffers eviction and initialization. From what I recall
the size of buffer lookup table is about two orders of magnitude lower
than shared buffers, so the overhead should not be that large even for
significant amount of buffers.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-28 09:24 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-09-29 06:51 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-01 09:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-09-29 06:51 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Sun, Sep 28, 2025 at 2:54 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Thu, Sep 18, 2025 at 10:25:29AM +0530, Ashutosh Bapat wrote:
> > Given these things, I think we should set up the buffer lookup table
> > to hold maximum entries required to expand the buffer pool to its
> > maximum, right at the beginning.
>
> Thanks for investigating. I think another option would be to rebuild the
> buffer lookup table (create a new table based on the new size and copy
> the data over from the original one) as part of the resize procedure,
> alongsize with buffers eviction and initialization. From what I recall
> the size of buffer lookup table is about two orders of magnitude lower
> than shared buffers, so the overhead should not be that large even for
> significant amount of buffers.
The proposal will work but will require significant work:
1. The pointer to the shared buffer lookup table will change. The
change needs to be absorbed by all the processes at the same time; we
can not have few processes accessing old lookup table and few
processes new one. That has potential to make many processes wait for
a very long time. That can be fixed by accessing a new pointer when
the next buffer lookup access happens by modifying BufTable*
functions. But that means an extra condition checks and some extra
code in those hot paths. Not sure whether that's acceptable.
2. The memory consumed by the old buffer lookup table will need to be
"freed" to the OS. The only way to do so is by having a new memory
segment (which can be unmapped) or unmapping portions of segment
dedicated to the buffer lookup table. That's some more synchronization
and additional wait times for backends.
3. When the new shared buffer lookup table will be built, processes
may be able to access it in shared mode but they may not be able to
make changes to it (or else we need to make corresponding changes to
new table as well). That means more restrictions on the running
backends.
I am not saying that we can not implement your idea, but maybe we
could do that incrementally after basic resizing is in place.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-28 09:24 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-29 06:51 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-10-01 09:10 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-10-01 10:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-10-01 09:10 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Mon, Sep 29, 2025 at 12:21:08PM +0530, Ashutosh Bapat wrote:
> On Sun, Sep 28, 2025 at 2:54 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> >
> > > On Thu, Sep 18, 2025 at 10:25:29AM +0530, Ashutosh Bapat wrote:
> > > Given these things, I think we should set up the buffer lookup table
> > > to hold maximum entries required to expand the buffer pool to its
> > > maximum, right at the beginning.
> >
> > Thanks for investigating. I think another option would be to rebuild the
> > buffer lookup table (create a new table based on the new size and copy
> > the data over from the original one) as part of the resize procedure,
> > alongsize with buffers eviction and initialization. From what I recall
> > the size of buffer lookup table is about two orders of magnitude lower
> > than shared buffers, so the overhead should not be that large even for
> > significant amount of buffers.
>
> The proposal will work but will require significant work:
>
> 1. The pointer to the shared buffer lookup table will change.
Which pointers you mean? AFAICT no operation on the buffer lookup table
returns a pointer (they work with buffer id or a hash) and keys are
compared by value as well.
> we can not have few processes accessing old lookup table and few
> processes new one. That has potential to make many processes wait for
> a very long time.
As I've mentioned above, size of the buffer lookup table is few
magnitudes lower than shared buffers, so I doubt about "a very long
time". But it can be measured.
> 2. The memory consumed by the old buffer lookup table will need to be
> "freed" to the OS. The only way to do so is by having a new memory
> segment
Shared buffer lookup table already lives in it's own segment as
implemented in the current patch, so I don't see any problem here.
I see you folks are inclined to keep some small segments static and
allocate maximum allowed memory for it. It's an option, at the end of
the day we need to experiment and measure both approaches.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-28 09:24 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-29 06:51 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-01 09:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-10-01 10:20 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-01 10:42 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-10-01 10:20 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
On Wed, Oct 1, 2025 at 2:40 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
>
> > On Mon, Sep 29, 2025 at 12:21:08PM +0530, Ashutosh Bapat wrote:
> > On Sun, Sep 28, 2025 at 2:54 PM Dmitry Dolgov <9erthalion6@gmail.com> wrote:
> > >
> > > > On Thu, Sep 18, 2025 at 10:25:29AM +0530, Ashutosh Bapat wrote:
> > > > Given these things, I think we should set up the buffer lookup table
> > > > to hold maximum entries required to expand the buffer pool to its
> > > > maximum, right at the beginning.
> > >
> > > Thanks for investigating. I think another option would be to rebuild the
> > > buffer lookup table (create a new table based on the new size and copy
> > > the data over from the original one) as part of the resize procedure,
> > > alongsize with buffers eviction and initialization. From what I recall
> > > the size of buffer lookup table is about two orders of magnitude lower
> > > than shared buffers, so the overhead should not be that large even for
> > > significant amount of buffers.
> >
> > The proposal will work but will require significant work:
> >
> > 1. The pointer to the shared buffer lookup table will change.
>
> Which pointers you mean? AFAICT no operation on the buffer lookup table
> returns a pointer (they work with buffer id or a hash) and keys are
> compared by value as well.
The buffer lookup table itself.
/* Pass location of hashtable header to hash_create */
infoP->hctl = (HASHHDR *) location;
>
> > we can not have few processes accessing old lookup table and few
> > processes new one. That has potential to make many processes wait for
> > a very long time.
>
> As I've mentioned above, size of the buffer lookup table is few
> magnitudes lower than shared buffers, so I doubt about "a very long
> time". But it can be measured.
>
> > 2. The memory consumed by the old buffer lookup table will need to be
> > "freed" to the OS. The only way to do so is by having a new memory
> > segment
>
> Shared buffer lookup table already lives in it's own segment as
> implemented in the current patch, so I don't see any problem here.
The table is not a single chunk of memory. It's a few chunks spread
across the shared memory segment. Freeing a lookup table is like
freeing those chunks. We have ways to free tail parts of shared memory
segments, but not chunks in-between.
--
Best Wishes,
Ashutosh Bapat
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:05 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Re: Changing shared_buffers without restart Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-28 09:24 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-29 06:51 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-01 09:10 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-10-01 10:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-10-01 10:42 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-10-01 10:42 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Wed, Oct 01, 2025 at 03:50:17PM +0530, Ashutosh Bapat wrote:
> The buffer lookup table itself.
> /* Pass location of hashtable header to hash_create */
> infoP->hctl = (HASHHDR *) location;
How does this affect any users of the lookup table, if they do not even
get to see those?
> > Shared buffer lookup table already lives in it's own segment as
> > implemented in the current patch, so I don't see any problem here.
>
> The table is not a single chunk of memory. It's a few chunks spread
> across the shared memory segment. Freeing a lookup table is like
> freeing those chunks. We have ways to free tail parts of shared memory
> segments, but not chunks in-between.
Right, and the idea was to rebuild it completely to fit into the new
size, not just chunk-by-chunk.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-04-17 23:05 ` Ni Ku <jakkuniku@gmail.com>
2025-04-21 09:33 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
1 sibling, 1 reply; 167+ messages in thread
From: Ni Ku @ 2025-04-17 23:05 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
Hi Ashutosh / Dmitry,
Thanks for the information and discussions, it's been very helpful.
I also have a related question about how ftruncate() is used in the patch.
In my testing I also see that when using ftruncate to shrink a shared
segment, the memory is freed immediately after the call, even if other
processes still have that memory mapped, and they will hit SIGBUS if they
try to access that memory again as the manpage says.
So am I correct to think that, to support the bufferpool shrinking case, it
would not be safe to call ftruncate in AnonymousShmemResize as-is, since at
that point other processes may still be using pages that belong to the
truncated memory?
It appears that for shrinking we should only call ftruncate when we're sure
no process will access those pages again (eg, all processes have handled
the resize interrupt signal barrier). I suppose this can be done by the
resize coordinator after synchronizing with all the other processes.
But in that case it seems we cannot use the postmaster as the coordinator
then? b/c I see some code comments saying the postmaster does not have
waiting infrastructure... (maybe even if the postmaster has waiting infra
we don't want to use it anyway since it can be blocked for a long time and
won't be able to serve other requests).
Regards,
Jack Ng
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 23:05 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
@ 2025-04-21 09:33 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-06 04:23 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-21 09:33 UTC (permalink / raw)
To: Ni Ku <jakkuniku@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>
> On Thu, Apr 17, 2025 at 07:05:36PM GMT, Ni Ku wrote:
> I also have a related question about how ftruncate() is used in the patch.
> In my testing I also see that when using ftruncate to shrink a shared
> segment, the memory is freed immediately after the call, even if other
> processes still have that memory mapped, and they will hit SIGBUS if they
> try to access that memory again as the manpage says.
>
> So am I correct to think that, to support the bufferpool shrinking case, it
> would not be safe to call ftruncate in AnonymousShmemResize as-is, since at
> that point other processes may still be using pages that belong to the
> truncated memory?
> It appears that for shrinking we should only call ftruncate when we're sure
> no process will access those pages again (eg, all processes have handled
> the resize interrupt signal barrier). I suppose this can be done by the
> resize coordinator after synchronizing with all the other processes.
> But in that case it seems we cannot use the postmaster as the coordinator
> then? b/c I see some code comments saying the postmaster does not have
> waiting infrastructure... (maybe even if the postmaster has waiting infra
> we don't want to use it anyway since it can be blocked for a long time and
> won't be able to serve other requests).
There is already a coordination infrastructure, implemented in the patch
0006, which will take care of this and prevent access to the shared
memory until everything is resized.
^ permalink raw reply [nested|flat] 167+ messages in thread
* RE: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 23:05 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-04-21 09:33 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-05-06 04:23 ` Jack Ng <Jack.Ng@huawei.com>
2025-05-06 08:05 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Jack Ng @ 2025-05-06 04:23 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers; Robert Haas <robertmhaas@gmail.com>; Ni Ku <jakkuniku@gmail.com>
Thanks Dmitry. Right, the coordination mechanism in v4-0006 works as expected in various tests (sorry, I misunderstood some details initially).
I also want to report a couple of minor issues found during testing (which you may be aware of already):
1. For memory segments other the first one ('main'), the start address passed to mmap may not be aligned to 4KB or huge page size (since reserved_offset may not be aligned) and cause mmap to fail.
2. Since the ratio for main/desc/iocv/checkpt/strategy in SHMEM_RESIZE_RATIO are relatively small, I think we need to guard against the case where 'max_available_memory' is too small for the required sizes of these segments (from CalculateShmemSize).
Like when max_available_memory=default and shared_numbers=128kB, 'main' still needs ~109MB, but since only 10% of max_available_memory is reserved for it (~102MB) and start address of the next segment is calculated based on reserved_offset, this would cause the mappings to overlap and memory problems later (I hit this after fixing 1.)
I suppose we can change the minimum value of max_available_memory to be large enough, and may also adjust the ratios in SHMEM_RESIZE_RATIO to ensure the reserved space of those segments are sufficient.
Regards,
Jack Ng
-----Original Message-----
From: Dmitry Dolgov <9erthalion6@gmail.com>
Sent: Monday, April 21, 2025 5:33 AM
To: Ni Ku <jakkuniku@gmail.com>
Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers@postgresql.org; Robert Haas <robertmhaas@gmail.com>
Subject: Re: Changing shared_buffers without restart
> On Thu, Apr 17, 2025 at 07:05:36PM GMT, Ni Ku wrote:
> I also have a related question about how ftruncate() is used in the patch.
> In my testing I also see that when using ftruncate to shrink a shared
> segment, the memory is freed immediately after the call, even if other
> processes still have that memory mapped, and they will hit SIGBUS if
> they try to access that memory again as the manpage says.
>
> So am I correct to think that, to support the bufferpool shrinking
> case, it would not be safe to call ftruncate in AnonymousShmemResize
> as-is, since at that point other processes may still be using pages
> that belong to the truncated memory?
> It appears that for shrinking we should only call ftruncate when we're
> sure no process will access those pages again (eg, all processes have
> handled the resize interrupt signal barrier). I suppose this can be
> done by the resize coordinator after synchronizing with all the other processes.
> But in that case it seems we cannot use the postmaster as the
> coordinator then? b/c I see some code comments saying the postmaster
> does not have waiting infrastructure... (maybe even if the postmaster
> has waiting infra we don't want to use it anyway since it can be
> blocked for a long time and won't be able to serve other requests).
There is already a coordination infrastructure, implemented in the patch 0006, which will take care of this and prevent access to the shared memory until everything is resized.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 23:05 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-04-21 09:33 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-06 04:23 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
@ 2025-05-06 08:05 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-07 05:34 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-05-06 08:05 UTC (permalink / raw)
To: Jack Ng <Jack.Ng@huawei.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers; Robert Haas <robertmhaas@gmail.com>; Ni Ku <jakkuniku@gmail.com>
> On Tue, May 06, 2025 at 04:23:07AM GMT, Jack Ng wrote:
> Thanks Dmitry. Right, the coordination mechanism in v4-0006 works as expected in various tests (sorry, I misunderstood some details initially).
Great, thanks for checking.
> I also want to report a couple of minor issues found during testing (which you may be aware of already):
>
> 1. For memory segments other the first one ('main'), the start address passed to mmap may not be aligned to 4KB or huge page size (since reserved_offset may not be aligned) and cause mmap to fail.
>
> 2. Since the ratio for main/desc/iocv/checkpt/strategy in SHMEM_RESIZE_RATIO are relatively small, I think we need to guard against the case where 'max_available_memory' is too small for the required sizes of these segments (from CalculateShmemSize).
> Like when max_available_memory=default and shared_numbers=128kB, 'main' still needs ~109MB, but since only 10% of max_available_memory is reserved for it (~102MB) and start address of the next segment is calculated based on reserved_offset, this would cause the mappings to overlap and memory problems later (I hit this after fixing 1.)
> I suppose we can change the minimum value of max_available_memory to be large enough, and may also adjust the ratios in SHMEM_RESIZE_RATIO to ensure the reserved space of those segments are sufficient.
Yeah, good points. I've introduced max_available_memory expecting some
heated discussions about it, and thus didn't put lots of efforts into
covering all the possible scenarios. But now I'm reworking it along the
lines suggested by Thomas, and will address those as well. Thanks!
^ permalink raw reply [nested|flat] 167+ messages in thread
* RE: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 23:05 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-04-21 09:33 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-06 04:23 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
2025-05-06 08:05 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-05-07 05:34 ` Jack Ng <Jack.Ng@huawei.com>
2025-05-09 14:43 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Jack Ng @ 2025-05-07 05:34 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; pgsql-hackers; Robert Haas <robertmhaas@gmail.com>; Ni Ku <jakkuniku@gmail.com>
> all the possible scenarios. But now I'm reworking it along the lines suggested
> by Thomas, and will address those as well. Thanks!
Thanks for the info, Dmitry.
Just want to confirm my understanding of Thomas' suggestion and your discussions... I think the simpler and more portable solution goes something like the following?
* For each BP resource segment (main, desc, buffers, etc):
1. create an anonymous file as backing
2. mmap a large reserved shared memory area with PROTO_READ/WRITE + MAP_NORESERVE using the anon fd
3. use ftruncate to back the in-use region (and maybe posix_fallocate too to avoid SIGBUS on alloc failure during first-touch), but no need to create a memory mapping for it
4. also no need to create a separate mapping for the reserved region (already covered by the mapping created in 2.)
|-- Memory mapping (MAP_NORESERVE) for BUFFER --|
|-- In-use region --|----- Reserved region -----|
* During resize, simply calculate the new size and call ftruncate on each segment to adjust memory accordingly, no need to mmap/munmap or modify any memory mapping.
I tried this approach with a test program (with huge pages), and both expand and shrink seem to work as expected --for shrink, the memory is freed right after the resize ftruncate.
Regards,
Jack Ng
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 23:05 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-04-21 09:33 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-06 04:23 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
2025-05-06 08:05 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-07 05:34 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
@ 2025-05-09 14:43 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-05-13 05:03 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
0 siblings, 1 reply; 167+ messages in thread
From: Ashutosh Bapat @ 2025-05-09 14:43 UTC (permalink / raw)
To: Jack Ng <Jack.Ng@huawei.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers; Robert Haas <robertmhaas@gmail.com>; Ni Ku <jakkuniku@gmail.com>
On Wed, May 7, 2025 at 11:04 AM Jack Ng <Jack.Ng@huawei.com> wrote:
> > all the possible scenarios. But now I'm reworking it along the lines
> suggested
> > by Thomas, and will address those as well. Thanks!
>
> Thanks for the info, Dmitry.
> Just want to confirm my understanding of Thomas' suggestion and your
> discussions... I think the simpler and more portable solution goes
> something like the following?
>
> * For each BP resource segment (main, desc, buffers, etc):
> 1. create an anonymous file as backing
> 2. mmap a large reserved shared memory area with PROTO_READ/WRITE +
> MAP_NORESERVE using the anon fd
> 3. use ftruncate to back the in-use region (and maybe posix_fallocate
> too to avoid SIGBUS on alloc failure during first-touch), but no need to
> create a memory mapping for it
> 4. also no need to create a separate mapping for the reserved region
> (already covered by the mapping created in 2.)
>
> |-- Memory mapping (MAP_NORESERVE) for BUFFER --|
> |-- In-use region --|----- Reserved region -----|
>
> * During resize, simply calculate the new size and call ftruncate on each
> segment to adjust memory accordingly, no need to mmap/munmap or modify any
> memory mapping.
>
>
That's same as my understanding.
> I tried this approach with a test program (with huge pages), and both
> expand and shrink seem to work as expected --for shrink, the memory is
> freed right after the resize ftruncate.
>
> I thought I had shared a test program upthread, but I don't find it now.
Attached here. Can you please share your test program?
There are concerns around portability of this approach, though.
--
Best Wishes,
Ashutosh Bapat
Attachments:
[text/x-csrc] mfdtruncate.c (1.9K, ../../CAExHW5sNCQdZsUH8qPZ5c1qJHWzhP5K3LJkC+MOE+jX6vTKMyA@mail.gmail.com/3-mfdtruncate.c)
download | inline:
#define _GNU_SOURCE 1 /* See feature_test_macros(7) */
#include <errno.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mman.h>
#include <unistd.h>
#include <stdbool.h>
#define MBSIZE(x) ((x) * 1024 * 1024)
int
main(int argc, char **argv)
{
int flags = MAP_NORESERVE;
void *memaddr;
pid_t pid = getpid();
char *localmem;
int fd = memfd_create("mmap_fd_exp", 0);
size_t size = MBSIZE(300);
localmem = malloc(size);
memset(localmem, 1, size);
if (fd < 0)
{
printf("memfd_create failed with errno %d\n", errno);
exit(__LINE__);
}
if (ftruncate(fd, size) < 0)
{
printf("ftruncate failed with errno %d on fd = %d and size = %ld\n", errno, fd, size);
exit(__LINE__);
}
memaddr = mmap(NULL, size, PROT_WRITE | PROT_READ, MAP_SHARED | MAP_NORESERVE, fd, 0);
if (memaddr == MAP_FAILED)
{
printf("mmap failed with error %m\n");
exit(__LINE__);
}
memset(memaddr, 1, size);
if (memcmp(memaddr, localmem, size) != 0)
{
printf("mmap memory and local memory are not equal upto size = %ld\n", size);
exit(__LINE__);
}
/* causes a segmentation fault: memset(memaddr, 0, maxsize + 1); */
size = MBSIZE(100);
if (ftruncate(fd, size) < 0)
{
printf("ftruncate failed with errno %d on fd = %d and size = %ld\n", errno, fd, size);
exit(__LINE__);
}
memset(memaddr, 1, size);
if (memcmp(memaddr, localmem, size) != 0)
{
printf("mmap memory and local memory are not equal upto size = %ld\n", size);
exit(__LINE__);
}
/* causes a segmentation fault: memset(memaddr, 1, MBSIZE(200)); */
size = MBSIZE(200);
if (ftruncate(fd, size) < 0)
{
printf("ftruncate failed with errno %d on fd = %d and size = %ld\n", errno, fd, size);
exit(__LINE__);
}
memset(memaddr, 1, size);
if (memcmp(memaddr, localmem, size) != 0)
{
printf("mmap memory and local memory are not equal upto size = %ld\n", size);
exit(__LINE__);
}
}
^ permalink raw reply [nested|flat] 167+ messages in thread
* RE: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 12:01 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 06:20 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 09:52 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 23:05 ` Re: Changing shared_buffers without restart Ni Ku <jakkuniku@gmail.com>
2025-04-21 09:33 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-06 04:23 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
2025-05-06 08:05 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-07 05:34 ` RE: Changing shared_buffers without restart Jack Ng <Jack.Ng@huawei.com>
2025-05-09 14:43 ` Re: Changing shared_buffers without restart Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
@ 2025-05-13 05:03 ` Jack Ng <Jack.Ng@huawei.com>
0 siblings, 0 replies; 167+ messages in thread
From: Jack Ng @ 2025-05-13 05:03 UTC (permalink / raw)
To: Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers; Robert Haas <robertmhaas@gmail.com>; Ni Ku <jakkuniku@gmail.com>
Hi Ashutosh,
> > * During resize, simply calculate the new size and call ftruncate on each
> > segment to adjust memory accordingly, no need to mmap/munmap or modify any
> > memory mapping.
> >
> >
> That's same as my understanding.
Great, thanks for confirming!
> I thought I had shared a test program upthread, but I don't find it now. Attached here. Can you please share your test program?
Sure, mine is attached here (it’s based on another test program you shared before :-)
Regards,
Jack Ng
#define _GNU_SOURCE
#include <fcntl.h>
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <sys/mman.h>
#include <linux/memfd.h>
#include <string.h>
#include <errno.h>
#define HUGE_PG_SZ ((size_t)(2048*1024))
unsigned long long MAX_SZ = (100 * ((size_t)(1024 * 1024) * (size_t)(1024 * 1024))); // 100 TB
int main(int argc, char *argv[]) {
int fd_1 = -1;
int fd_2 = -1;
size_t old_size = atoi(argv[1]) * HUGE_PG_SZ;
size_t new_size = atoi(argv[2]) * HUGE_PG_SZ;
// Probe for a valid address that can grow up to MAX_SZ
void *probe = mmap(NULL, MAX_SZ,
PROT_READ | PROT_WRITE,
MAP_ANONYMOUS | MAP_NORESERVE | MAP_SHARED | MAP_HUGETLB, -1, 0);
if (probe == MAP_FAILED) {
perror("mmap");
exit(1);
}
if (munmap(probe, MAX_SZ) == -1) {
perror("munmap");
exit(1);
}
/*
|----- Memory mapping (MAP_NORESERVE) #1 -----|----- Memory mapping (MAP_NORESERVE) #2 -----|
|-- In-use region 1 --|-- Reserved region 1 --|-- In-use region 2 --|-- Reserved region 2 --|
*/
// Create anon mapping #1 that's backed by an anon file fd_1
fd_1 = memfd_create("test_hugepage_map1", MFD_HUGETLB | MFD_HUGE_2MB | MFD_CLOEXEC);
if (fd_1 == -1) {
perror("memfd_create");
exit(1);
}
void *map1_addr = mmap(probe, MAX_SZ/2,
PROT_READ | PROT_WRITE,
MAP_NORESERVE | MAP_SHARED | MAP_HUGETLB, fd_1, 0);
if (map1_addr == MAP_FAILED) {
perror("mmap");
exit(1);
}
// Create in-use region 1
if (ftruncate(fd_1, old_size) == -1) {
perror("ftruncate");
exit(1);
}
// The difference between ftruncate and fallocate is that ftruncate does not actually
// allocate any memory yet (rather it's allocated on "first touch") but fallocate
// does and also initializes the memory with zeros.
/*if (fallocate(fd_1, 0, 0, old_size) == -1) {
perror("fallocate");
exit(1);
}*/
// Optionally, we can protect the reserved region just to be safe
if (mprotect(map1_addr + old_size, MAX_SZ/2 - old_size, PROT_NONE) == -1) {
perror("mprotect");
exit(1);
}
// Create anon mapping #2 that's backed by an anon file fd_2
fd_2 = memfd_create("test_hugepage_map2", MFD_HUGETLB | MFD_HUGE_2MB | MFD_CLOEXEC);
if (fd_2 == -1) {
perror("memfd_create");
exit(1);
}
void *map2_addr = mmap(probe + MAX_SZ/2, MAX_SZ/2,
PROT_READ | PROT_WRITE,
MAP_NORESERVE | MAP_SHARED | MAP_HUGETLB, fd_2, 0);
if (map2_addr == MAP_FAILED) {
perror("mmap");
exit(1);
}
// Create in-use region 2
if (ftruncate(fd_2, old_size) == -1) {
perror("ftruncate");
exit(1);
}
/*if (fallocate(fd_2, 0, 0, old_size) == -1) {
perror("fallocate");
exit(1);
}*/
if (mprotect(map2_addr + old_size, MAX_SZ/2 - old_size, PROT_NONE) == -1) {
perror("mprotect");
exit(1);
}
unsigned char *in_use1 = (unsigned char*)map1_addr;
unsigned char *in_use2 = (unsigned char*)map2_addr;
printf("Press Enter to fill in-use segment 1 memory...\n");
getchar();
memset(map1_addr, 0xee, old_size);
printf("Press Enter to fill in-use segment 2 memory...\n");
getchar();
memset(map2_addr, 0xdd, old_size);
for (int i=0; i < old_size; i++) {
if (in_use1[i] != 0xee) {
printf("seg1 content changed!\n");
exit(1);
}
}
for (int i=0; i < old_size; i++) {
if (in_use2[i] != 0xdd) {
printf("seg2 content changed!\n");
exit(1);
}
}
printf("Press Enter to resize in-use segment 1 memory...\n");
getchar();
// "Unprotect" the old reserved region first
if (mprotect(map1_addr + old_size, MAX_SZ/2 - old_size, PROT_READ | PROT_WRITE) == -1) {
perror("mprotect");
exit(1);
}
if (ftruncate(fd_1, new_size) == -1) {
perror("fd_1 ftruncate");
exit(1);
}
// Protect the new reserved region
if (mprotect(map1_addr + new_size, MAX_SZ/2 - new_size, PROT_NONE) == -1) {
perror("mprotect");
exit(1);
}
printf("Press Enter to resize in-use segment 2 memory...\n");
getchar();
if (mprotect(map2_addr + old_size, MAX_SZ/2 - old_size, PROT_READ | PROT_WRITE) == -1) {
perror("mprotect");
exit(1);
}
if (ftruncate(fd_2, new_size) == -1) {
perror("fd_2 ftruncate");
exit(1);
}
if (mprotect(map2_addr + new_size, MAX_SZ/2 - new_size, PROT_NONE) == -1) {
perror("mprotect");
exit(1);
}
size_t check_size = (old_size <= new_size ? old_size : new_size);
for (int i=0; i < check_size; i++) {
if (in_use1[i] != 0xee) {
printf("new seg1 content changed!\n");
exit(1);
}
}
for (int i=0; i < check_size; i++) {
if (in_use2[i] != 0xdd) {
printf("new seg2 content changed!\n");
exit(1);
}
}
memset(map1_addr, 0xaa, new_size);
for (int i=0; i < new_size; i++) {
if (in_use1[i] != 0xaa) {
printf("new seg1 corrupted!\n");
exit(1);
}
}
memset(map2_addr, 0xbb, new_size);
for (int i=0; i < new_size; i++) {
if (in_use2[i] != 0xbb) {
printf("new seg2 corrupted!\n");
exit(1);
}
}
printf("Press Enter to shrink seg 1 to 0...\n");
getchar();
if (ftruncate(fd_1, 0) == -1) {
perror("ftruncate");
exit(1);
}
close(fd_1);
printf("Press Enter to shrink seg 2 to 0...\n");
getchar();
if (ftruncate(fd_2, 0) == -1) {
perror("ftruncate");
exit(1);
}
close(fd_2);
printf("Press Enter to unmap first mapping...\n");
getchar();
if (munmap(map1_addr, MAX_SZ/2) == -1) {
perror("munmap");
exit(1);
}
printf("Press Enter to unmap first mapping...\n");
getchar();
if (munmap(map2_addr, MAX_SZ/2) == -1) {
perror("munmap");
exit(1);
}
printf("Press Enter to termiante program...\n");
getchar();
return 0;
}
Attachments:
[text/plain] test_hugepage_mappings.c (5.7K, ../../73f05b290b584e0581c4a96a33daa77b@huawei.com/3-test_hugepage_mappings.c)
download | inline:
#define _GNU_SOURCE
#include <fcntl.h>
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <sys/mman.h>
#include <linux/memfd.h>
#include <string.h>
#include <errno.h>
#define HUGE_PG_SZ ((size_t)(2048*1024))
unsigned long long MAX_SZ = (100 * ((size_t)(1024 * 1024) * (size_t)(1024 * 1024))); // 100 TB
int main(int argc, char *argv[]) {
int fd_1 = -1;
int fd_2 = -1;
size_t old_size = atoi(argv[1]) * HUGE_PG_SZ;
size_t new_size = atoi(argv[2]) * HUGE_PG_SZ;
// Probe for a valid address that can grow up to MAX_SZ
void *probe = mmap(NULL, MAX_SZ,
PROT_READ | PROT_WRITE,
MAP_ANONYMOUS | MAP_NORESERVE | MAP_SHARED | MAP_HUGETLB, -1, 0);
if (probe == MAP_FAILED) {
perror("mmap");
exit(1);
}
if (munmap(probe, MAX_SZ) == -1) {
perror("munmap");
exit(1);
}
/*
|----- Memory mapping (MAP_NORESERVE) #1 -----|----- Memory mapping (MAP_NORESERVE) #2 -----|
|-- In-use region 1 --|-- Reserved region 1 --|-- In-use region 2 --|-- Reserved region 2 --|
*/
// Create anon mapping #1 that's backed by an anon file fd_1
fd_1 = memfd_create("test_hugepage_map1", MFD_HUGETLB | MFD_HUGE_2MB | MFD_CLOEXEC);
if (fd_1 == -1) {
perror("memfd_create");
exit(1);
}
void *map1_addr = mmap(probe, MAX_SZ/2,
PROT_READ | PROT_WRITE,
MAP_NORESERVE | MAP_SHARED | MAP_HUGETLB, fd_1, 0);
if (map1_addr == MAP_FAILED) {
perror("mmap");
exit(1);
}
// Create in-use region 1
if (ftruncate(fd_1, old_size) == -1) {
perror("ftruncate");
exit(1);
}
// The difference between ftruncate and fallocate is that ftruncate does not actually
// allocate any memory yet (rather it's allocated on "first touch") but fallocate
// does and also initializes the memory with zeros.
/*if (fallocate(fd_1, 0, 0, old_size) == -1) {
perror("fallocate");
exit(1);
}*/
// Optionally, we can protect the reserved region just to be safe
if (mprotect(map1_addr + old_size, MAX_SZ/2 - old_size, PROT_NONE) == -1) {
perror("mprotect");
exit(1);
}
// Create anon mapping #2 that's backed by an anon file fd_2
fd_2 = memfd_create("test_hugepage_map2", MFD_HUGETLB | MFD_HUGE_2MB | MFD_CLOEXEC);
if (fd_2 == -1) {
perror("memfd_create");
exit(1);
}
void *map2_addr = mmap(probe + MAX_SZ/2, MAX_SZ/2,
PROT_READ | PROT_WRITE,
MAP_NORESERVE | MAP_SHARED | MAP_HUGETLB, fd_2, 0);
if (map2_addr == MAP_FAILED) {
perror("mmap");
exit(1);
}
// Create in-use region 2
if (ftruncate(fd_2, old_size) == -1) {
perror("ftruncate");
exit(1);
}
/*if (fallocate(fd_2, 0, 0, old_size) == -1) {
perror("fallocate");
exit(1);
}*/
if (mprotect(map2_addr + old_size, MAX_SZ/2 - old_size, PROT_NONE) == -1) {
perror("mprotect");
exit(1);
}
unsigned char *in_use1 = (unsigned char*)map1_addr;
unsigned char *in_use2 = (unsigned char*)map2_addr;
printf("Press Enter to fill in-use segment 1 memory...\n");
getchar();
memset(map1_addr, 0xee, old_size);
printf("Press Enter to fill in-use segment 2 memory...\n");
getchar();
memset(map2_addr, 0xdd, old_size);
for (int i=0; i < old_size; i++) {
if (in_use1[i] != 0xee) {
printf("seg1 content changed!\n");
exit(1);
}
}
for (int i=0; i < old_size; i++) {
if (in_use2[i] != 0xdd) {
printf("seg2 content changed!\n");
exit(1);
}
}
printf("Press Enter to resize in-use segment 1 memory...\n");
getchar();
// "Unprotect" the old reserved region first
if (mprotect(map1_addr + old_size, MAX_SZ/2 - old_size, PROT_READ | PROT_WRITE) == -1) {
perror("mprotect");
exit(1);
}
if (ftruncate(fd_1, new_size) == -1) {
perror("fd_1 ftruncate");
exit(1);
}
// Protect the new reserved region
if (mprotect(map1_addr + new_size, MAX_SZ/2 - new_size, PROT_NONE) == -1) {
perror("mprotect");
exit(1);
}
printf("Press Enter to resize in-use segment 2 memory...\n");
getchar();
if (mprotect(map2_addr + old_size, MAX_SZ/2 - old_size, PROT_READ | PROT_WRITE) == -1) {
perror("mprotect");
exit(1);
}
if (ftruncate(fd_2, new_size) == -1) {
perror("fd_2 ftruncate");
exit(1);
}
if (mprotect(map2_addr + new_size, MAX_SZ/2 - new_size, PROT_NONE) == -1) {
perror("mprotect");
exit(1);
}
size_t check_size = (old_size <= new_size ? old_size : new_size);
for (int i=0; i < check_size; i++) {
if (in_use1[i] != 0xee) {
printf("new seg1 content changed!\n");
exit(1);
}
}
for (int i=0; i < check_size; i++) {
if (in_use2[i] != 0xdd) {
printf("new seg2 content changed!\n");
exit(1);
}
}
memset(map1_addr, 0xaa, new_size);
for (int i=0; i < new_size; i++) {
if (in_use1[i] != 0xaa) {
printf("new seg1 corrupted!\n");
exit(1);
}
}
memset(map2_addr, 0xbb, new_size);
for (int i=0; i < new_size; i++) {
if (in_use2[i] != 0xbb) {
printf("new seg2 corrupted!\n");
exit(1);
}
}
printf("Press Enter to shrink seg 1 to 0...\n");
getchar();
if (ftruncate(fd_1, 0) == -1) {
perror("ftruncate");
exit(1);
}
close(fd_1);
printf("Press Enter to shrink seg 2 to 0...\n");
getchar();
if (ftruncate(fd_2, 0) == -1) {
perror("ftruncate");
exit(1);
}
close(fd_2);
printf("Press Enter to unmap first mapping...\n");
getchar();
if (munmap(map1_addr, MAX_SZ/2) == -1) {
perror("munmap");
exit(1);
}
printf("Press Enter to unmap first mapping...\n");
getchar();
if (munmap(map2_addr, MAX_SZ/2) == -1) {
perror("munmap");
exit(1);
}
printf("Press Enter to termiante program...\n");
getchar();
return 0;
}
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-17 11:21 ` Konstantin Knizhnik <knizhnik@garret.ru>
2025-04-17 21:26 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2 siblings, 1 reply; 167+ messages in thread
From: Konstantin Knizhnik @ 2025-04-17 11:21 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; pgsql-hackers; +Cc: Robert Haas <robertmhaas@gmail.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
On 25/02/2025 11:52 am, Dmitry Dolgov wrote:
>> On Fri, Oct 18, 2024 at 09:21:19PM GMT, Dmitry Dolgov wrote:
>> TL;DR A PoC for changing shared_buffers without PostgreSQL restart, via
>> changing shared memory mapping layout. Any feedback is appreciated.
Hi Dmitry,
I am sorry that I have not participated in the discussion in this thread
from the very beginning, although I am also very interested in dynamic
shared buffer resizing and evn proposed my own implementation of it:
https://github.com/knizhnik/postgres/pull/2 based on memory ballooning
and using `madvise`. And it really works (returns unused memory to the
system).
This PoC allows me to understand the main drawbacks of this approach:
1. Performance of Postgres CLOCK page eviction algorithm depends on
number of shared buffers. My first native attempt just to mark unused
buffers as invalid cause significant degrade of performance
pgbench -c 32 -j 4 -T 100 -P1 -M prepared -S
(here shared_buffers - is maximal shared buffers size and
`available_buffers` - is used part:
| shared_buffers | available_buffers | TPS | | ------------------|
---------------------------- | ---- | | 128MB | -1 | 280k | | 1GB | -1 |
324k | | 2GB | -1 | 358k | | 32GB | -1 | 350k | | 2GB | 128Mb | 130k | |
2GB | 1Gb | 311k | | 32GB | 128Mb | 13k | | 32GB | 1Gb | 140k | | 32GB |
2Gb | 348k |
My first thought is to replace clock with LRU based in double-linked
list. As far as there is no lockless double-list implementation,
it need some global lock. This lock can become bottleneck. The standard
solution is partitioning: use N LRU lists instead of 1.
Just as partitioned has table used by buffer manager to lockup buffers.
Actually we can use the same partitions locks to protect LRU list.
But it not clear what to do with ring buffers (strategies).So I decided
not to perform such revolution in bufmgr, but optimize clock to more
efficiently split reserved buffers.
Just add|skip_count|field to buffer descriptor. And it helps! Now the
worst case shared_buffer/available_buffers = 32Gb/128Mb
shows the same performance 280k as shared_buffers=128Mb without ballooning.
2. There are several data structures i Postgres which size depends on
number of buffers.
In my patch I used in some cases dynamic shared buffer size, but if this
structure has to be allocated in shared memory then still maximal size
has to be used. We have the buffers themselves (8 kB per buffer), then
the main BufferDescriptors array (64 B), the BufferIOCVArray (16 B),
checkpoint's CkptBufferIds (20 B), and the hashmap on the buffer cache
(24B+8B/entry).
128 bytes per 8kb bytes seems to large overhead (~1%) but but it may be
quote noticeable with size differences larger than 2 orders of magnitude:
E.g. to support scaling to from 0.5Gb to 128GB , with 128 bytes/buffer
we'd have ~2GiB of static overhead on only 0.5GiB of actual buffers.
3. `madvise` is not portable.
Certainly you have moved much further in your proposal comparing with my
PoC (including huge pages support).
But it is still not quite clear to me how you are going to solve the
problems with large memory overhead in case of ~100x times variation of
shared buffers size.
I
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 11:21 ` Re: Changing shared_buffers without restart Konstantin Knizhnik <knizhnik@garret.ru>
@ 2025-04-17 21:26 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 07:06 ` Re: Changing shared_buffers without restart Konstantin Knizhnik <knizhnik@garret.ru>
0 siblings, 1 reply; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-17 21:26 UTC (permalink / raw)
To: Konstantin Knizhnik <knizhnik@garret.ru>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
> On Thu, Apr 17, 2025 at 02:21:07PM GMT, Konstantin Knizhnik wrote:
>
> 1. Performance of Postgres CLOCK page eviction algorithm depends on number
> of shared buffers. My first native attempt just to mark unused buffers as
> invalid cause significant degrade of performance
Thanks for sharing!
Right, but it concerns the case when the number of shared buffers is
high, independently from whether it was changed online or with a
restart, correct? In that case it's out of scope for this patch.
> 2. There are several data structures i Postgres which size depends on number
> of buffers.
> In my patch I used in some cases dynamic shared buffer size, but if this
> structure has to be allocated in shared memory then still maximal size has
> to be used. We have the buffers themselves (8 kB per buffer), then the main
> BufferDescriptors array (64 B), the BufferIOCVArray (16 B), checkpoint's
> CkptBufferIds (20 B), and the hashmap on the buffer cache (24B+8B/entry).
> 128 bytes per 8kb bytes seems to large overhead (~1%) but but it may be
> quote noticeable with size differences larger than 2 orders of magnitude:
> E.g. to support scaling to from 0.5Gb to 128GB , with 128 bytes/buffer we'd
> have ~2GiB of static overhead on only 0.5GiB of actual buffers.
Not sure what do you mean by using a maximal size, can you elaborate.
In the current patch those structures are allocated as before, except
each goes into a separate segment -- without any extra memory overhead
as far as I see.
> 3. `madvise` is not portable.
The current implementation doesn't rely on madvise so far (it might for
shared memory shrinking), but yeah there are plenty of other not very
portable things (MAP_FIXED, memfd_create). All of that is mentioned in
the corresponding patches as a limitation.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 11:21 ` Re: Changing shared_buffers without restart Konstantin Knizhnik <knizhnik@garret.ru>
2025-04-17 21:26 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
@ 2025-04-18 07:06 ` Konstantin Knizhnik <knizhnik@garret.ru>
2025-04-21 09:38 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 167+ messages in thread
From: Konstantin Knizhnik @ 2025-04-18 07:06 UTC (permalink / raw)
To: Dmitry Dolgov <9erthalion6@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
On 18/04/2025 12:26 am, Dmitry Dolgov wrote:
>> On Thu, Apr 17, 2025 at 02:21:07PM GMT, Konstantin Knizhnik wrote:
>>
>> 1. Performance of Postgres CLOCK page eviction algorithm depends on number
>> of shared buffers. My first native attempt just to mark unused buffers as
>> invalid cause significant degrade of performance
> Thanks for sharing!
>
> Right, but it concerns the case when the number of shared buffers is
> high, independently from whether it was changed online or with a
> restart, correct? In that case it's out of scope for this patch.
>
>> 2. There are several data structures i Postgres which size depends on number
>> of buffers.
>> In my patch I used in some cases dynamic shared buffer size, but if this
>> structure has to be allocated in shared memory then still maximal size has
>> to be used. We have the buffers themselves (8 kB per buffer), then the main
>> BufferDescriptors array (64 B), the BufferIOCVArray (16 B), checkpoint's
>> CkptBufferIds (20 B), and the hashmap on the buffer cache (24B+8B/entry).
>> 128 bytes per 8kb bytes seems to large overhead (~1%) but but it may be
>> quote noticeable with size differences larger than 2 orders of magnitude:
>> E.g. to support scaling to from 0.5Gb to 128GB , with 128 bytes/buffer we'd
>> have ~2GiB of static overhead on only 0.5GiB of actual buffers.
> Not sure what do you mean by using a maximal size, can you elaborate.
>
> In the current patch those structures are allocated as before, except
> each goes into a separate segment -- without any extra memory overhead
> as far as I see.
Thank you for explanation. I am sorry that I have not precisely
investigated your patch before writing: it seems to be that you are are
placing in separate segment only content of shared buffers.
Now I see that I was wrong and it is actually the main difference with
memory ballooning approach I have used. As far as you are are allocating
buffers descriptors and hash table in the same segment,
there is no extra memory overhead.
The only drawback is that we are loosing content of shared buffers in
case of resize. It may be sadly, but not looks like there is no better
alternative.
But there are still some dependencies on shared buffers size which are
not addressed in this PR.
I am not sure how critical they are and is it possible to do something
here, but at least I want to enumerate them:
1. Checkpointer: maximal number of checkpointer requests depends on
NBuffers. So if we start with small shared buffers and then upscale, it
may cause the too frequent checkpoints:
Size
CheckpointerShmemSize(void)
...
size = add_size(size, mul_size(NBuffers,
sizeof(CheckpointerRequest)));
CheckpointerShmemInit(void)
CheckpointerShmem->max_requests = NBuffers;
2. XLOG: number of xlog buffers is calculated depending on number of
shared buffers:
XLOGChooseNumBuffers(void)
{
...
xbuffers = NBuffers / 32;
Should not cause some errors, but may be not so efficient if once again
we start we tiny shared buffers.
3. AIO: AIO max concurrency is also calculated based on number of shared
buffers:
AioChooseMaxConcurrency(void)
{
...
max_proportional_pins = NBuffers / max_backends;
For small shared buffers (i.e. 1Mb, there will be no concurrency at all).
So none of this issues can cause some error, just some inefficient behavior.
But if we want to start with very small shared buffers and then increase
them on demand,
then it can be a problem.
In all this three cases NBuffers is used not just to calculate some
threshold value, but also determine size of the structure in shared memory.
The straightforward solution is to place them in the same segment as
shared buffers. But I am not sure how difficult it will be to implement.
^ permalink raw reply [nested|flat] 167+ messages in thread
* Re: Changing shared_buffers without restart
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-25 09:52 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 11:21 ` Re: Changing shared_buffers without restart Konstantin Knizhnik <knizhnik@garret.ru>
2025-04-17 21:26 ` Re: Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 07:06 ` Re: Changing shared_buffers without restart Konstantin Knizhnik <knizhnik@garret.ru>
@ 2025-04-21 09:38 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 167+ messages in thread
From: Dmitry Dolgov @ 2025-04-21 09:38 UTC (permalink / raw)
To: Konstantin Knizhnik <knizhnik@garret.ru>; +Cc: pgsql-hackers@postgresql.org, Robert Haas <robertmhaas@gmail.com>; Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
> On Fri, Apr 18, 2025 at 10:06:23AM GMT, Konstantin Knizhnik wrote:
> The only drawback is that we are loosing content of shared buffers in case
> of resize. It may be sadly, but not looks like there is no better
> alternative.
No, why would we loose the content? If we do mremap, it will leave the
content as it is. If we do munmap/mmap with an anonymous backing file,
it will also keep the content in memory. The same with another proposal
about using ftruncate/fallocate only, both will leave the content
untouch unless told to do otherwise.
> But there are still some dependencies on shared buffers size which are not
> addressed in this PR.
> I am not sure how critical they are and is it possible to do something here,
> but at least I want to enumerate them:
Righ, I'm aware about those (except the AIO one, which was added after
the first version of the patch), and didn't address them yet due to the
same reason you've mentioned -- they're not hard errors, rather
inefficiencies. But thanks for the reminder, I keep those in the back of
my mind, and when the rest of the design will be settled down, I'll try
to address them as well.
^ permalink raw reply [nested|flat] 167+ messages in thread
end of thread, other threads:[~2026-09-22 07:02 UTC | newest]
Thread overview: 167+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2024-10-18 19:21 Changing shared_buffers without restart Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-01 15:27 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-06 19:10 ` Vladlen Popolitov <v.popolitov@postgrespro.ru>
2024-11-08 16:43 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-07 01:05 ` Thomas Munro <thomas.munro@gmail.com>
2024-11-08 16:40 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-19 12:57 ` Peter Eisentraut <peter@eisentraut.org>
2024-11-19 13:29 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-21 07:55 ` Peter Eisentraut <peter@eisentraut.org>
2025-04-17 15:54 ` Thomas Munro <thomas.munro@gmail.com>
2025-04-18 01:27 ` Thomas Munro <thomas.munro@gmail.com>
2024-11-26 07:53 ` Peter Eisentraut <peter@eisentraut.org>
2024-11-25 19:33 ` Robert Haas <robertmhaas@gmail.com>
2024-11-26 19:17 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 15:20 ` Robert Haas <robertmhaas@gmail.com>
2024-11-27 20:48 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-27 21:05 ` Robert Haas <robertmhaas@gmail.com>
2024-11-27 21:28 ` Jelte Fennema-Nio <postgres@jeltef.nl>
2024-11-28 01:26 ` Robert Haas <robertmhaas@gmail.com>
2024-11-27 21:41 ` Andres Freund <andres@anarazel.de>
2024-11-28 01:28 ` Robert Haas <robertmhaas@gmail.com>
2024-11-28 16:30 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-28 17:18 ` Robert Haas <robertmhaas@gmail.com>
2024-11-28 18:13 ` Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-28 18:57 ` Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 00:56 ` Matthias van de Meent <boekewurm+postgres@gmail.com>
2024-11-29 01:42 ` Tom Lane <tgl@sss.pgh.pa.us>
2024-11-29 16:47 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-02 19:17 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-12-03 14:31 ` Robert Haas <robertmhaas@gmail.com>
2024-12-17 14:10 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-01-13 08:11 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2024-11-28 18:45 ` Dmitry Dolgov <9erthalion6@gmail.com>
2024-11-29 18:17 ` Andres Freund <andres@anarazel.de>
2025-02-25 09:52 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-27 08:28 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-02-28 11:52 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-02-28 12:01 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-03-20 08:55 ` Ni Ku <jakkuniku@gmail.com>
2025-03-20 10:21 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-03-21 08:48 ` Ni Ku <jakkuniku@gmail.com>
2025-03-21 09:31 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-03-21 10:30 ` Ni Ku <jakkuniku@gmail.com>
2025-04-07 06:20 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-07 08:43 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-09 05:42 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-09 07:45 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-09 07:50 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-09 08:19 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-11 14:34 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-11 15:01 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 05:10 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-14 07:20 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-14 08:58 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 09:52 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-04-17 21:16 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 09:17 ` Thomas Munro <thomas.munro@gmail.com>
2025-04-18 11:02 ` Andres Freund <andres@anarazel.de>
2025-04-18 11:05 ` Thomas Munro <thomas.munro@gmail.com>
2025-04-21 09:29 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-21 14:16 ` Thomas Munro <thomas.munro@gmail.com>
2025-06-10 11:09 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-16 12:39 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-06-20 10:19 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-06-20 10:22 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-02 12:35 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-04 00:06 ` Tomas Vondra <tomas@vondra.me>
2025-07-04 14:41 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-04 15:23 ` Tomas Vondra <tomas@vondra.me>
2025-07-05 10:35 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-07 11:57 ` Tomas Vondra <tomas@vondra.me>
2025-07-07 13:06 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-07 13:42 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-07 13:58 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:01 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-06 13:21 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-13 18:37 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 04:55 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:10 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 08:25 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-07-14 08:54 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 09:24 ` Thom Brown <thom@linux.com>
2025-07-14 09:32 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 12:56 ` Andres Freund <andres@anarazel.de>
2025-07-14 13:08 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:14 ` Andres Freund <andres@anarazel.de>
2025-07-14 13:20 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 13:42 ` Andres Freund <andres@anarazel.de>
2025-07-14 14:01 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:22 ` Burd, Greg <greg@burd.me>
2025-07-14 14:43 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 14:23 ` Andres Freund <andres@anarazel.de>
2025-07-14 14:39 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:11 ` Andres Freund <andres@anarazel.de>
2025-07-14 15:18 ` Jack Ng <Jack.Ng@huawei.com>
2025-07-14 16:32 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-15 22:52 ` Jack Ng <Jack.Ng@huawei.com>
2025-07-16 14:48 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:35 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-14 15:10 ` Jack Ng <Jack.Ng@huawei.com>
2025-07-14 22:55 ` Jim Nasby <jnasby@upgrade.com>
2025-07-16 14:52 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-07-16 15:44 ` Andres Freund <andres@anarazel.de>
2025-09-18 04:47 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 04:55 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-18 13:52 ` Andres Freund <andres@anarazel.de>
2025-09-18 14:05 ` Andres Freund <andres@anarazel.de>
2025-09-26 18:04 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-26 18:36 ` Andres Freund <andres@anarazel.de>
2025-09-29 06:57 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-13 15:58 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-14 08:35 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-10-16 16:25 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-17 10:33 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-11-14 11:53 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-01-28 13:19 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-06 09:25 ` Bowen Shi <zxwsbg12138@gmail.com>
2026-02-06 10:00 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-07 23:44 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 15:15 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 22:14 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-10 15:23 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 15:39 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-12 20:42 ` Andres Freund <andres@anarazel.de>
2026-02-09 22:37 ` Andres Freund <andres@anarazel.de>
2026-02-10 05:32 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:43 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-12 20:12 ` Andres Freund <andres@anarazel.de>
2026-02-13 11:49 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-09 13:41 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-09 14:29 ` Andres Freund <andres@anarazel.de>
2026-02-10 12:50 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 06:17 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-10 14:37 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-10 15:21 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-02-12 14:12 ` Jakub Wartak <jakub.wartak@enterprisedb.com>
2026-02-13 11:52 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-07-24 12:56 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 11:56 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2026-08-17 14:26 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-20 09:11 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-20 11:13 ` Yuhang Qiu <iamqyh@gmail.com>
2026-08-26 16:36 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-08-27 09:01 ` Yuhang Qiu <iamqyh@gmail.com>
2026-09-07 13:00 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-08 03:09 ` Yuhang Qiu <iamqyh@gmail.com>
2026-09-08 11:24 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-11 10:56 ` Yuhang Qiu <iamqyh@gmail.com>
2026-09-17 16:02 ` Palak Chaturvedi <chaturvedipalak1911@gmail.com>
2026-09-22 03:45 ` Yuhang Qiu <iamqyh@gmail.com>
2026-09-22 07:02 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-09-28 09:24 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-09-29 06:51 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-01 09:10 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-10-01 10:20 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-10-01 10:42 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-17 23:05 ` Ni Ku <jakkuniku@gmail.com>
2025-04-21 09:33 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-06 04:23 ` Jack Ng <Jack.Ng@huawei.com>
2025-05-06 08:05 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-05-07 05:34 ` Jack Ng <Jack.Ng@huawei.com>
2025-05-09 14:43 ` Ashutosh Bapat <ashutosh.bapat.oss@gmail.com>
2025-05-13 05:03 ` Jack Ng <Jack.Ng@huawei.com>
2025-04-17 11:21 ` Konstantin Knizhnik <knizhnik@garret.ru>
2025-04-17 21:26 ` Dmitry Dolgov <9erthalion6@gmail.com>
2025-04-18 07:06 ` Konstantin Knizhnik <knizhnik@garret.ru>
2025-04-21 09:38 ` Dmitry Dolgov <9erthalion6@gmail.com>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox