Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from ) id 1u1i4y-006o17-UA for pgsql-hackers@arkaria.postgresql.org; Mon, 07 Apr 2025 08:43:41 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.94.2) (envelope-from ) id 1u1i4w-00DzMD-NI for pgsql-hackers@arkaria.postgresql.org; Mon, 07 Apr 2025 08:43:39 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from <9erthalion6@gmail.com>) id 1u1i4v-00DzLU-RQ for pgsql-hackers@lists.postgresql.org; Mon, 07 Apr 2025 08:43:38 +0000 Received: from mail-ed1-x533.google.com ([2a00:1450:4864:20::533]) by makus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.96) (envelope-from <9erthalion6@gmail.com>) id 1u1i4n-003R8h-1U for pgsql-hackers@postgresql.org; Mon, 07 Apr 2025 08:43:37 +0000 Received: by mail-ed1-x533.google.com with SMTP id 4fb4d7f45d1cf-5e6194e9d2cso7798982a12.2 for ; Mon, 07 Apr 2025 01:43:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1744015408; x=1744620208; darn=postgresql.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=lTadLQ2ngHHwwLGPJ7oKESv46Wn3+ar6YtsDcfhgNJs=; b=E4BxMXqz9EBPCC/gPnW3Rq3yGgwm1sr4loCcXLXpaM6pJv+ER33Em46q0zIcr99j1j 0cfeJ0+8qgGQuh/LPMn9bcmiRN/xyIIFNXYm4S9itXlQ0x7nsvkRTiUuhK2AYq1x1ldA et+RiWMvYlCQF7UQVnqvZ1hGqxA8O+5AcsZAb1McuCSor3aEvFXFBcYxZJZPvJFomIue ijuj+o3kcGGBmiJSy6kEyFNOBzJoMbtx5EJLeU23KSDZWR7OpHs1Ul6Yqjhz+LfF0/dm FFf7kcLLolFLysGbZXGwkt5YmMlXlfIll8p5p6GQRsQu29vVfDAs2nsg73GBZJc0Vgo9 Km9A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1744015408; x=1744620208; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=lTadLQ2ngHHwwLGPJ7oKESv46Wn3+ar6YtsDcfhgNJs=; b=cvZKjfweTVXFvtZ7UOOsVSpeoMgJLTt88u46gfi2rqogYwxsg4bbSO8fdekPW74jy9 sC+wDJT1RZf1A1mvs/kpvQWPwIeUtz3UYBAPoAn8Em+8P2aMogOhvz3DQPZBtGFoaLCM uR9LZiFZ+p3hXUmXQQbYBNdhLc76eMjk0dyULU12X6mFBwDG/jhm7TkZiFJMXwn52OzG HmOpjSgueqmMhM49JIax5HH+YdPKxJlKW5mDMyXsphrdrCQk/v9Q4Cvbyu8XqzeE0uVl 4xb0VB6Hr44vkeTsQg5Jm2CisDQyJzs5AH3kOlZLULUrSKJw/tljAtg5+BfeQ9k2BVem R2Vg== X-Gm-Message-State: AOJu0Yw9lG4CMzKFSuP0clNp2A50ZsawuiVEh31Fm+k7uJKHRVmvptRL QQU8DS3YthwQ1sriTUDbB/MrlRBg3FBHezmU00AqvXDIM5ygumzPtl7/rA== X-Gm-Gg: ASbGncvV62PL0AvQ2SmXOKf9jrrJkcQ2D9YJlbTWJ7nj8UwmIs1zYyuix9EcIA+5eRN hF4ifxl0KwNt4/Z4B0J+MTesjqmUbaSHUcG3jvK/yO3SuHCIy7FiP6ug5Mm6oUo0OYXy1cs36rz uYWSGpb59StLX1wqpwxh/rvOBgKzHbiyK4iGTHZNY+A7sSD19bR9AKfDYEyuDoOu62/YHfqVhmH 0/pQbgTtOjjBUOAcfvQrIt36T3Qn6ZZLMKd+IKdkLz3BD4SkrkWhxVUfvZRTRkAlBLLl3dEEisE 5dhwI3uqK63qLmZt/n8gkDQAqZb0igj17olz3L+8a+zEcxkhmG5ci+qpMiehiHgWOgU6N8b7PBy ygTnDJW9Ca9Y3AZQWLHM7QX68PItYJv6X6iXT6Qw7cPpNhtvKbPA8nVxrCA== X-Google-Smtp-Source: AGHT+IESFJcq5Vfx2DspRTlLneLs5N2NxSphS1KZ5Gkm38SckRHsW575g5i/z5ZHs4T90ziUuRLF7w== X-Received: by 2002:a05:6402:2b99:b0:5eb:ca9b:523a with SMTP id 4fb4d7f45d1cf-5f0b3bf72c9mr9524851a12.20.1744015407009; Mon, 07 Apr 2025 01:43:27 -0700 (PDT) Received: from ddolgov-thinkpadt14sgen1.rmtde.csb (dslb-178-005-227-122.178.005.pools.vodafone-ip.de. [178.5.227.122]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-5f087714e1csm6188788a12.7.2025.04.07.01.43.23 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 07 Apr 2025 01:43:26 -0700 (PDT) Date: Mon, 7 Apr 2025 10:43:22 +0200 From: Dmitry Dolgov <9erthalion6@gmail.com> To: Ashutosh Bapat Cc: pgsql-hackers@postgresql.org, Robert Haas Subject: Re: Changing shared_buffers without restart Message-ID: References: MIME-Version: 1.0 Content-Type: multipart/mixed; boundary="vninua6xybvzgrci" Content-Disposition: inline In-Reply-To: List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: inline > On Mon, Apr 07, 2025 at 11:50:46AM GMT, Ashutosh Bapat wrote: > This is because the BarrierArriveAndWait() only waits for all the > attached backends. It doesn't wait for backends which are yet to > attach. I think what we want is *all* the backends should execute all > the phases synchronously and wait for others to finish. If we don't do > that, there's a possibility that some of them would see inconsistent > buffer states or even worse may not have necessary memory mapped and > resized - thus causing segfaults. Am I correct? > > I think what needs to be done is that every backend should wait for other > backends to attach themselves to the barrier before moving to the > first phase. One way I can think of is we use two signal barriers - > one to ensure that all the backends have attached themselves and > second for the actual resizing. But then the postmaster needs to wait for > all the processes to process the first signal barrier. A postmaster can > not wait on anything. Maybe there's a way to poll, but I didn't find > it. Does that mean that we have to make some other backend a coordinator? Yes, you're right, plain dynamic Barrier does not ensure all available processes will be synchronized. I was aware about the scenario you describe, it's mentioned in commentaries for the resize function. I was under the impression this should be enough, but after some more thinking I'm not so sure anymore. Let me try to structure it as a list of possible corner cases that we need to worry about: * New backend spawned while we're busy resizing shared memory. Those should wait until the resizing is complete and get the new size as well. * Old backend receives a resize message, but exits before attempting to resize. Those should be excluded from coordination. * A backend is blocked and not responding before or after the ProcSignalBarrier message was sent. I'm thinking about a failure situation, when one rogue backend is doing something without checking for interrupts. We need to wait for those to become responsive, and potentially abort shared memory resize after some timeout. * Backends join the barrier in disjoint groups with some time in between, which is longer than what it takes to resize shared memory. That means that relying only on the shared dynamic barrier is not enough -- it will only synchronize resize procedure withing those groups. Out of those I think the third poses some problems, e.g. if we shrinking the shared memory, but one backend is accessing buffer pool without checking for interrupts. In the v3 implementation this won't be handled correctly, other backends will ignore such rogue process. Independently from that we could reason about the logic much easier if it's guaranteed that all the process to resize shared memory will wait for each other to start simultaneously. Looks like to achieve that we need a slightly different combination of a global Barrier and ProcSignalBarrier mechanism. We can't use ProcSignalBarrier as it is, because processes need to wait for each other, and at the same time finish processing to bump the generation. We also can't use a simple dynamic Barrier due to possibility of disjoint groups of processes. A static Barrier is also not easier, because we would need somehow to know exact number of processes, which might change over time. I think a relatively elegant solution is to extend ProcSignalBarrier mechanism to track not only pss_barrierGeneration, as a sign that everything was processed, but also something like pss_barrierReceivedGeneration, indicating that the message was received everywhere but not processed yet. That would be enough to allow processes to wait until the resize message was received everywhere, then use a global Barrier to wait until all processes are finished. It's somehow similar to your proposal to use two signals, but has less implementation overhead. This would also allow different solutions regarding error handling. E.g. we could do an unbounded waiting for all processes we expect to resize, assuming that the user will be able to intervene and fix an issue if there is any. Or we can do a timed waiting, and abort the resize after some timeout of not all processes are ready yet. In the new v4 version of the patch the first option is implemented. On top of that there are following changes: * Shared memory address space is now reserved for future usage, making shared memory segments clash (e.g. due to memory allocation) impossible. There is a new GUC to control how much space to reserve, which is called max_available_memory -- on the assumption that most of the time it would make sense to set its value to the total amount of memory on the machine. I'm open for suggestions regarding the name. * There is one more patch to address hugepages remap. As mentioned in this thread above, Linux kernel has certain limitations when it comes to mremap for segments allocated with huge pages. To work around it's possible to replace mremap with a sequence of unmap and map again, relying on the anon file behind the segment to keep the memory content. I haven't found any downsides of this approach so far, but it makes the anonymous file patch 0007 mandatory. --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="v4-0001-Allow-to-use-multiple-shared-memory-mappings.patch" From 15b87a1cb89d3f31b656e27d07ead5aa935f1643 Mon Sep 17 00:00:00 2001 From: Dmitrii Dolgov <9erthalion6@gmail.com> Date: Fri, 28 Feb 2025 19:54:47 +0100 Subject: [PATCH v4 1/8] Allow to use multiple shared memory mappings Currently all the work with shared memory is done via a single anonymous memory mapping, which limits ways how the shared memory could be organized. Introduce possibility to allocate multiple shared memory mappings, where a single mapping is associated with a specified shared memory segment. There is only fixed amount of available segments, currently only one main shared memory segment is allocated. A new shared memory API is introduces, extended with a segment as a new parameter. As a path of least resistance, the original API is kept in place, utilizing the main shared memory segment. --- src/backend/port/posix_sema.c | 4 +- src/backend/port/sysv_sema.c | 4 +- src/backend/port/sysv_shmem.c | 138 ++++++++++++++++++++--------- src/backend/port/win32_sema.c | 2 +- src/backend/storage/ipc/ipc.c | 4 +- src/backend/storage/ipc/ipci.c | 63 +++++++------ src/backend/storage/ipc/shmem.c | 141 +++++++++++++++++++++--------- src/backend/storage/lmgr/lwlock.c | 13 ++- src/include/storage/ipc.h | 2 +- src/include/storage/pg_sema.h | 2 +- src/include/storage/pg_shmem.h | 18 ++++ src/include/storage/shmem.h | 12 +++ 12 files changed, 278 insertions(+), 125 deletions(-) diff --git a/src/backend/port/posix_sema.c b/src/backend/port/posix_sema.c index 269c7460817..401e1113fa1 100644 --- a/src/backend/port/posix_sema.c +++ b/src/backend/port/posix_sema.c @@ -193,7 +193,7 @@ PGSemaphoreShmemSize(int maxSemas) * we don't have to expose the counters to other processes.) */ void -PGReserveSemaphores(int maxSemas) +PGReserveSemaphores(int maxSemas, int shmem_segment) { struct stat statbuf; @@ -220,7 +220,7 @@ PGReserveSemaphores(int maxSemas) * ShmemAlloc() won't be ready yet. */ sharedSemas = (PGSemaphore) - ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas)); + ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment); #endif numSems = 0; diff --git a/src/backend/port/sysv_sema.c b/src/backend/port/sysv_sema.c index f7c8638aec5..b6301463ac7 100644 --- a/src/backend/port/sysv_sema.c +++ b/src/backend/port/sysv_sema.c @@ -313,7 +313,7 @@ PGSemaphoreShmemSize(int maxSemas) * have clobbered.) */ void -PGReserveSemaphores(int maxSemas) +PGReserveSemaphores(int maxSemas, int shmem_segment) { struct stat statbuf; @@ -334,7 +334,7 @@ PGReserveSemaphores(int maxSemas) * ShmemAlloc() won't be ready yet. */ sharedSemas = (PGSemaphore) - ShmemAllocUnlocked(PGSemaphoreShmemSize(maxSemas)); + ShmemAllocUnlockedInSegment(PGSemaphoreShmemSize(maxSemas), shmem_segment); numSharedSemas = 0; maxSharedSemas = maxSemas; diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c index 197926d44f6..56af0231d24 100644 --- a/src/backend/port/sysv_shmem.c +++ b/src/backend/port/sysv_shmem.c @@ -94,8 +94,19 @@ typedef enum unsigned long UsedShmemSegID = 0; void *UsedShmemSegAddr = NULL; -static Size AnonymousShmemSize; -static void *AnonymousShmem = NULL; +typedef struct AnonymousMapping +{ + int shmem_segment; + Size shmem_size; /* Size of the mapping */ + Pointer shmem; /* Pointer to the start of the mapped memory */ + Pointer seg_addr; /* SysV shared memory for the header */ + unsigned long seg_id; /* IPC key */ +} AnonymousMapping; + +static AnonymousMapping Mappings[ANON_MAPPINGS]; + +/* Keeps track of used mapping segments */ +static int next_free_segment = 0; static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size); static void IpcMemoryDetach(int status, Datum shmaddr); @@ -104,6 +115,28 @@ static IpcMemoryState PGSharedMemoryAttach(IpcMemoryId shmId, void *attachAt, PGShmemHeader **addr); +static const char* +MappingName(int shmem_segment) +{ + switch (shmem_segment) + { + case MAIN_SHMEM_SEGMENT: + return "main"; + default: + return "unknown"; + } +} + +static void +DebugMappings() +{ + for(int i = 0; i < next_free_segment; i++) + { + AnonymousMapping m = Mappings[i]; + elog(DEBUG1, "Mapping[%s]: addr %p, size %zu", + MappingName(i), m.shmem, m.shmem_size); + } +} /* * InternalIpcMemoryCreate(memKey, size) @@ -591,14 +624,13 @@ check_huge_page_size(int *newval, void **extra, GucSource source) /* * Creates an anonymous mmap()ed shared memory segment. * - * Pass the requested size in *size. This function will modify *size to the - * actual size of the allocation, if it ends up allocating a segment that is - * larger than requested. + * This function will modify mapping size to the actual size of the allocation, + * if it ends up allocating a segment that is larger than requested. */ -static void * -CreateAnonymousSegment(Size *size) +static void +CreateAnonymousSegment(AnonymousMapping *mapping) { - Size allocsize = *size; + Size allocsize = mapping->shmem_size; void *ptr = MAP_FAILED; int mmap_errno = 0; @@ -623,8 +655,11 @@ CreateAnonymousSegment(Size *size) PG_MMAP_FLAGS | mmap_flags, -1, 0); mmap_errno = errno; if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED) - elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m", - allocsize); + { + DebugMappings(); + elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m", + MappingName(mapping->shmem_segment), allocsize); + } } #endif @@ -642,7 +677,7 @@ CreateAnonymousSegment(Size *size) * Use the original size, not the rounded-up value, when falling back * to non-huge pages. */ - allocsize = *size; + allocsize = mapping->shmem_size; ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE, PG_MMAP_FLAGS, -1, 0); mmap_errno = errno; @@ -651,8 +686,10 @@ CreateAnonymousSegment(Size *size) if (ptr == MAP_FAILED) { errno = mmap_errno; + DebugMappings(); ereport(FATAL, - (errmsg("could not map anonymous shared memory: %m"), + (errmsg("segment[%s]: could not map anonymous shared memory: %m", + MappingName(mapping->shmem_segment)), (mmap_errno == ENOMEM) ? errhint("This error usually means that PostgreSQL's request " "for a shared memory segment exceeded available memory, " @@ -663,8 +700,8 @@ CreateAnonymousSegment(Size *size) allocsize) : 0)); } - *size = allocsize; - return ptr; + mapping->shmem = ptr; + mapping->shmem_size = allocsize; } /* @@ -674,13 +711,18 @@ CreateAnonymousSegment(Size *size) static void AnonymousShmemDetach(int status, Datum arg) { - /* Release anonymous shared memory block, if any. */ - if (AnonymousShmem != NULL) + for(int i = 0; i < next_free_segment; i++) { - if (munmap(AnonymousShmem, AnonymousShmemSize) < 0) - elog(LOG, "munmap(%p, %zu) failed: %m", - AnonymousShmem, AnonymousShmemSize); - AnonymousShmem = NULL; + AnonymousMapping m = Mappings[i]; + + /* Release anonymous shared memory block, if any. */ + if (m.shmem != NULL) + { + if (munmap(m.shmem, m.shmem_size) < 0) + elog(LOG, "munmap(%p, %zu) failed: %m", + m.shmem, m.shmem_size); + m.shmem = NULL; + } } } @@ -705,6 +747,7 @@ PGSharedMemoryCreate(Size size, PGShmemHeader *hdr; struct stat statbuf; Size sysvsize; + AnonymousMapping *mapping = &Mappings[next_free_segment]; /* * We use the data directory's ID info (inode and device numbers) to @@ -733,11 +776,15 @@ PGSharedMemoryCreate(Size size, /* Room for a header? */ Assert(size > MAXALIGN(sizeof(PGShmemHeader))); + mapping->shmem_size = size; + mapping->shmem_segment = next_free_segment; if (shared_memory_type == SHMEM_TYPE_MMAP) { - AnonymousShmem = CreateAnonymousSegment(&size); - AnonymousShmemSize = size; + /* On success, mapping data will be modified. */ + CreateAnonymousSegment(mapping); + + next_free_segment++; /* Register on-exit routine to unmap the anonymous segment */ on_shmem_exit(AnonymousShmemDetach, (Datum) 0); @@ -760,7 +807,7 @@ PGSharedMemoryCreate(Size size, * loop simultaneously. (CreateDataDirLockFile() does not entirely ensure * that, but prefer fixing it over coping here.) */ - NextShmemSegID = statbuf.st_ino; + NextShmemSegID = statbuf.st_ino + next_free_segment; for (;;) { @@ -852,13 +899,13 @@ PGSharedMemoryCreate(Size size, /* * Initialize space allocation status for segment. */ - hdr->totalsize = size; + hdr->totalsize = mapping->shmem_size; hdr->freeoffset = MAXALIGN(sizeof(PGShmemHeader)); *shim = hdr; /* Save info for possible future use */ - UsedShmemSegAddr = memAddress; - UsedShmemSegID = (unsigned long) NextShmemSegID; + mapping->seg_addr = memAddress; + mapping->seg_id = (unsigned long) NextShmemSegID; /* * If AnonymousShmem is NULL here, then we're not using anonymous shared @@ -866,10 +913,10 @@ PGSharedMemoryCreate(Size size, * block. Otherwise, the System V shared memory block is only a shim, and * we must return a pointer to the real block. */ - if (AnonymousShmem == NULL) + if (mapping->shmem == NULL) return hdr; - memcpy(AnonymousShmem, hdr, sizeof(PGShmemHeader)); - return (PGShmemHeader *) AnonymousShmem; + memcpy(mapping->shmem, hdr, sizeof(PGShmemHeader)); + return (PGShmemHeader *) mapping->shmem; } #ifdef EXEC_BACKEND @@ -969,23 +1016,28 @@ PGSharedMemoryNoReAttach(void) void PGSharedMemoryDetach(void) { - if (UsedShmemSegAddr != NULL) + for(int i = 0; i < next_free_segment; i++) { - if ((shmdt(UsedShmemSegAddr) < 0) + AnonymousMapping m = Mappings[i]; + + if (m.seg_addr != NULL) + { + if ((shmdt(m.seg_addr) < 0) #if defined(EXEC_BACKEND) && defined(__CYGWIN__) - /* Work-around for cygipc exec bug */ - && shmdt(NULL) < 0 + /* Work-around for cygipc exec bug */ + && shmdt(NULL) < 0 #endif - ) - elog(LOG, "shmdt(%p) failed: %m", UsedShmemSegAddr); - UsedShmemSegAddr = NULL; - } + ) + elog(LOG, "shmdt(%p) failed: %m", m.seg_addr); + m.seg_addr = NULL; + } - if (AnonymousShmem != NULL) - { - if (munmap(AnonymousShmem, AnonymousShmemSize) < 0) - elog(LOG, "munmap(%p, %zu) failed: %m", - AnonymousShmem, AnonymousShmemSize); - AnonymousShmem = NULL; + if (m.shmem != NULL) + { + if (munmap(m.shmem, m.shmem_size) < 0) + elog(LOG, "munmap(%p, %zu) failed: %m", + m.shmem, m.shmem_size); + m.shmem = NULL; + } } } diff --git a/src/backend/port/win32_sema.c b/src/backend/port/win32_sema.c index 5854ad1f54d..e7365ff8060 100644 --- a/src/backend/port/win32_sema.c +++ b/src/backend/port/win32_sema.c @@ -44,7 +44,7 @@ PGSemaphoreShmemSize(int maxSemas) * process exits. */ void -PGReserveSemaphores(int maxSemas) +PGReserveSemaphores(int maxSemas, int shmem_segment) { mySemSet = (HANDLE *) malloc(maxSemas * sizeof(HANDLE)); if (mySemSet == NULL) diff --git a/src/backend/storage/ipc/ipc.c b/src/backend/storage/ipc/ipc.c index 567739b5be9..5b55bec8d9d 100644 --- a/src/backend/storage/ipc/ipc.c +++ b/src/backend/storage/ipc/ipc.c @@ -61,6 +61,8 @@ static void proc_exit_prepare(int code); * but provide some additional features we need --- in particular, * we want to register callbacks to invoke when we are disconnecting * from a broken shared-memory context but not exiting the postmaster. + * Maximum number of such exit callbacks depends on the number of shared + * segments. * * Callback functions can take zero, one, or two args: the first passed * arg is the integer exitcode, the second is the Datum supplied when @@ -68,7 +70,7 @@ static void proc_exit_prepare(int code); * ---------------------------------------------------------------- */ -#define MAX_ON_EXITS 20 +#define MAX_ON_EXITS 40 struct ONEXIT { diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c index 2fa045e6b0f..8b38e985327 100644 --- a/src/backend/storage/ipc/ipci.c +++ b/src/backend/storage/ipc/ipci.c @@ -86,7 +86,7 @@ RequestAddinShmemSpace(Size size) * required. */ Size -CalculateShmemSize(int *num_semaphores) +CalculateShmemSize(int *num_semaphores, int shmem_segment) { Size size; int numSemas; @@ -206,33 +206,38 @@ CreateSharedMemoryAndSemaphores(void) Assert(!IsUnderPostmaster); - /* Compute the size of the shared-memory block */ - size = CalculateShmemSize(&numSemas); - elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size); - - /* - * Create the shmem segment - */ - seghdr = PGSharedMemoryCreate(size, &shim); - - /* - * Make sure that huge pages are never reported as "unknown" while the - * server is running. - */ - Assert(strcmp("unknown", - GetConfigOption("huge_pages_status", false, false)) != 0); - - InitShmemAccess(seghdr); - - /* - * Create semaphores - */ - PGReserveSemaphores(numSemas); - - /* - * Set up shared memory allocation mechanism - */ - InitShmemAllocation(); + for(int segment = 0; segment < ANON_MAPPINGS; segment++) + { + /* Compute the size of the shared-memory block */ + size = CalculateShmemSize(&numSemas, segment); + elog(DEBUG3, "invoking IpcMemoryCreate(size=%zu)", size); + + /* + * Create the shmem segment. + * + * XXX: Do multiple shims are needed, one per segment? + */ + seghdr = PGSharedMemoryCreate(size, &shim); + + /* + * Make sure that huge pages are never reported as "unknown" while the + * server is running. + */ + Assert(strcmp("unknown", + GetConfigOption("huge_pages_status", false, false)) != 0); + + InitShmemAccessInSegment(seghdr, segment); + + /* + * Create semaphores + */ + PGReserveSemaphores(numSemas, segment); + + /* + * Set up shared memory allocation mechanism + */ + InitShmemAllocationInSegment(segment); + } /* Initialize subsystems */ CreateOrAttachShmemStructs(); @@ -363,7 +368,7 @@ InitializeShmemGUCs(void) /* * Calculate the shared memory size and round up to the nearest megabyte. */ - size_b = CalculateShmemSize(&num_semas); + size_b = CalculateShmemSize(&num_semas, MAIN_SHMEM_SEGMENT); size_mb = add_size(size_b, (1024 * 1024) - 1) / (1024 * 1024); sprintf(buf, "%zu", size_mb); SetConfigOption("shared_memory_size", buf, diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c index 895a43fb39e..389abc82519 100644 --- a/src/backend/storage/ipc/shmem.c +++ b/src/backend/storage/ipc/shmem.c @@ -75,19 +75,19 @@ #include "utils/builtins.h" static void *ShmemAllocRaw(Size size, Size *allocated_size); +static void *ShmemAllocRawInSegment(Size size, Size *allocated_size, + int shmem_segment); /* shared memory global variables */ -static PGShmemHeader *ShmemSegHdr; /* shared mem segment header */ +ShmemSegment Segments[ANON_MAPPINGS]; -static void *ShmemBase; /* start address of shared memory */ - -static void *ShmemEnd; /* end+1 address of shared memory */ - -slock_t *ShmemLock; /* spinlock for shared memory and LWLock - * allocation */ - -static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */ +/* + * Primary index hashtable for shmem, for simplicity we use a single for all + * shared memory segments. There can be performance consequences of that, and + * an alternative option would be to have one index per shared memory segments. + */ +static HTAB *ShmemIndex = NULL; /* @@ -96,9 +96,17 @@ static HTAB *ShmemIndex = NULL; /* primary index hashtable for shmem */ void InitShmemAccess(PGShmemHeader *seghdr) { - ShmemSegHdr = seghdr; - ShmemBase = seghdr; - ShmemEnd = (char *) ShmemBase + seghdr->totalsize; + InitShmemAccessInSegment(seghdr, MAIN_SHMEM_SEGMENT); +} + +void +InitShmemAccessInSegment(PGShmemHeader *seghdr, int shmem_segment) +{ + PGShmemHeader *shmhdr = (PGShmemHeader *) seghdr; + ShmemSegment *seg = &Segments[shmem_segment]; + seg->ShmemSegHdr = shmhdr; + seg->ShmemBase = (void *) shmhdr; + seg->ShmemEnd = (char *) seg->ShmemBase + shmhdr->totalsize; } /* @@ -109,7 +117,13 @@ InitShmemAccess(PGShmemHeader *seghdr) void InitShmemAllocation(void) { - PGShmemHeader *shmhdr = ShmemSegHdr; + InitShmemAllocationInSegment(MAIN_SHMEM_SEGMENT); +} + +void +InitShmemAllocationInSegment(int shmem_segment) +{ + PGShmemHeader *shmhdr = Segments[shmem_segment].ShmemSegHdr; char *aligned; Assert(shmhdr != NULL); @@ -118,9 +132,9 @@ InitShmemAllocation(void) * Initialize the spinlock used by ShmemAlloc. We must use * ShmemAllocUnlocked, since obviously ShmemAlloc can't be called yet. */ - ShmemLock = (slock_t *) ShmemAllocUnlocked(sizeof(slock_t)); + Segments[shmem_segment].ShmemLock = (slock_t *) ShmemAllocUnlockedInSegment(sizeof(slock_t), shmem_segment); - SpinLockInit(ShmemLock); + SpinLockInit(Segments[shmem_segment].ShmemLock); /* * Allocations after this point should go through ShmemAlloc, which @@ -145,11 +159,17 @@ InitShmemAllocation(void) */ void * ShmemAlloc(Size size) +{ + return ShmemAllocInSegment(size, MAIN_SHMEM_SEGMENT); +} + +void * +ShmemAllocInSegment(Size size, int shmem_segment) { void *newSpace; Size allocated_size; - newSpace = ShmemAllocRaw(size, &allocated_size); + newSpace = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment); if (!newSpace) ereport(ERROR, (errcode(ERRCODE_OUT_OF_MEMORY), @@ -179,6 +199,12 @@ ShmemAllocNoError(Size size) */ static void * ShmemAllocRaw(Size size, Size *allocated_size) +{ + return ShmemAllocRawInSegment(size, allocated_size, MAIN_SHMEM_SEGMENT); +} + +static void * +ShmemAllocRawInSegment(Size size, Size *allocated_size, int shmem_segment) { Size newStart; Size newFree; @@ -198,22 +224,22 @@ ShmemAllocRaw(Size size, Size *allocated_size) size = CACHELINEALIGN(size); *allocated_size = size; - Assert(ShmemSegHdr != NULL); + Assert(Segments[shmem_segment].ShmemSegHdr != NULL); - SpinLockAcquire(ShmemLock); + SpinLockAcquire(Segments[shmem_segment].ShmemLock); - newStart = ShmemSegHdr->freeoffset; + newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset; newFree = newStart + size; - if (newFree <= ShmemSegHdr->totalsize) + if (newFree <= Segments[shmem_segment].ShmemSegHdr->totalsize) { - newSpace = (char *) ShmemBase + newStart; - ShmemSegHdr->freeoffset = newFree; + newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart; + Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree; } else newSpace = NULL; - SpinLockRelease(ShmemLock); + SpinLockRelease(Segments[shmem_segment].ShmemLock); /* note this assert is okay with newSpace == NULL */ Assert(newSpace == (void *) CACHELINEALIGN(newSpace)); @@ -231,6 +257,12 @@ ShmemAllocRaw(Size size, Size *allocated_size) */ void * ShmemAllocUnlocked(Size size) +{ + return ShmemAllocUnlockedInSegment(size, MAIN_SHMEM_SEGMENT); +} + +void * +ShmemAllocUnlockedInSegment(Size size, int shmem_segment) { Size newStart; Size newFree; @@ -241,19 +273,19 @@ ShmemAllocUnlocked(Size size) */ size = MAXALIGN(size); - Assert(ShmemSegHdr != NULL); + Assert(Segments[shmem_segment].ShmemSegHdr != NULL); - newStart = ShmemSegHdr->freeoffset; + newStart = Segments[shmem_segment].ShmemSegHdr->freeoffset; newFree = newStart + size; - if (newFree > ShmemSegHdr->totalsize) + if (newFree > Segments[shmem_segment].ShmemSegHdr->totalsize) ereport(ERROR, (errcode(ERRCODE_OUT_OF_MEMORY), errmsg("out of shared memory (%zu bytes requested)", size))); - ShmemSegHdr->freeoffset = newFree; + Segments[shmem_segment].ShmemSegHdr->freeoffset = newFree; - newSpace = (char *) ShmemBase + newStart; + newSpace = (char *) Segments[shmem_segment].ShmemBase + newStart; Assert(newSpace == (void *) MAXALIGN(newSpace)); @@ -268,7 +300,13 @@ ShmemAllocUnlocked(Size size) bool ShmemAddrIsValid(const void *addr) { - return (addr >= ShmemBase) && (addr < ShmemEnd); + return ShmemAddrIsValidInSegment(addr, MAIN_SHMEM_SEGMENT); +} + +bool +ShmemAddrIsValidInSegment(const void *addr, int shmem_segment) +{ + return (addr >= Segments[shmem_segment].ShmemBase) && (addr < Segments[shmem_segment].ShmemEnd); } /* @@ -329,6 +367,18 @@ ShmemInitHash(const char *name, /* table string name for shmem index */ long max_size, /* max size of the table */ HASHCTL *infoP, /* info about key and bucket size */ int hash_flags) /* info about infoP */ +{ + return ShmemInitHashInSegment(name, init_size, max_size, infoP, hash_flags, + MAIN_SHMEM_SEGMENT); +} + +HTAB * +ShmemInitHashInSegment(const char *name, /* table string name for shmem index */ + long init_size, /* initial table size */ + long max_size, /* max size of the table */ + HASHCTL *infoP, /* info about key and bucket size */ + int hash_flags, /* info about infoP */ + int shmem_segment) /* in which segment to keep the table */ { bool found; void *location; @@ -345,9 +395,9 @@ ShmemInitHash(const char *name, /* table string name for shmem index */ hash_flags |= HASH_SHARED_MEM | HASH_ALLOC | HASH_DIRSIZE; /* look it up in the shmem index */ - location = ShmemInitStruct(name, + location = ShmemInitStructInSegment(name, hash_get_shared_size(infoP, hash_flags), - &found); + &found, shmem_segment); /* * if it already exists, attach to it rather than allocate and initialize @@ -380,6 +430,13 @@ ShmemInitHash(const char *name, /* table string name for shmem index */ */ void * ShmemInitStruct(const char *name, Size size, bool *foundPtr) +{ + return ShmemInitStructInSegment(name, size, foundPtr, MAIN_SHMEM_SEGMENT); +} + +void * +ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, + int shmem_segment) { ShmemIndexEnt *result; void *structPtr; @@ -388,7 +445,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr) if (!ShmemIndex) { - PGShmemHeader *shmemseghdr = ShmemSegHdr; + PGShmemHeader *shmemseghdr = Segments[shmem_segment].ShmemSegHdr; /* Must be trying to create/attach to ShmemIndex itself */ Assert(strcmp(name, "ShmemIndex") == 0); @@ -411,7 +468,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr) * process can be accessing shared memory yet. */ Assert(shmemseghdr->index == NULL); - structPtr = ShmemAlloc(size); + structPtr = ShmemAllocInSegment(size, shmem_segment); shmemseghdr->index = structPtr; *foundPtr = false; } @@ -428,8 +485,8 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr) LWLockRelease(ShmemIndexLock); ereport(ERROR, (errcode(ERRCODE_OUT_OF_MEMORY), - errmsg("could not create ShmemIndex entry for data structure \"%s\"", - name))); + errmsg("could not create ShmemIndex entry for data structure \"%s\" in segment %d", + name, shmem_segment))); } if (*foundPtr) @@ -454,7 +511,7 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr) Size allocated_size; /* It isn't in the table yet. allocate and initialize it */ - structPtr = ShmemAllocRaw(size, &allocated_size); + structPtr = ShmemAllocRawInSegment(size, &allocated_size, shmem_segment); if (structPtr == NULL) { /* out of memory; remove the failed ShmemIndex entry */ @@ -473,14 +530,13 @@ ShmemInitStruct(const char *name, Size size, bool *foundPtr) LWLockRelease(ShmemIndexLock); - Assert(ShmemAddrIsValid(structPtr)); + Assert(ShmemAddrIsValidInSegment(structPtr, shmem_segment)); Assert(structPtr == (void *) CACHELINEALIGN(structPtr)); return structPtr; } - /* * Add two Size values, checking for overflow */ @@ -537,10 +593,11 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS) /* output all allocated entries */ memset(nulls, 0, sizeof(nulls)); + /* XXX: take all shared memory segments into account. */ while ((ent = (ShmemIndexEnt *) hash_seq_search(&hstat)) != NULL) { values[0] = CStringGetTextDatum(ent->key); - values[1] = Int64GetDatum((char *) ent->location - (char *) ShmemSegHdr); + values[1] = Int64GetDatum((char *) ent->location - (char *) Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr); values[2] = Int64GetDatum(ent->size); values[3] = Int64GetDatum(ent->allocated_size); named_allocated += ent->allocated_size; @@ -552,15 +609,15 @@ pg_get_shmem_allocations(PG_FUNCTION_ARGS) /* output shared memory allocated but not counted via the shmem index */ values[0] = CStringGetTextDatum(""); nulls[1] = true; - values[2] = Int64GetDatum(ShmemSegHdr->freeoffset - named_allocated); + values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset - named_allocated); values[3] = values[2]; tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls); /* output as-of-yet unused shared memory */ nulls[0] = true; - values[1] = Int64GetDatum(ShmemSegHdr->freeoffset); + values[1] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset); nulls[1] = false; - values[2] = Int64GetDatum(ShmemSegHdr->totalsize - ShmemSegHdr->freeoffset); + values[2] = Int64GetDatum(Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->totalsize - Segments[MAIN_SHMEM_SEGMENT].ShmemSegHdr->freeoffset); values[3] = values[2]; tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls); diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c index 3df29658f18..8241c061507 100644 --- a/src/backend/storage/lmgr/lwlock.c +++ b/src/backend/storage/lmgr/lwlock.c @@ -80,6 +80,8 @@ #include "pg_trace.h" #include "pgstat.h" #include "port/pg_bitutils.h" +#include "postmaster/postmaster.h" +#include "storage/pg_shmem.h" #include "storage/proc.h" #include "storage/proclist.h" #include "storage/procnumber.h" @@ -618,10 +620,15 @@ LWLockNewTrancheId(void) int *LWLockCounter; LWLockCounter = (int *) ((char *) MainLWLockArray - sizeof(int)); - /* We use the ShmemLock spinlock to protect LWLockCounter */ - SpinLockAcquire(ShmemLock); + /* + * We use the ShmemLock spinlock to protect LWLockCounter. + * + * XXX: Looks like this is the only use of Segments outside of shmem.c, + * it's maybe worth it to reshape this part to hide Segments structure. + */ + SpinLockAcquire(Segments[MAIN_SHMEM_SEGMENT].ShmemLock); result = (*LWLockCounter)++; - SpinLockRelease(ShmemLock); + SpinLockRelease(Segments[MAIN_SHMEM_SEGMENT].ShmemLock); return result; } diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h index 3baf418b3d1..6ebda479ced 100644 --- a/src/include/storage/ipc.h +++ b/src/include/storage/ipc.h @@ -77,7 +77,7 @@ extern void check_on_shmem_exit_lists_are_empty(void); /* ipci.c */ extern PGDLLIMPORT shmem_startup_hook_type shmem_startup_hook; -extern Size CalculateShmemSize(int *num_semaphores); +extern Size CalculateShmemSize(int *num_semaphores, int shmem_segment); extern void CreateSharedMemoryAndSemaphores(void); #ifdef EXEC_BACKEND extern void AttachSharedMemoryStructs(void); diff --git a/src/include/storage/pg_sema.h b/src/include/storage/pg_sema.h index fa6ca35a51f..8ae9637fcd0 100644 --- a/src/include/storage/pg_sema.h +++ b/src/include/storage/pg_sema.h @@ -41,7 +41,7 @@ typedef HANDLE PGSemaphore; extern Size PGSemaphoreShmemSize(int maxSemas); /* Module initialization (called during postmaster start or shmem reinit) */ -extern void PGReserveSemaphores(int maxSemas); +extern void PGReserveSemaphores(int maxSemas, int shmem_segment); /* Allocate a PGSemaphore structure with initial count 1 */ extern PGSemaphore PGSemaphoreCreate(void); diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h index b99ebc9e86f..138078c29c5 100644 --- a/src/include/storage/pg_shmem.h +++ b/src/include/storage/pg_shmem.h @@ -25,6 +25,7 @@ #define PG_SHMEM_H #include "storage/dsm_impl.h" +#include "storage/spin.h" typedef struct PGShmemHeader /* standard header for all Postgres shmem */ { @@ -41,6 +42,20 @@ typedef struct PGShmemHeader /* standard header for all Postgres shmem */ #endif } PGShmemHeader; +typedef struct ShmemSegment +{ + PGShmemHeader *ShmemSegHdr; /* shared mem segment header */ + void *ShmemBase; /* start address of shared memory */ + void *ShmemEnd; /* end+1 address of shared memory */ + slock_t *ShmemLock; /* spinlock for shared memory and LWLock + * allocation */ +} ShmemSegment; + +/* Number of available segments for anonymous memory mappings */ +#define ANON_MAPPINGS 1 + +extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS]; + /* GUC variables */ extern PGDLLIMPORT int shared_memory_type; extern PGDLLIMPORT int huge_pages; @@ -90,4 +105,7 @@ extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2); extern void PGSharedMemoryDetach(void); extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags); +/* The main segment, contains everything except buffer blocks and related data. */ +#define MAIN_SHMEM_SEGMENT 0 + #endif /* PG_SHMEM_H */ diff --git a/src/include/storage/shmem.h b/src/include/storage/shmem.h index 904a336b851..5929f140236 100644 --- a/src/include/storage/shmem.h +++ b/src/include/storage/shmem.h @@ -29,15 +29,27 @@ extern PGDLLIMPORT slock_t *ShmemLock; struct PGShmemHeader; /* avoid including storage/pg_shmem.h here */ extern void InitShmemAccess(struct PGShmemHeader *seghdr); +extern void InitShmemAccessInSegment(struct PGShmemHeader *seghdr, + int shmem_segment); extern void InitShmemAllocation(void); +extern void InitShmemAllocationInSegment(int shmem_segment); extern void *ShmemAlloc(Size size); +extern void *ShmemAllocInSegment(Size size, int shmem_segment); extern void *ShmemAllocNoError(Size size); extern void *ShmemAllocUnlocked(Size size); +extern void *ShmemAllocUnlockedInSegment(Size size, int shmem_segment); extern bool ShmemAddrIsValid(const void *addr); +extern bool ShmemAddrIsValidInSegment(const void *addr, int shmem_segment); extern void InitShmemIndex(void); +extern void InitVariableShmemIndex(void); extern HTAB *ShmemInitHash(const char *name, long init_size, long max_size, HASHCTL *infoP, int hash_flags); +extern HTAB *ShmemInitHashInSegment(const char *name, long init_size, + long max_size, HASHCTL *infoP, + int hash_flags, int shmem_segment); extern void *ShmemInitStruct(const char *name, Size size, bool *foundPtr); +extern void *ShmemInitStructInSegment(const char *name, Size size, + bool *foundPtr, int shmem_segment); extern Size add_size(Size s1, Size s2); extern Size mul_size(Size s1, Size s2); base-commit: 5e1915439085014140314979c4dd5e23bd677cac -- 2.45.1 --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="v4-0002-Address-space-reservation-for-shared-memory.patch" From eae77d430e6e6cc3ec95b2cf613e4b3ae095e75e Mon Sep 17 00:00:00 2001 From: Dmitrii Dolgov <9erthalion6@gmail.com> Date: Wed, 16 Oct 2024 20:21:33 +0200 Subject: [PATCH v4 2/8] Address space reservation for shared memory Currently the kernel is responsible to chose an address, where to place each shared memory mapping, which is the lowest possible address that do not clash with any other mappings. This is considered to be the most portable approach, but one of the downsides is that there is no place to resize allocated mappings anymore. Here is how it looks like for one mapping in /proc/$PID/maps, /dev/zero represents the anonymous shared memory we talk about: 00400000-00490000 /path/bin/postgres ... 012d9000-0133e000 [heap] 7f443a800000-7f470a800000 /dev/zero (deleted) 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2 ... 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted) By specifying the mapping address directly it's possible to place the mapping in a way that leaves room for resizing. The idea is: * To reserve some address space via mmap'ing a large chunk of memory with PROT_NONE and MAP_NORESERVE. This way we prepare a playground for preparing shared memory layout without risking anything interfering with that. * To slice the reserved space up into sections, one to use for each shared segment. * Allocate shared memory segments out of corresponding slices and leaving unclaimed space in between them. This is implemented via mmap'ing memory at a specified address from the reserved space with MAP_FIXED. The result looks like this: 012d9000-0133e000 [heap] 7f443a800000-7f444196c000 /dev/zero (deleted) 7f444196c000-7f470a800000 # reserved space 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2 Things like address space randomization should not be a problem in this context, since the randomization is applied to the mmap base, which is one per process. This approach also do not impact the actual memory usage as reported by the kernel. Here is the output of /proc/$PID/status for the master version with shared_buffers = 128 MB: // Peak virtual memory size, which is described as total pages // mapped in mm_struct. It corresponds to the mapped reserved space // and is the only number that grows with it. VmPeak: 2043192 kB // Size of memory portions. It contains RssAnon + RssFile + RssShmem VmRSS: 22908 kB // Size of resident anonymous memory RssAnon: 768 kB // Size of resident file mappings RssFile: 10364 kB // Size of resident shmem memory (includes SysV shm, mapping of tmpfs and // shared anonymous mappings) RssShmem: 11776 kB Here is the same for the patch when reserving 20GB of space: VmPeak: 21250648 kB VmRSS: 22948 kB RssAnon: 768 kB RssFile: 10404 kB RssShmem: 11776 kB Cgroup v2 doesn't have any problems with that as well. To verify a new cgroup was created with the memory limit 256 MB, then PostgreSQL was launched withing this cgroup with shared_buffers = 128 MB: $ cd /sys/fs/cgroup $ mkdir postgres $ cd postres $ echo 268435456 > memory.max $ echo $MASTER_PID_SHELL > cgroup.procs # postgres from the master branch has being successfully launched # from that shell $ cat memory.current 17465344 (~16.6 MB) # stop postgres $ echo $PATCH_PID_SHELL > cgroup.procs # postgres from the patch has being successfully launched from that shell $ cat memory.current 17637376 (~16.8 MB) To control the amount of space reserved a new GUC max_available_memory is introduced. Ideally it should be based on the maximum available memory, hense the name. --- src/backend/port/sysv_shmem.c | 284 ++++++++++++++++++++++++---- src/backend/port/win32_shmem.c | 2 +- src/backend/storage/ipc/ipci.c | 5 +- src/backend/utils/init/globals.c | 1 + src/backend/utils/misc/guc_tables.c | 14 ++ src/include/storage/pg_shmem.h | 4 +- 6 files changed, 271 insertions(+), 39 deletions(-) diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c index 56af0231d24..a0f03ff868f 100644 --- a/src/backend/port/sysv_shmem.c +++ b/src/backend/port/sysv_shmem.c @@ -108,6 +108,66 @@ static AnonymousMapping Mappings[ANON_MAPPINGS]; /* Keeps track of used mapping segments */ static int next_free_segment = 0; +/* + * Anonymous mapping placing (/dev/zero (deleted) below) looks like this: + * + * 00400000-00490000 /path/bin/postgres + * ... + * 012d9000-0133e000 [heap] + * 7f443a800000-7f470a800000 /dev/zero (deleted) + * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive + * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2 + * ... + * 7f471aef2000-7f471aef9000 /dev/shm/PostgreSQL.3859891842 + * 7f471aef9000-7f471aefa000 /SYSV007dbf7d (deleted) + * ... + * + * We would like to place multiple mappings in such a way, that there will be + * enough space between them in the address space to be able to resize up to + * certain size, but without counting towards the total memory consumption. + * + * To achieve that we first reserve some shared memory address space by + * mmap'ing a segment of MaxAvailableMemory size with PROT_NONE and + * MAP_NORESERVE (these flags allow to make sure this space will not be used by + * anything else, yet do not count against memory limits). Having the reserved + * space, we allocate out of it actual chunks of shared memory as usual, + * updating a pointer to the current available reserved space for the next + * allocation with the gap between segments in mind. + * + * The result would look like this: + * + * 012d9000-0133e000 [heap] + * 7f4426f54000-7f442e010000 /dev/zero (deleted) + * 7f442e010000-7f443a800000 # reserved empty space + * 7f443a800000-7f444196c000 /dev/zero (deleted) + * 7f444196c000-7f470a800000 # reserved empty space + * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive + * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2 + * [...] + * + * The reserved space pointer is calculated to slice up the total reserved + * space into fixed fractions of address space for each segment, as specified + * in the SHMEM_RESIZE_RATIO array. + */ +static double SHMEM_RESIZE_RATIO[1] = { + 1.0, /* MAIN_SHMEM_SLOT */ +}; + +/* + * Offset from the beginning of the reserved space, which indicates currently + * available range. New shared memory segments have to be allocated at this + * offset related to the reserved space. + */ +static Size reserved_offset = 0; + +/* + * Flag telling that we have decided to use huge pages. + * + * XXX: It's possible to use GetConfigOption("huge_pages_status", false, false) + * instead, but it feels like an overkill. + */ +static bool huge_pages_on = false; + static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size); static void IpcMemoryDetach(int status, Datum shmaddr); static void IpcMemoryDelete(int status, Datum shmId); @@ -626,39 +686,198 @@ check_huge_page_size(int *newval, void **extra, GucSource source) * * This function will modify mapping size to the actual size of the allocation, * if it ends up allocating a segment that is larger than requested. + * + * Note that we do not switch from huge pages to regular pages in this + * function, this decision was already made in ReserveAnonymousMemory and we + * stick to it. */ static void -CreateAnonymousSegment(AnonymousMapping *mapping) +CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base) { Size allocsize = mapping->shmem_size; void *ptr = MAP_FAILED; int mmap_errno = 0; + int mmap_flags = PG_MMAP_FLAGS; #ifndef MAP_HUGETLB - /* PGSharedMemoryCreate should have dealt with this case */ - Assert(huge_pages != HUGE_PAGES_ON); + /* ReserveAnonymousMemory should have dealt with this case */ + Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on); #else - if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY) + if (huge_pages_on) { - /* - * Round up the request size to a suitable large value. - */ Size hugepagesize; - int mmap_flags; + /* Make sure nothing is messed up */ + Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY); + + /* Round up the request size to a suitable large value */ GetHugePageSize(&hugepagesize, &mmap_flags); if (allocsize % hugepagesize != 0) allocsize += hugepagesize - (allocsize % hugepagesize); + mmap_flags = PG_MMAP_FLAGS | mmap_flags; + } +#endif + + elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p", + MappingName(mapping->shmem_segment), allocsize, base + reserved_offset); + + /* + * Try to create mapping at an address out of the reserved range, which + * will allow to extend it later. Use reserved_offset to allocate the + * segment, then update currently available reserved range. + * + * If the last step has failed, fallback to the regular mapping + * creation and signal that shared buffers could not be resized without + * a restart. + */ + ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE, + mmap_flags | MAP_FIXED, -1, 0); + mmap_errno = errno; + + if (ptr == MAP_FAILED) + { + DebugMappings(); + elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p failed: %m, " + "fallback to the non-resizable allocation", + MappingName(mapping->shmem_segment), allocsize, base + reserved_offset); + ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE, - PG_MMAP_FLAGS | mmap_flags, -1, 0); + PG_MMAP_FLAGS, -1, 0); + mmap_errno = errno; + } + else + { + Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ; + + reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment]; + } + + if (ptr == MAP_FAILED) + { + errno = mmap_errno; + DebugMappings(); + ereport(FATAL, + (errmsg("segment[%s]: could not map anonymous shared memory: %m", + MappingName(mapping->shmem_segment)), + (mmap_errno == ENOMEM) ? + errhint("This error usually means that PostgreSQL's request " + "for a shared memory segment exceeded available memory, " + "swap space, or huge pages. To reduce the request size " + "(currently %zu bytes), reduce PostgreSQL's shared " + "memory usage, perhaps by reducing \"shared_buffers\" or " + "\"max_connections\".", + allocsize) : 0)); + } + + mapping->shmem = ptr; + mapping->shmem_size = allocsize; +} + +/* + * ReserveAnonymousMemory + * + * Reserve shared memory address space, from which shared memory segments are + * going to be sliced out. The goal of this exercise is to support segments + * resizing, for which we need a reserved space free of potential clashes with + * other mmap'd areas that are not under our control. Reservation is done via + * mmap, and will not allocate any memory until it will be actually used, and + * MAP_NORESERVE allows to make it not counting againt kernel reservation + * limits (e.g. in cgroups or for huge pages). Do not get confused because of + * MAP_NORESERVE -- we need to reserve some space, but not the actual memory, + * and that is that this flag is about. + * + * Note, that with MAP_NORESERVE a reservation with hugetlb will succeed even + * if there is actually not enough huge pages. Hence this function is + * responsible for deciding whether to use huge pages or not. To achieve that + * we need to probe first and try to allocate needed memory for all segments -- + * if this succeeds, we unmap the probe segment and use hugetlb; if it fails, + * we proceed with the regular memory. + */ +void * +ReserveAnonymousMemory(Size reserve_size) +{ + Size allocsize = reserve_size; + void *ptr = MAP_FAILED; + int mmap_errno = 0; + + /* Complain if hugepages demanded but we can't possibly support them */ +#if !defined(MAP_HUGETLB) + if (huge_pages == HUGE_PAGES_ON) + ereport(ERROR, + (errcode(ERRCODE_FEATURE_NOT_SUPPORTED), + errmsg("huge pages not supported on this platform"))); +#else + if (huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY) + { + Size hugepagesize, total_size = 0; + int mmap_flags; + + GetHugePageSize(&hugepagesize, &mmap_flags); + + /* + * Figure out how much memory is needed for all segments, keeping in + * mind that for every segment this value will be rounding up by the + * huge page size. The resulting value will be used to probe memory and + * decide whether we will allocate huge pages or not. + * + * We could actually have a mix and match of segments with and without + * huge pages. But in that case we need to have multiple reservation + * spaces to use corresponding memory (hugetlb adress space reserved + * for hugetlb segments, regular memory for others), and it doesn't + * seem to worth the complexity for now. + */ + for(int segment = 0; segment < ANON_MAPPINGS; segment++) + { + int numSemas; + Size segment_size = CalculateShmemSize(&numSemas, segment); + + if (segment_size % hugepagesize != 0) + segment_size += hugepagesize - (segment_size % hugepagesize); + + total_size += segment_size; + } + + /* Map total amount of memory to test its availability. */ + elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB", + total_size); + ptr = mmap(NULL, total_size, PROT_NONE, + PG_MMAP_FLAGS | MAP_ANONYMOUS | mmap_flags, -1, 0); mmap_errno = errno; if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED) { - DebugMappings(); - elog(DEBUG1, "segment[%s]: mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m", - MappingName(mapping->shmem_segment), allocsize); + /* No huge pages, we will go with the regular page size */ + elog(DEBUG1, "reserving space: probe mmap(%zu) with MAP_HUGETLB " + "failed, huge pages disabled: %m", total_size); + } + else + { + /* + * All fine, unmap the temporary segment and proceed with reserving + * using huge pages. + */ + if (munmap(ptr, total_size) < 0) + elog(LOG, "reservice space: munmap(%p, %zu) failed: %m", + ptr, total_size); + + /* Round up the requested size to a suitable large value. */ + if (allocsize % hugepagesize != 0) + allocsize += hugepagesize - (allocsize % hugepagesize); + + elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB", + allocsize); + ptr = mmap(NULL, allocsize, PROT_NONE, + PG_MMAP_FLAGS | MAP_ANONYMOUS | MAP_NORESERVE | mmap_flags, + -1, 0); + mmap_errno = errno; + + /* This should not happen, but handle errors anyway */ + if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED) + { + elog(DEBUG1, "reserving space: mmap(%zu) with MAP_HUGETLB " + "failed, huge pages disabled: %m", allocsize); + } } } #endif @@ -666,10 +885,12 @@ CreateAnonymousSegment(AnonymousMapping *mapping) /* * Report whether huge pages are in use. This needs to be tracked before * the second mmap() call if attempting to use huge pages failed - * previously. + * previously. At this point ptr is either pointing to the probe segment, + * if we couldn't mmap it, or the reservation space. */ SetConfigOption("huge_pages_status", (ptr == MAP_FAILED) ? "off" : "on", PGC_INTERNAL, PGC_S_DYNAMIC_DEFAULT); + huge_pages_on = ptr != MAP_FAILED; if (ptr == MAP_FAILED && huge_pages != HUGE_PAGES_ON) { @@ -677,10 +898,11 @@ CreateAnonymousSegment(AnonymousMapping *mapping) * Use the original size, not the rounded-up value, when falling back * to non-huge pages. */ - allocsize = mapping->shmem_size; - ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE, - PG_MMAP_FLAGS, -1, 0); - mmap_errno = errno; + allocsize = reserve_size; + + elog(DEBUG1, "reserving space: mmap(%zu)", allocsize); + ptr = mmap(NULL, allocsize, PROT_NONE, + MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0); } if (ptr == MAP_FAILED) @@ -688,20 +910,18 @@ CreateAnonymousSegment(AnonymousMapping *mapping) errno = mmap_errno; DebugMappings(); ereport(FATAL, - (errmsg("segment[%s]: could not map anonymous shared memory: %m", - MappingName(mapping->shmem_segment)), + (errmsg("reserving space: could not map anonymous shared " + "memory: %m"), (mmap_errno == ENOMEM) ? errhint("This error usually means that PostgreSQL's request " - "for a shared memory segment exceeded available memory, " - "swap space, or huge pages. To reduce the request size " - "(currently %zu bytes), reduce PostgreSQL's shared " - "memory usage, perhaps by reducing \"shared_buffers\" or " - "\"max_connections\".", + "for a reserved shared memory address space exceeded " + "available memory, swap space, or huge pages. To " + "reduce the request reservation size (currently %zu " + "bytes), reduce PostgreSQL's \"maximum_shared_buffers\".", allocsize) : 0)); } - mapping->shmem = ptr; - mapping->shmem_size = allocsize; + return ptr; } /* @@ -740,7 +960,7 @@ AnonymousShmemDetach(int status, Datum arg) */ PGShmemHeader * PGSharedMemoryCreate(Size size, - PGShmemHeader **shim) + PGShmemHeader **shim, Pointer base) { IpcMemoryKey NextShmemSegID; void *memAddress; @@ -760,14 +980,6 @@ PGSharedMemoryCreate(Size size, errmsg("could not stat data directory \"%s\": %m", DataDir))); - /* Complain if hugepages demanded but we can't possibly support them */ -#if !defined(MAP_HUGETLB) - if (huge_pages == HUGE_PAGES_ON) - ereport(ERROR, - (errcode(ERRCODE_FEATURE_NOT_SUPPORTED), - errmsg("huge pages not supported on this platform"))); -#endif - /* For now, we don't support huge pages in SysV memory */ if (huge_pages == HUGE_PAGES_ON && shared_memory_type != SHMEM_TYPE_MMAP) ereport(ERROR, @@ -782,7 +994,7 @@ PGSharedMemoryCreate(Size size, if (shared_memory_type == SHMEM_TYPE_MMAP) { /* On success, mapping data will be modified. */ - CreateAnonymousSegment(mapping); + CreateAnonymousSegment(mapping, base); next_free_segment++; diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c index 4dee856d6bd..ce719f1b412 100644 --- a/src/backend/port/win32_shmem.c +++ b/src/backend/port/win32_shmem.c @@ -205,7 +205,7 @@ EnableLockPagesPrivilege(int elevel) */ PGShmemHeader * PGSharedMemoryCreate(Size size, - PGShmemHeader **shim) + PGShmemHeader **shim, Pointer base) { void *memAddress; PGShmemHeader *hdr; diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c index 8b38e985327..076888c0172 100644 --- a/src/backend/storage/ipc/ipci.c +++ b/src/backend/storage/ipc/ipci.c @@ -203,9 +203,12 @@ CreateSharedMemoryAndSemaphores(void) PGShmemHeader *seghdr; Size size; int numSemas; + void *base; Assert(!IsUnderPostmaster); + base = ReserveAnonymousMemory((Size) MaxAvailableMemory * BLCKSZ); + for(int segment = 0; segment < ANON_MAPPINGS; segment++) { /* Compute the size of the shared-memory block */ @@ -217,7 +220,7 @@ CreateSharedMemoryAndSemaphores(void) * * XXX: Do multiple shims are needed, one per segment? */ - seghdr = PGSharedMemoryCreate(size, &shim); + seghdr = PGSharedMemoryCreate(size, &shim, base); /* * Make sure that huge pages are never reported as "unknown" while the diff --git a/src/backend/utils/init/globals.c b/src/backend/utils/init/globals.c index 2152aad97d9..1d42a5856c0 100644 --- a/src/backend/utils/init/globals.c +++ b/src/backend/utils/init/globals.c @@ -140,6 +140,7 @@ int max_parallel_maintenance_workers = 2; * register background workers. */ int NBuffers = 16384; +int MaxAvailableMemory = 131072; int MaxConnections = 100; int max_worker_processes = 8; int max_parallel_workers = 8; diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c index 4eaeca89f2c..dede37f7905 100644 --- a/src/backend/utils/misc/guc_tables.c +++ b/src/backend/utils/misc/guc_tables.c @@ -2364,6 +2364,20 @@ struct config_int ConfigureNamesInt[] = NULL, NULL, NULL }, + { + {"max_available_memory", PGC_SIGHUP, RESOURCES_MEM, + gettext_noop("Sets the upper limit for the shared_buffers value."), + gettext_noop("Shared memory could be resized at runtime, this " + "parameters sets the upper limit for it, beyond which " + "resizing would not be supported. Normally this value " + "would be the same as the total available memory."), + GUC_UNIT_BLOCKS + }, + &MaxAvailableMemory, + 131072, 16, INT_MAX / 2, + NULL, NULL, NULL + }, + { {"vacuum_buffer_usage_limit", PGC_USERSET, RESOURCES_MEM, gettext_noop("Sets the buffer pool size for VACUUM, ANALYZE, and autovacuum."), diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h index 138078c29c5..4a83e255652 100644 --- a/src/include/storage/pg_shmem.h +++ b/src/include/storage/pg_shmem.h @@ -60,6 +60,7 @@ extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS]; extern PGDLLIMPORT int shared_memory_type; extern PGDLLIMPORT int huge_pages; extern PGDLLIMPORT int huge_page_size; +extern PGDLLIMPORT int MaxAvailableMemory; /* Possible values for huge_pages and huge_pages_status */ typedef enum @@ -100,10 +101,11 @@ extern void PGSharedMemoryNoReAttach(void); #endif extern PGShmemHeader *PGSharedMemoryCreate(Size size, - PGShmemHeader **shim); + PGShmemHeader **shim, Pointer base); extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2); extern void PGSharedMemoryDetach(void); extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags); +void *ReserveAnonymousMemory(Size reserve_size); /* The main segment, contains everything except buffer blocks and related data. */ #define MAIN_SHMEM_SEGMENT 0 -- 2.45.1 --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="v4-0003-Introduce-multiple-shmem-segments-for-shared-buff.patch" From dca1257476fc4c0718fec35b11ba0f4a4e57151b Mon Sep 17 00:00:00 2001 From: Dmitrii Dolgov <9erthalion6@gmail.com> Date: Sat, 15 Mar 2025 16:38:59 +0100 Subject: [PATCH v4 3/8] Introduce multiple shmem segments for shared buffers Add more shmem segments to split shared buffers into following chunks: * BUFFERS_SHMEM_SEGMENT: contains buffer blocks * BUFFER_DESCRIPTORS_SHMEM_SEGMENT: contains buffer descriptors * BUFFER_IOCV_SHMEM_SEGMENT: contains condition variables for buffers * CHECKPOINT_BUFFERS_SHMEM_SEGMENT: contains checkpoint buffer ids * STRATEGY_SHMEM_SEGMENT: contains buffer strategy status Size of the corresponding shared data directly depends on NBuffers, meaning that if we would like to change NBuffers, they have to be resized correspondingly. Placing each of them in a separate shmem segment allows to achieve that. There are some asumptions made about each of shmem segments upper size limit. The buffer blocks have the largest, while the rest claim less extra room for resize. Ideally those limits have to be deduced from the maximum allowed shared memory. --- src/backend/port/sysv_shmem.c | 24 +++++++- src/backend/storage/buffer/buf_init.c | 79 +++++++++++++++++--------- src/backend/storage/buffer/buf_table.c | 6 +- src/backend/storage/buffer/freelist.c | 5 +- src/backend/storage/ipc/ipci.c | 2 +- src/include/storage/bufmgr.h | 2 +- src/include/storage/pg_shmem.h | 24 +++++++- 7 files changed, 105 insertions(+), 37 deletions(-) diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c index a0f03ff868f..f46d9d5d9cd 100644 --- a/src/backend/port/sysv_shmem.c +++ b/src/backend/port/sysv_shmem.c @@ -147,10 +147,18 @@ static int next_free_segment = 0; * * The reserved space pointer is calculated to slice up the total reserved * space into fixed fractions of address space for each segment, as specified - * in the SHMEM_RESIZE_RATIO array. + * in the SHMEM_RESIZE_RATIO array. E.g. we allow BUFFERS_SHMEM_SEGMENT to take + * up to 60% of the whole space when resizing, based on the fact that it most + * likely will be the main consumer of this memory. Those numbers are pulled + * out of thin air for now, makes sense to evaluate them more precise. */ -static double SHMEM_RESIZE_RATIO[1] = { - 1.0, /* MAIN_SHMEM_SLOT */ +static double SHMEM_RESIZE_RATIO[6] = { + 0.1, /* MAIN_SHMEM_SEGMENT */ + 0.6, /* BUFFERS_SHMEM_SEGMENT */ + 0.1, /* BUFFER_DESCRIPTORS_SHMEM_SEGMENT */ + 0.1, /* BUFFER_IOCV_SHMEM_SEGMENT */ + 0.05, /* CHECKPOINT_BUFFERS_SHMEM_SEGMENT */ + 0.05, /* STRATEGY_SHMEM_SEGMENT */ }; /* @@ -182,6 +190,16 @@ MappingName(int shmem_segment) { case MAIN_SHMEM_SEGMENT: return "main"; + case BUFFERS_SHMEM_SEGMENT: + return "buffers"; + case BUFFER_DESCRIPTORS_SHMEM_SEGMENT: + return "descriptors"; + case BUFFER_IOCV_SHMEM_SEGMENT: + return "iocv"; + case CHECKPOINT_BUFFERS_SHMEM_SEGMENT: + return "checkpoint"; + case STRATEGY_SHMEM_SEGMENT: + return "strategy"; default: return "unknown"; } diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c index ed1dc488a42..bd68b69ee98 100644 --- a/src/backend/storage/buffer/buf_init.c +++ b/src/backend/storage/buffer/buf_init.c @@ -62,7 +62,10 @@ CkptSortItem *CkptBufferIds; * Initialize shared buffer pool * * This is called once during shared-memory initialization (either in the - * postmaster, or in a standalone backend). + * postmaster, or in a standalone backend). Size of data structures initialized + * here depends on NBuffers, and to be able to change NBuffers without a + * restart we store each structure into a separate shared memory segment, which + * could be resized on demand. */ void BufferManagerShmemInit(void) @@ -74,22 +77,22 @@ BufferManagerShmemInit(void) /* Align descriptors to a cacheline boundary. */ BufferDescriptors = (BufferDescPadded *) - ShmemInitStruct("Buffer Descriptors", + ShmemInitStructInSegment("Buffer Descriptors", NBuffers * sizeof(BufferDescPadded), - &foundDescs); + &foundDescs, BUFFER_DESCRIPTORS_SHMEM_SEGMENT); /* Align buffer pool on IO page size boundary. */ BufferBlocks = (char *) TYPEALIGN(PG_IO_ALIGN_SIZE, - ShmemInitStruct("Buffer Blocks", + ShmemInitStructInSegment("Buffer Blocks", NBuffers * (Size) BLCKSZ + PG_IO_ALIGN_SIZE, - &foundBufs)); + &foundBufs, BUFFERS_SHMEM_SEGMENT)); /* Align condition variables to cacheline boundary. */ BufferIOCVArray = (ConditionVariableMinimallyPadded *) - ShmemInitStruct("Buffer IO Condition Variables", + ShmemInitStructInSegment("Buffer IO Condition Variables", NBuffers * sizeof(ConditionVariableMinimallyPadded), - &foundIOCV); + &foundIOCV, BUFFER_IOCV_SHMEM_SEGMENT); /* * The array used to sort to-be-checkpointed buffer ids is located in @@ -99,8 +102,9 @@ BufferManagerShmemInit(void) * painful. */ CkptBufferIds = (CkptSortItem *) - ShmemInitStruct("Checkpoint BufferIds", - NBuffers * sizeof(CkptSortItem), &foundBufCkpt); + ShmemInitStructInSegment("Checkpoint BufferIds", + NBuffers * sizeof(CkptSortItem), &foundBufCkpt, + CHECKPOINT_BUFFERS_SHMEM_SEGMENT); if (foundDescs || foundBufs || foundIOCV || foundBufCkpt) { @@ -156,33 +160,54 @@ BufferManagerShmemInit(void) * BufferManagerShmemSize * * compute the size of shared memory for the buffer pool including - * data pages, buffer descriptors, hash tables, etc. + * data pages, buffer descriptors, hash tables, etc. based on the + * shared memory segment. The main segment must not allocate anything + * related to buffers, every other segment will receive part of the + * data. */ Size -BufferManagerShmemSize(void) +BufferManagerShmemSize(int shmem_segment) { Size size = 0; - /* size of buffer descriptors */ - size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded))); - /* to allow aligning buffer descriptors */ - size = add_size(size, PG_CACHE_LINE_SIZE); + if (shmem_segment == MAIN_SHMEM_SEGMENT) + return size; - /* size of data pages, plus alignment padding */ - size = add_size(size, PG_IO_ALIGN_SIZE); - size = add_size(size, mul_size(NBuffers, BLCKSZ)); + if (shmem_segment == BUFFER_DESCRIPTORS_SHMEM_SEGMENT) + { + /* size of buffer descriptors */ + size = add_size(size, mul_size(NBuffers, sizeof(BufferDescPadded))); + /* to allow aligning buffer descriptors */ + size = add_size(size, PG_CACHE_LINE_SIZE); + } - /* size of stuff controlled by freelist.c */ - size = add_size(size, StrategyShmemSize()); + if (shmem_segment == BUFFERS_SHMEM_SEGMENT) + { + /* size of data pages, plus alignment padding */ + size = add_size(size, PG_IO_ALIGN_SIZE); + size = add_size(size, mul_size(NBuffers, BLCKSZ)); + } - /* size of I/O condition variables */ - size = add_size(size, mul_size(NBuffers, - sizeof(ConditionVariableMinimallyPadded))); - /* to allow aligning the above */ - size = add_size(size, PG_CACHE_LINE_SIZE); + if (shmem_segment == STRATEGY_SHMEM_SEGMENT) + { + /* size of stuff controlled by freelist.c */ + size = add_size(size, StrategyShmemSize()); + } - /* size of checkpoint sort array in bufmgr.c */ - size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem))); + if (shmem_segment == BUFFER_IOCV_SHMEM_SEGMENT) + { + /* size of I/O condition variables */ + size = add_size(size, mul_size(NBuffers, + sizeof(ConditionVariableMinimallyPadded))); + /* to allow aligning the above */ + size = add_size(size, PG_CACHE_LINE_SIZE); + } + + if (shmem_segment == CHECKPOINT_BUFFERS_SHMEM_SEGMENT) + { + /* size of checkpoint sort array in bufmgr.c */ + size = add_size(size, mul_size(NBuffers, sizeof(CkptSortItem))); + } return size; } diff --git a/src/backend/storage/buffer/buf_table.c b/src/backend/storage/buffer/buf_table.c index a50955d5286..a9952b36eba 100644 --- a/src/backend/storage/buffer/buf_table.c +++ b/src/backend/storage/buffer/buf_table.c @@ -22,6 +22,7 @@ #include "postgres.h" #include "storage/buf_internals.h" +#include "storage/pg_shmem.h" /* entry for buffer lookup hashtable */ typedef struct @@ -59,10 +60,11 @@ InitBufTable(int size) info.entrysize = sizeof(BufferLookupEnt); info.num_partitions = NUM_BUFFER_PARTITIONS; - SharedBufHash = ShmemInitHash("Shared Buffer Lookup Table", + SharedBufHash = ShmemInitHashInSegment("Shared Buffer Lookup Table", size, size, &info, - HASH_ELEM | HASH_BLOBS | HASH_PARTITION); + HASH_ELEM | HASH_BLOBS | HASH_PARTITION, + STRATEGY_SHMEM_SEGMENT); } /* diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index 336715b6c63..81543cb5ced 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -19,6 +19,7 @@ #include "port/atomics.h" #include "storage/buf_internals.h" #include "storage/bufmgr.h" +#include "storage/pg_shmem.h" #include "storage/proc.h" #define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var)))) @@ -491,9 +492,9 @@ StrategyInitialize(bool init) * Get or create the shared strategy control block */ StrategyControl = (BufferStrategyControl *) - ShmemInitStruct("Buffer Strategy Status", + ShmemInitStructInSegment("Buffer Strategy Status", sizeof(BufferStrategyControl), - &found); + &found, STRATEGY_SHMEM_SEGMENT); if (!found) { diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c index 076888c0172..9d00b80b4f8 100644 --- a/src/backend/storage/ipc/ipci.c +++ b/src/backend/storage/ipc/ipci.c @@ -113,7 +113,7 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment) sizeof(ShmemIndexEnt))); size = add_size(size, dsm_estimate_size()); size = add_size(size, DSMRegistryShmemSize()); - size = add_size(size, BufferManagerShmemSize()); + size = add_size(size, BufferManagerShmemSize(shmem_segment)); size = add_size(size, LockManagerShmemSize()); size = add_size(size, PredicateLockShmemSize()); size = add_size(size, ProcGlobalShmemSize()); diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h index f2192ceb271..1977001e533 100644 --- a/src/include/storage/bufmgr.h +++ b/src/include/storage/bufmgr.h @@ -308,7 +308,7 @@ extern bool EvictUnpinnedBuffer(Buffer buf); /* in buf_init.c */ extern void BufferManagerShmemInit(void); -extern Size BufferManagerShmemSize(void); +extern Size BufferManagerShmemSize(int); /* in localbuf.c */ extern void AtProcExit_LocalBuffers(void); diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h index 4a83e255652..c5009a1cd73 100644 --- a/src/include/storage/pg_shmem.h +++ b/src/include/storage/pg_shmem.h @@ -52,7 +52,7 @@ typedef struct ShmemSegment } ShmemSegment; /* Number of available segments for anonymous memory mappings */ -#define ANON_MAPPINGS 1 +#define ANON_MAPPINGS 6 extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS]; @@ -107,7 +107,29 @@ extern void PGSharedMemoryDetach(void); extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags); void *ReserveAnonymousMemory(Size reserve_size); +/* + * To be able to dynamically resize largest parts of the data stored in shared + * memory, we split it into multiple shared memory mappings segments. Each + * segment contains only certain part of the data, which size depends on + * NBuffers. + */ + /* The main segment, contains everything except buffer blocks and related data. */ #define MAIN_SHMEM_SEGMENT 0 +/* Buffer blocks */ +#define BUFFERS_SHMEM_SEGMENT 1 + +/* Buffer descriptors */ +#define BUFFER_DESCRIPTORS_SHMEM_SEGMENT 2 + +/* Condition variables for buffers */ +#define BUFFER_IOCV_SHMEM_SEGMENT 3 + +/* Checkpoint BufferIds */ +#define CHECKPOINT_BUFFERS_SHMEM_SEGMENT 4 + +/* Buffer strategy status */ +#define STRATEGY_SHMEM_SEGMENT 5 + #endif /* PG_SHMEM_H */ -- 2.45.1 --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="v4-0004-Introduce-pending-flag-for-GUC-assign-hooks.patch" From 24704e57aea0ee94fbfb37ca3a4ea4fcf050a738 Mon Sep 17 00:00:00 2001 From: Dmitrii Dolgov <9erthalion6@gmail.com> Date: Sun, 6 Apr 2025 16:40:32 +0200 Subject: [PATCH v4 4/8] Introduce pending flag for GUC assign hooks Currently an assing hook can perform some preprocessing of a new value, but it cannot change the behavior, which dictates that the new value will be applied immediately after the hook. Certain GUC options (like shared_buffers, coming in subsequent patches) may need coordinating work between backends to change, meaning we cannot apply it right away. Add a new flag "pending" for an assign hook to allow the hook indicate exactly that. If the pending flag is set after the hook, the new value will not be applied and it's handling becomes the hook's implementation responsibility. Note, that this also requires changes in the way how GUCs are getting reported, but the patch does not cover that yet. --- src/backend/access/transam/xlog.c | 2 +- src/backend/commands/variable.c | 6 +-- src/backend/libpq/pqcomm.c | 8 ++-- src/backend/tcop/postgres.c | 2 +- src/backend/utils/misc/guc.c | 59 +++++++++++++++++++--------- src/backend/utils/misc/stack_depth.c | 2 +- src/include/utils/guc.h | 2 +- src/include/utils/guc_hooks.h | 20 +++++----- 8 files changed, 61 insertions(+), 40 deletions(-) diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c index ec40c0b7c42..9aa426992a2 100644 --- a/src/backend/access/transam/xlog.c +++ b/src/backend/access/transam/xlog.c @@ -2321,7 +2321,7 @@ CalculateCheckpointSegments(void) } void -assign_max_wal_size(int newval, void *extra) +assign_max_wal_size(int newval, void *extra, bool *pending) { max_wal_size_mb = newval; CalculateCheckpointSegments(); diff --git a/src/backend/commands/variable.c b/src/backend/commands/variable.c index a9f2a3a3062..e715a6f01c2 100644 --- a/src/backend/commands/variable.c +++ b/src/backend/commands/variable.c @@ -1143,7 +1143,7 @@ check_cluster_name(char **newval, void **extra, GucSource source) * GUC assign_hook for maintenance_io_concurrency */ void -assign_maintenance_io_concurrency(int newval, void *extra) +assign_maintenance_io_concurrency(int newval, void *extra, bool *pending) { /* * Reconfigure recovery prefetching, because a setting it depends on @@ -1161,13 +1161,13 @@ assign_maintenance_io_concurrency(int newval, void *extra) * they may be assigned in either order. */ void -assign_io_max_combine_limit(int newval, void *extra) +assign_io_max_combine_limit(int newval, void *extra, bool *pending) { io_max_combine_limit = newval; io_combine_limit = Min(io_max_combine_limit, io_combine_limit_guc); } void -assign_io_combine_limit(int newval, void *extra) +assign_io_combine_limit(int newval, void *extra, bool *pending) { io_combine_limit_guc = newval; io_combine_limit = Min(io_max_combine_limit, io_combine_limit_guc); diff --git a/src/backend/libpq/pqcomm.c b/src/backend/libpq/pqcomm.c index e5171467de1..2a6a587ef76 100644 --- a/src/backend/libpq/pqcomm.c +++ b/src/backend/libpq/pqcomm.c @@ -1952,7 +1952,7 @@ pq_settcpusertimeout(int timeout, Port *port) * GUC assign_hook for tcp_keepalives_idle */ void -assign_tcp_keepalives_idle(int newval, void *extra) +assign_tcp_keepalives_idle(int newval, void *extra, bool *pending) { /* * The kernel API provides no way to test a value without setting it; and @@ -1985,7 +1985,7 @@ show_tcp_keepalives_idle(void) * GUC assign_hook for tcp_keepalives_interval */ void -assign_tcp_keepalives_interval(int newval, void *extra) +assign_tcp_keepalives_interval(int newval, void *extra, bool *pending) { /* See comments in assign_tcp_keepalives_idle */ (void) pq_setkeepalivesinterval(newval, MyProcPort); @@ -2008,7 +2008,7 @@ show_tcp_keepalives_interval(void) * GUC assign_hook for tcp_keepalives_count */ void -assign_tcp_keepalives_count(int newval, void *extra) +assign_tcp_keepalives_count(int newval, void *extra, bool *pending) { /* See comments in assign_tcp_keepalives_idle */ (void) pq_setkeepalivescount(newval, MyProcPort); @@ -2031,7 +2031,7 @@ show_tcp_keepalives_count(void) * GUC assign_hook for tcp_user_timeout */ void -assign_tcp_user_timeout(int newval, void *extra) +assign_tcp_user_timeout(int newval, void *extra, bool *pending) { /* See comments in assign_tcp_keepalives_idle */ (void) pq_settcpusertimeout(newval, MyProcPort); diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c index 6ae9f38f0c8..b1fba850f02 100644 --- a/src/backend/tcop/postgres.c +++ b/src/backend/tcop/postgres.c @@ -3593,7 +3593,7 @@ check_log_stats(bool *newval, void **extra, GucSource source) /* GUC assign hook for transaction_timeout */ void -assign_transaction_timeout(int newval, void *extra) +assign_transaction_timeout(int newval, void *extra, bool *pending) { if (IsTransactionState()) { diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c index 667df448732..bb681f5bc60 100644 --- a/src/backend/utils/misc/guc.c +++ b/src/backend/utils/misc/guc.c @@ -1679,6 +1679,7 @@ InitializeOneGUCOption(struct config_generic *gconf) struct config_int *conf = (struct config_int *) gconf; int newval = conf->boot_val; void *extra = NULL; + bool pending = false; Assert(newval >= conf->min); Assert(newval <= conf->max); @@ -1687,9 +1688,13 @@ InitializeOneGUCOption(struct config_generic *gconf) elog(FATAL, "failed to initialize %s to %d", conf->gen.name, newval); if (conf->assign_hook) - conf->assign_hook(newval, extra); - *conf->variable = conf->reset_val = newval; - conf->gen.extra = conf->reset_extra = extra; + conf->assign_hook(newval, extra, &pending); + + if (!pending) + { + *conf->variable = conf->reset_val = newval; + conf->gen.extra = conf->reset_extra = extra; + } break; } case PGC_REAL: @@ -2041,13 +2046,18 @@ ResetAllOptions(void) case PGC_INT: { struct config_int *conf = (struct config_int *) gconf; + bool pending = false; if (conf->assign_hook) conf->assign_hook(conf->reset_val, - conf->reset_extra); - *conf->variable = conf->reset_val; - set_extra_field(&conf->gen, &conf->gen.extra, - conf->reset_extra); + conf->reset_extra, + &pending); + if (!pending) + { + *conf->variable = conf->reset_val; + set_extra_field(&conf->gen, &conf->gen.extra, + conf->reset_extra); + } break; } case PGC_REAL: @@ -2424,16 +2434,21 @@ AtEOXact_GUC(bool isCommit, int nestLevel) struct config_int *conf = (struct config_int *) gconf; int newval = newvalue.val.intval; void *newextra = newvalue.extra; + bool pending = false; if (*conf->variable != newval || conf->gen.extra != newextra) { if (conf->assign_hook) - conf->assign_hook(newval, newextra); - *conf->variable = newval; - set_extra_field(&conf->gen, &conf->gen.extra, - newextra); - changed = true; + conf->assign_hook(newval, newextra, &pending); + + if (!pending) + { + *conf->variable = newval; + set_extra_field(&conf->gen, &conf->gen.extra, + newextra); + changed = true; + } } break; } @@ -3850,18 +3865,24 @@ set_config_with_handle(const char *name, config_handle *handle, if (changeVal) { + bool pending = false; + /* Save old value to support transaction abort */ if (!makeDefault) push_old_value(&conf->gen, action); if (conf->assign_hook) - conf->assign_hook(newval, newextra); - *conf->variable = newval; - set_extra_field(&conf->gen, &conf->gen.extra, - newextra); - set_guc_source(&conf->gen, source); - conf->gen.scontext = context; - conf->gen.srole = srole; + conf->assign_hook(newval, newextra, &pending); + + if (!pending) + { + *conf->variable = newval; + set_extra_field(&conf->gen, &conf->gen.extra, + newextra); + set_guc_source(&conf->gen, source); + conf->gen.scontext = context; + conf->gen.srole = srole; + } } if (makeDefault) { diff --git a/src/backend/utils/misc/stack_depth.c b/src/backend/utils/misc/stack_depth.c index 8f7cf531fbc..ef59ae62008 100644 --- a/src/backend/utils/misc/stack_depth.c +++ b/src/backend/utils/misc/stack_depth.c @@ -156,7 +156,7 @@ check_max_stack_depth(int *newval, void **extra, GucSource source) /* GUC assign hook for max_stack_depth */ void -assign_max_stack_depth(int newval, void *extra) +assign_max_stack_depth(int newval, void *extra, bool *pending) { ssize_t newval_bytes = newval * (ssize_t) 1024; diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h index f619100467d..8802ad8a3cb 100644 --- a/src/include/utils/guc.h +++ b/src/include/utils/guc.h @@ -187,7 +187,7 @@ typedef bool (*GucStringCheckHook) (char **newval, void **extra, GucSource sourc typedef bool (*GucEnumCheckHook) (int *newval, void **extra, GucSource source); typedef void (*GucBoolAssignHook) (bool newval, void *extra); -typedef void (*GucIntAssignHook) (int newval, void *extra); +typedef void (*GucIntAssignHook) (int newval, void *extra, bool *pending); typedef void (*GucRealAssignHook) (double newval, void *extra); typedef void (*GucStringAssignHook) (const char *newval, void *extra); typedef void (*GucEnumAssignHook) (int newval, void *extra); diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h index 799fa7ace68..c8300cffa8e 100644 --- a/src/include/utils/guc_hooks.h +++ b/src/include/utils/guc_hooks.h @@ -81,14 +81,14 @@ extern bool check_log_stats(bool *newval, void **extra, GucSource source); extern bool check_log_timezone(char **newval, void **extra, GucSource source); extern void assign_log_timezone(const char *newval, void *extra); extern const char *show_log_timezone(void); -extern void assign_maintenance_io_concurrency(int newval, void *extra); -extern void assign_io_max_combine_limit(int newval, void *extra); -extern void assign_io_combine_limit(int newval, void *extra); +extern void assign_maintenance_io_concurrency(int newval, void *extra, bool *pending); +extern void assign_io_max_combine_limit(int newval, void *extra, bool *pending); +extern void assign_io_combine_limit(int newval, void *extra, bool *pending); extern bool check_max_slot_wal_keep_size(int *newval, void **extra, GucSource source); -extern void assign_max_wal_size(int newval, void *extra); +extern void assign_max_wal_size(int newval, void *extra, bool *pending); extern bool check_max_stack_depth(int *newval, void **extra, GucSource source); -extern void assign_max_stack_depth(int newval, void *extra); +extern void assign_max_stack_depth(int newval, void *extra, bool *pending); extern bool check_multixact_member_buffers(int *newval, void **extra, GucSource source); extern bool check_multixact_offset_buffers(int *newval, void **extra, @@ -143,13 +143,13 @@ extern void assign_synchronous_standby_names(const char *newval, void *extra); extern void assign_synchronous_commit(int newval, void *extra); extern void assign_syslog_facility(int newval, void *extra); extern void assign_syslog_ident(const char *newval, void *extra); -extern void assign_tcp_keepalives_count(int newval, void *extra); +extern void assign_tcp_keepalives_count(int newval, void *extra, bool *pending); extern const char *show_tcp_keepalives_count(void); -extern void assign_tcp_keepalives_idle(int newval, void *extra); +extern void assign_tcp_keepalives_idle(int newval, void *extra, bool *pending); extern const char *show_tcp_keepalives_idle(void); -extern void assign_tcp_keepalives_interval(int newval, void *extra); +extern void assign_tcp_keepalives_interval(int newval, void *extra, bool *pending); extern const char *show_tcp_keepalives_interval(void); -extern void assign_tcp_user_timeout(int newval, void *extra); +extern void assign_tcp_user_timeout(int newval, void *extra, bool *pending); extern const char *show_tcp_user_timeout(void); extern bool check_temp_buffers(int *newval, void **extra, GucSource source); extern bool check_temp_tablespaces(char **newval, void **extra, @@ -165,7 +165,7 @@ extern bool check_transaction_buffers(int *newval, void **extra, GucSource sourc extern bool check_transaction_deferrable(bool *newval, void **extra, GucSource source); extern bool check_transaction_isolation(int *newval, void **extra, GucSource source); extern bool check_transaction_read_only(bool *newval, void **extra, GucSource source); -extern void assign_transaction_timeout(int newval, void *extra); +extern void assign_transaction_timeout(int newval, void *extra, bool *pending); extern const char *show_unix_socket_permissions(void); extern bool check_wal_buffers(int *newval, void **extra, GucSource source); extern bool check_wal_consistency_checking(char **newval, void **extra, -- 2.45.1 --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="v4-0005-Introduce-pss_barrierReceivedGeneration.patch" From 619b10ec409185995a4a3ffd56972f1efa493c45 Mon Sep 17 00:00:00 2001 From: Dmitrii Dolgov <9erthalion6@gmail.com> Date: Fri, 4 Apr 2025 21:46:14 +0200 Subject: [PATCH v4 5/8] Introduce pss_barrierReceivedGeneration Currently WaitForProcSignalBarrier allows to make sure the message sent via EmitProcSignalBarrier was processed by all ProcSignal mechanism participants. Add pss_barrierReceivedGeneration alongside with pss_barrierGeneration, which will be updated when a process has received the message, but not processed it yet. This makes it possible to support a new mode of waiting, when ProcSignal participants want to synchronize message processing. To do that, a participant can wait via WaitForProcSignalBarrierReceived when processing a message, effectively making sure that all processes are going to start processing ProcSignalBarrier simultaneously. --- src/backend/storage/ipc/procsignal.c | 67 ++++++++++++++++++++++------ src/include/storage/procsignal.h | 1 + 2 files changed, 54 insertions(+), 14 deletions(-) diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c index b7c39a4c5f0..8e313ad9bf8 100644 --- a/src/backend/storage/ipc/procsignal.c +++ b/src/backend/storage/ipc/procsignal.c @@ -58,7 +58,10 @@ * of it. For such use cases, we set a bit in pss_barrierCheckMask and then * increment the current "barrier generation"; when the new barrier generation * (or greater) appears in the pss_barrierGeneration flag of every process, - * we know that the message has been received everywhere. + * we know that the message has been received and processed everywhere. In case + * if we only need to know only that the message was received everywhere (e.g. + * receiving processes need to handle the message in a coordinated fashion) + * use pss_barrierReceivedGeneration in the same way. */ typedef struct { @@ -70,6 +73,7 @@ typedef struct /* Barrier-related fields (not protected by pss_mutex) */ pg_atomic_uint64 pss_barrierGeneration; + pg_atomic_uint64 pss_barrierReceivedGeneration; pg_atomic_uint32 pss_barrierCheckMask; ConditionVariable pss_barrierCV; } ProcSignalSlot; @@ -151,6 +155,8 @@ ProcSignalShmemInit(void) slot->pss_cancel_key_len = 0; MemSet(slot->pss_signalFlags, 0, sizeof(slot->pss_signalFlags)); pg_atomic_init_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX); + pg_atomic_init_u64(&slot->pss_barrierReceivedGeneration, + PG_UINT64_MAX); pg_atomic_init_u32(&slot->pss_barrierCheckMask, 0); ConditionVariableInit(&slot->pss_barrierCV); } @@ -198,6 +204,8 @@ ProcSignalInit(char *cancel_key, int cancel_key_len) barrier_generation = pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration); pg_atomic_write_u64(&slot->pss_barrierGeneration, barrier_generation); + pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, + barrier_generation); if (cancel_key_len > 0) memcpy(slot->pss_cancel_key, cancel_key, cancel_key_len); @@ -262,6 +270,7 @@ CleanupProcSignalState(int status, Datum arg) * no barrier waits block on it. */ pg_atomic_write_u64(&slot->pss_barrierGeneration, PG_UINT64_MAX); + pg_atomic_write_u64(&slot->pss_barrierReceivedGeneration, PG_UINT64_MAX); SpinLockRelease(&slot->pss_mutex); @@ -415,12 +424,8 @@ EmitProcSignalBarrier(ProcSignalBarrierType type) return generation; } -/* - * WaitForProcSignalBarrier - wait until it is guaranteed that all changes - * requested by a specific call to EmitProcSignalBarrier() have taken effect. - */ -void -WaitForProcSignalBarrier(uint64 generation) +static void +WaitForProcSignalBarrierInternal(uint64 generation, bool receivedOnly) { Assert(generation <= pg_atomic_read_u64(&ProcSignal->psh_barrierGeneration)); @@ -435,12 +440,17 @@ WaitForProcSignalBarrier(uint64 generation) uint64 oldval; /* - * It's important that we check only pss_barrierGeneration here and - * not pss_barrierCheckMask. Bits in pss_barrierCheckMask get cleared - * before the barrier is actually absorbed, but pss_barrierGeneration + * It's important that we check only pss_barrierGeneration & + * pss_barrierGeneration here and not pss_barrierCheckMask. Bits in + * pss_barrierCheckMask get cleared before the barrier is actually + * absorbed, but pss_barrierGeneration & pss_barrierReceivedGeneration * is updated only afterward. */ - oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration); + if (receivedOnly) + oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration); + else + oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration); + while (oldval < generation) { if (ConditionVariableTimedSleep(&slot->pss_barrierCV, @@ -449,7 +459,11 @@ WaitForProcSignalBarrier(uint64 generation) ereport(LOG, (errmsg("still waiting for backend with PID %d to accept ProcSignalBarrier", (int) pg_atomic_read_u32(&slot->pss_pid)))); - oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration); + + if (receivedOnly) + oldval = pg_atomic_read_u64(&slot->pss_barrierReceivedGeneration); + else + oldval = pg_atomic_read_u64(&slot->pss_barrierGeneration); } ConditionVariableCancelSleep(); } @@ -463,12 +477,33 @@ WaitForProcSignalBarrier(uint64 generation) * The caller is probably calling this function because it wants to read * the shared state or perform further writes to shared state once all * backends are known to have absorbed the barrier. However, the read of - * pss_barrierGeneration was performed unlocked; insert a memory barrier - * to separate it from whatever follows. + * pss_barrierGeneration & pss_barrierReceivedGeneration was performed + * unlocked; insert a memory barrier to separate it from whatever follows. */ pg_memory_barrier(); } +/* + * WaitForProcSignalBarrier - wait until it is guaranteed that all changes + * requested by a specific call to EmitProcSignalBarrier() have taken effect. + */ +void +WaitForProcSignalBarrier(uint64 generation) +{ + WaitForProcSignalBarrierInternal(generation, false); +} + +/* + * WaitForProcSignalBarrierReceived - wait until it is guaranteed that all + * backends have observed the message sent by a specific call to + * EmitProcSignalBarrier(). + */ +void +WaitForProcSignalBarrierReceived(uint64 generation) +{ + WaitForProcSignalBarrierInternal(generation, true); +} + /* * Handle receipt of an interrupt indicating a global barrier event. * @@ -522,6 +557,10 @@ ProcessProcSignalBarrier(void) if (local_gen == shared_gen) return; + /* The message is observed, record that */ + pg_atomic_write_u64(&MyProcSignalSlot->pss_barrierReceivedGeneration, + shared_gen); + /* * Get and clear the flags that are set for this backend. Note that * pg_atomic_exchange_u32 is a full barrier, so we're guaranteed that the diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h index 016dfd9b3f6..defd8b66a19 100644 --- a/src/include/storage/procsignal.h +++ b/src/include/storage/procsignal.h @@ -79,6 +79,7 @@ extern void SendCancelRequest(int backendPID, char *cancel_key, int cancel_key_l extern uint64 EmitProcSignalBarrier(ProcSignalBarrierType type); extern void WaitForProcSignalBarrier(uint64 generation); +extern void WaitForProcSignalBarrierReceived(uint64 generation); extern void ProcessProcSignalBarrier(void); extern void procsignal_sigusr1_handler(SIGNAL_ARGS); -- 2.45.1 --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="v4-0006-Allow-to-resize-shared-memory-without-restart.patch" From 886a3ea87408e628bea08a9c77116343616ad032 Mon Sep 17 00:00:00 2001 From: Dmitrii Dolgov <9erthalion6@gmail.com> Date: Sun, 6 Apr 2025 16:47:16 +0200 Subject: [PATCH v4 6/8] Allow to resize shared memory without restart Add assing hook for shared_buffers to resize shared memory using space, introduced in the previous commits without requiring PostgreSQL restart. Essentially the implementation is based on two mechanisms: a ProcSignalBarrier is used to make sure all processes are starting the resize procedure simultaneously, and a global Barrier is used to coordinate after that and make sure all finished processes are waiting for others that are in progress. The resize process looks like this: * The GUC assign hook sets a flag to let the Postmaster know that resize was requested. * Postmaster verifies the flag in the event loop, and starts the resize by emitting a ProcSignal barrier. * All processes, that participate in ProcSignal mechanism, begin to process ProcSignal barrier. First a process waits until all processes have confirmed they received the message and can start simultaneously. * Every process recalculates shared memory size based on the new NBuffers and extend it using mremap. One elected process signals the postmaster to do the same. * When finished, every process waits on a global ShmemControl barrier, untill all others are finished as well. This way we ensure three stages with clear boundaries: before the resize, when all processes use old NBuffers; during the resize, when processes have mix of old and new NBuffers, and wait until it's done; after the resize, when all processes use new NBuffers. * After all processes are using new value, one of them will initialize new shared structures (buffer blocks, descriptors, etc) as needed and broadcast new value of NBuffers via ShmemControl in shared memory. Other backends are waiting for this operation to finish as well. Then the barrier is lifted and everything goes as usual. Since resizing takes time, we need to take into account that during that time: - New backends can be spawned. They will check status of the barrier early during the bootstrap, and wait until everything is over to work with the new NBuffers value. - Old backends can exit before attempting to resize. Synchronization used between backends relies on ProcSignalBarrier and waits for all participants received the message at the beginning to gather all existing backends. - Some backends might be blocked and not responsing either before or after receiving the message. In the first case such backend still have ProcSignalSlot and should be waited for, in the second case shared barrier will make sure we still waiting for those backends. In any case there is an unbounded wait. - Backends might join barrier in disjoint groups with some time in between. That means that relying only on the shared dynamic barrier is not enough -- it will only synchronize resize procedure withing those groups. That's why we wait first for all participants of ProcSignal mechanism who received the message. Here is how it looks like after raising shared_buffers from 128 MB to 512 MB and calling pg_reload_conf(): -- 128 MB 7f90cde00000-7f90d4fa6000 /dev/zero (deleted) 7f90d4fa6000-7f914de00000 7f914de00000-7f915cfa8000 /dev/zero (deleted) ^ buffers mapping, ~241 MB 7f915cfa8000-7f944de00000 7f944de00000-7f94550a8000 /dev/zero (deleted) 7f94550a8000-7f94cde00000 7f94cde00000-7f94d4fe8000 /dev/zero (deleted) 7f94d4fe8000-7f954de00000 7f954de00000-7f9554ff6000 /dev/zero (deleted) 7f9554ff6000-7f958de00000 7f958de00000-7f959508a000 /dev/zero (deleted) 7f959508a000-7f95cde00000 -- 512 MB 7f90cde00000-7f90d5126000 /dev/zero (deleted) 7f90d5126000-7f914de00000 7f914de00000-7f9175128000 /dev/zero (deleted) ^ buffers mapping, ~627 MB 7f9175128000-7f944de00000 7f944de00000-7f9455528000 /dev/zero (deleted) 7f9455528000-7f94cde00000 7f94cde00000-7f94d5228000 /dev/zero (deleted) 7f94d5228000-7f954de00000 7f954de00000-7f9555266000 /dev/zero (deleted) 7f9555266000-7f958de00000 7f958de00000-7f95954aa000 /dev/zero (deleted) 7f95954aa000-7f95cde00000 The implementation supports only increasing of shared_buffers. For decreasing the value a similar procedure is needed. But the buffer blocks with data have to be drained first, so that the actual data set fits into the new smaller space. From experiment it turns out that shared mappings have to be extended separately for each process that uses them. Another rough edge is that a backend blocked on ReadCommand will not apply shared_buffers change until it receives something. Note, that mremap is Linux specific, thus the implementation not very portable. Authors: Dmitrii Dolgov, Ashutosh Bapat --- src/backend/port/sysv_shmem.c | 413 ++++++++++++++++++ src/backend/postmaster/postmaster.c | 18 + src/backend/storage/buffer/buf_init.c | 75 ++-- src/backend/storage/ipc/ipci.c | 18 +- src/backend/storage/ipc/procsignal.c | 46 ++ src/backend/storage/ipc/shmem.c | 23 +- src/backend/tcop/postgres.c | 10 + .../utils/activity/wait_event_names.txt | 3 + src/backend/utils/misc/guc_tables.c | 4 +- src/include/miscadmin.h | 1 + src/include/storage/bufmgr.h | 2 +- src/include/storage/ipc.h | 3 + src/include/storage/lwlocklist.h | 1 + src/include/storage/pg_shmem.h | 26 ++ src/include/storage/pmsignal.h | 1 + src/include/storage/procsignal.h | 1 + src/tools/pgindent/typedefs.list | 1 + 17 files changed, 603 insertions(+), 43 deletions(-) diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c index f46d9d5d9cd..a3437973784 100644 --- a/src/backend/port/sysv_shmem.c +++ b/src/backend/port/sysv_shmem.c @@ -30,13 +30,19 @@ #include "miscadmin.h" #include "port/pg_bitutils.h" #include "portability/mem.h" +#include "storage/bufmgr.h" #include "storage/dsm.h" #include "storage/fd.h" #include "storage/ipc.h" +#include "storage/lwlock.h" #include "storage/pg_shmem.h" +#include "storage/pmsignal.h" +#include "storage/procsignal.h" +#include "storage/shmem.h" #include "utils/guc.h" #include "utils/guc_hooks.h" #include "utils/pidfile.h" +#include "utils/wait_event.h" /* @@ -105,6 +111,13 @@ typedef struct AnonymousMapping static AnonymousMapping Mappings[ANON_MAPPINGS]; +/* Flag telling postmaster that resize is needed */ +volatile bool pending_pm_shmem_resize = false; + +/* Keeps track of the previous NBuffers value */ +static int NBuffersOld = -1; +static int NBuffersPending = -1; + /* Keeps track of used mapping segments */ static int next_free_segment = 0; @@ -176,6 +189,49 @@ static Size reserved_offset = 0; */ static bool huge_pages_on = false; +/* + * Flag telling that we have prepared the memory layout to be resizable. If + * false after all shared memory segments creation, it means we failed to setup + * needed layout and falled back to the regular non-resizable approach. + */ +static bool shmem_resizable = false; + +/* + * Currently broadcasted value of NBuffers in shared memory. + * + * Most of the time this value is going to be equal to NBuffers. But if + * postmaster is resizing shared memory and a new backend was created + * at the same time, there is a possibility for the new backend to inherit the + * old NBuffers value, but miss the resize signal if ProcSignal infrastructure + * was not initialized yet. Consider this situation: + * + * Postmaster ------> New Backend + * | | + * | Launch + * | | + * | Inherit NBuffers + * | | + * Resize NBuffers | + * | | + * Emit Barrier | + * | Init ProcSignal + * | | + * Finish resize | + * | | + * New NBuffers Old NBuffers + * + * In this case the backend is not yet ready to receive a signal from + * EmitProcSignalBarrier, and will be ignored. The same happens if ProcSignal + * is initialized even later, after the resizing was finished. + * + * To address resulting inconsistency, postmaster broadcasts the current + * NBuffers value via shared memory. Every new backend has to verify this value + * before it will access the buffer pool: if it differs from its own value, + * this indicates a shared memory resize has happened and the backend has to + * first synchronize with rest of the pack. + */ +ShmemControl *ShmemCtrl = NULL; + static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size); static void IpcMemoryDetach(int status, Datum shmaddr); static void IpcMemoryDelete(int status, Datum shmId); @@ -769,6 +825,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base) { Size total_reserved = (Size) MaxAvailableMemory * BLCKSZ; + shmem_resizable = true; reserved_offset += total_reserved * SHMEM_RESIZE_RATIO[next_free_segment]; } @@ -964,6 +1021,315 @@ AnonymousShmemDetach(int status, Datum arg) } } +/* + * Resize all shared memory segments based on the current NBuffers value, which + * is is applied from NBuffersPending. The actual segment resizing is done via + * mremap, which will fail if is not sufficient space to expand the mapping. + * When finished, based on the new and old values initialize new buffer blocks + * if any. + * + * If reinitializing took place, as the last step this function does buffers + * reinitialization as well and broadcasts the new value of NSharedBuffers. All + * of that needs to be done only by one backend, the first one that managed to + * grab the ShmemResizeLock. + */ +bool +AnonymousShmemResize(void) +{ + int numSemas; + bool reinit = false; + void *ptr = MAP_FAILED; + NBuffers = NBuffersPending; + + elog(DEBUG1, "Resize shmem from %d to %d", NBuffersOld, NBuffers); + + /* + * XXX: Where to reset the flag is still an open question. E.g. do we + * consider a no-op when NBuffers is equal to NBuffersOld a genuine resize + * and reset the flag? + */ + pending_pm_shmem_resize = false; + + /* + * XXX: Currently only increasing of shared_buffers is supported. For + * decreasing something similar has to be done, but buffer blocks with + * data have to be drained first. + */ + if(NBuffersOld > NBuffers) + return false; + + for(int i = 0; i < next_free_segment; i++) + { + /* Note that CalculateShmemSize indirectly depends on NBuffers */ + Size new_size = CalculateShmemSize(&numSemas, i); + AnonymousMapping *m = &Mappings[i]; + + if (m->shmem == NULL) + continue; + + if (m->shmem_size == new_size) + continue; + + /* Clean up some reserved space to resize into */ + if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1) + ereport(FATAL, + (errcode(ERRCODE_SYSTEM_ERROR), + errmsg("could not unmap %zu from reserved shared memory %p: %m", + new_size - m->shmem_size, m->shmem))); + + /* Claim the unused space */ + elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p", + MappingName(m->shmem_segment), m->shmem_size, + new_size, m->shmem); + + ptr = mremap(m->shmem, m->shmem_size, new_size, 0); + if (ptr == MAP_FAILED) + ereport(FATAL, + (errcode(ERRCODE_SYSTEM_ERROR), + errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m", + MappingName(m->shmem_segment), m->shmem, NBuffers, + new_size))); + + reinit = true; + m->shmem_size = new_size; + } + + if (reinit) + { + if(IsUnderPostmaster && + LWLockConditionalAcquire(ShmemResizeLock, LW_EXCLUSIVE)) + { + /* + * If the new NBuffers was already broadcasted, the buffer pool was + * already initialized before. + * + * Since we're not on a hot path, we use lwlocks and do not need to + * involve memory barrier. + */ + if(pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers) != NBuffers) + { + /* + * Allow the first backend that managed to get the lock to + * reinitialize the new portion of buffer pool. Every other + * process will wait on the shared barrier for that to finish, + * since it's a part of the SHMEM_RESIZE_DONE phase. + * + * Note that it's enough when only one backend will do that, + * even the ShmemInitStruct part. The reason is that resized + * shared memory will maintain the same addresses, meaning that + * all the pointers are still valid, and we only need to update + * structures size in the ShmemIndex once -- any other backend + * will pick up this shared structure from the index. + * + * XXX: This is the right place for buffer eviction as well. + */ + BufferManagerShmemInit(NBuffersOld); + + /* If all fine, broadcast the new value */ + pg_atomic_write_u32(&ShmemCtrl->NSharedBuffers, NBuffers); + } + + LWLockRelease(ShmemResizeLock); + } + } + + return true; +} + +/* + * We are asked to resize shared memory. Wait for all ProcSignal participants + * to join the barrier, then do the resize and wait on the barrier until all + * participating finish resizing as well -- otherwise we face danger of + * inconsistency between backends. + * + * XXX: If a backend is blocked on ReadCommand in PostgresMain, it will not + * proceed with AnonymousShmemResize after receiving SIGHUP, until something + * will be sent. + */ +bool +ProcessBarrierShmemResize(Barrier *barrier) +{ + elog(DEBUG1, "Handle a barrier for shmem resizing from %d to %d, %d", + NBuffersOld, NBuffersPending, pending_pm_shmem_resize); + + /* Wait until we have seen the new NBuffers value */ + if (!pending_pm_shmem_resize) + return false; + + /* + * First thing to do after attaching to the barrier is to wait for others. + * We can't simply use BarrierArriveAndWait, because backends might arrive + * here in disjoint groups, e.g. first two backends, pause, then second two + * backends. If the resize is quick enough that can lead to a situation + * when the first group is already finished before the second has appeared, + * and the barrier will only synchonize withing those groups. + */ + if (BarrierAttach(barrier) == SHMEM_RESIZE_REQUESTED) + WaitForProcSignalBarrierReceived( + pg_atomic_read_u64(&ShmemCtrl->Generation)); + + /* + * Now start the procedure, and elect one backend to ping postmaster to do + * the same. + * + * XXX: If we need to be able to abort resizing, this has to be done later, + * after the SHMEM_RESIZE_DONE. + */ + if (BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_START)) + { + Assert(IsUnderPostmaster); + SendPostmasterSignal(PMSIGNAL_SHMEM_RESIZE); + } + + AnonymousShmemResize(); + + /* The second phase means the resize has finished, SHMEM_RESIZE_DONE */ + BarrierArriveAndWait(barrier, WAIT_EVENT_SHMEM_RESIZE_DONE); + + BarrierDetach(barrier); + return true; +} + +/* + * GUC assign hook for shared_buffers. It's recommended for an assign hook to + * be as minimal as possible, thus we just request shared memory resize and + * remember the previous value. + */ +void +assign_shared_buffers(int newval, void *extra, bool *pending) +{ + elog(DEBUG1, "Received SIGHUP for shmem resizing"); + + /* Request shared memory resize only when it was initialized */ + if (next_free_segment != 0) + { + elog(DEBUG1, "Set pending signal"); + pending_pm_shmem_resize = true; + *pending = true; + NBuffersPending = newval; + } + + NBuffersOld = NBuffers; +} + +/* + * Test if we have somehow missed a shmem resize signal and NBuffers value + * differs from NSharedBuffers. If yes, catchup and do resize. + */ +void +AdjustShmemSize(void) +{ + uint32 NSharedBuffers = pg_atomic_read_u32(&ShmemCtrl->NSharedBuffers); + + if (NSharedBuffers != NBuffers) + { + /* + * If the broadcasted shared_buffers is different from the one we see, + * it could be that the backend has missed a resize signal. To avoid + * any inconsistency, adjust the shared mappings, before having a + * chance to access the buffer pool. + */ + ereport(LOG, + (errmsg("shared_buffers has been changed from %d to %d, " + "resize shared memory", + NBuffers, NSharedBuffers))); + NBuffers = NSharedBuffers; + AnonymousShmemResize(); + } +} + +/* + * Start resizing procedure, making sure all existing processes will have + * consistent view of shared memory size. Must be called only in postmaster. + */ +void +CoordinateShmemResize(void) +{ + elog(DEBUG1, "Coordinating shmem resize from %d to %d", + NBuffersOld, NBuffers); + Assert(!IsUnderPostmaster); + + /* + * We use dynamic barrier to help dealing with backends that were spawned + * during the resize. + */ + BarrierInit(&ShmemCtrl->Barrier, 0); + + /* + * If the value did not change, or shared memory segments are not + * initialized yet, skip the resize. + */ + if (NBuffersPending == NBuffersOld || next_free_segment == 0) + { + elog(DEBUG1, "Skip resizing, new %d, old %d, free segment %d", + NBuffers, NBuffersOld, next_free_segment); + return; + } + + /* + * Shared memory resize requires some coordination done by postmaster, + * and consists of three phases: + * + * - Before the resize all existing backends have the same old NBuffers. + * - When resize is in progress, backends are expected to have a + * mixture of old a new values. They're not allowed to touch buffer + * pool during this time frame. + * - After resize has been finished, all existing backends, that can access + * the buffer pool, are expected to have the same new value of NBuffers. + * + * Those phases are ensured by joining the shared barrier associated with + * the procedure. Since resizing takes time, we need to take into account + * that during that time: + * + * - New backends can be spawned. They will check status of the barrier + * early during the bootstrap, and wait until everything is over to work + * with the new NBuffers value. + * + * - Old backends can exit before attempting to resize. Synchronization + * used between backends relies on ProcSignalBarrier and waits for all + * participants received the message at the beginning to gather all + * existing backends. + * + * - Some backends might be blocked and not responsing either before or + * after receiving the message. In the first case such backend still + * have ProcSignalSlot and should be waited for, in the second case + * shared barrier will make sure we still waiting for those backends. In + * any case there is an unbounded wait. + * + * - Backends might join barrier in disjoint groups with some time in + * between. That means that relying only on the shared dynamic barrier is + * not enough -- it will only synchronize resize procedure withing those + * groups. That's why we wait first for all participants of ProcSignal + * mechanism who received the message. + */ + elog(DEBUG1, "Emit a barrier for shmem resizing"); + pg_atomic_init_u64(&ShmemCtrl->Generation, + EmitProcSignalBarrier(PROCSIGNAL_BARRIER_SHMEM_RESIZE)); + + /* To order everything after setting Generation value */ + pg_memory_barrier(); + + /* + * After that postmaster waits for PMSIGNAL_SHMEM_RESIZE as a sign that all + * the rest of the pack has started the procedure and it can resize shared + * memory as well. + * + * Normally we would call WaitForProcSignalBarrier here to wait until every + * backend has reported on the ProcSignalBarrier. But for shared memory + * resize we don't need this, as every participating backend will + * synchronize on the ProcSignal barrier. In fact even if we would like to + * wait here, it wouldn't be possible -- we're in the postmaster, without + * any waiting infrastructure available. + * + * If at some point it will turn out that waiting is essential, we would + * need to consider some alternatives. E.g. it could be a designated + * coordination process, which is not a postmaster. Another option would be + * to introduce a CoordinateShmemResize lock and allow only one process to + * take it (this probably would have to be something different than + * LWLocks, since they block interrupts, and coordination relies on them). + */ +} + /* * PGSharedMemoryCreate * @@ -1271,3 +1637,50 @@ PGSharedMemoryDetach(void) } } } + +void +WaitOnShmemBarrier() +{ + Barrier *barrier = &ShmemCtrl->Barrier; + + /* Nothing to do if resizing is not started */ + if (BarrierPhase(barrier) < SHMEM_RESIZE_START) + return; + + BarrierAttach(barrier); + + /* Otherwise wait through all available phases */ + while (BarrierPhase(barrier) < SHMEM_RESIZE_DONE) + { + ereport(LOG, (errmsg("ProcSignal barrier is in phase %d, waiting", + BarrierPhase(barrier)))); + + BarrierArriveAndWait(barrier, 0); + } + + BarrierDetach(barrier); +} + +void +ShmemControlInit(void) +{ + bool foundShmemCtrl; + + ShmemCtrl = (ShmemControl *) + ShmemInitStruct("Shmem Control", sizeof(ShmemControl), + &foundShmemCtrl); + + if (!foundShmemCtrl) + { + /* + * The barrier is missing here, it will be initialized right before + * starting the resizing process as a convenient way to reset it. + */ + + /* Initialize with the currently known value */ + pg_atomic_init_u32(&ShmemCtrl->NSharedBuffers, NBuffers); + + /* shmem_resizable should be initialized by now */ + ShmemCtrl->Resizable = shmem_resizable; + } +} diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c index 3fe45de5da0..196f233fe0e 100644 --- a/src/backend/postmaster/postmaster.c +++ b/src/backend/postmaster/postmaster.c @@ -425,6 +425,7 @@ static void process_pm_pmsignal(void); static void process_pm_child_exit(void); static void process_pm_reload_request(void); static void process_pm_shutdown_request(void); +static void process_pm_shmem_resize(void); static void dummy_handler(SIGNAL_ARGS); static void CleanupBackend(PMChild *bp, int exitstatus); static void HandleChildCrash(int pid, int exitstatus, const char *procname); @@ -1693,6 +1694,9 @@ ServerLoop(void) if (pending_pm_pmsignal) process_pm_pmsignal(); + if (pending_pm_shmem_resize) + process_pm_shmem_resize(); + if (events[i].events & WL_SOCKET_ACCEPT) { ClientSocket s; @@ -2038,6 +2042,17 @@ process_pm_reload_request(void) } } +static void +process_pm_shmem_resize(void) +{ + /* + * Failure to resize is considered to be fatal and will not be + * retried, which means we can disable pending flag right here. + */ + pending_pm_shmem_resize = false; + CoordinateShmemResize(); +} + /* * pg_ctl uses SIGTERM, SIGINT and SIGQUIT to request different types of * shutdown. @@ -3851,6 +3866,9 @@ process_pm_pmsignal(void) request_state_update = true; } + if (CheckPostmasterSignal(PMSIGNAL_SHMEM_RESIZE)) + AnonymousShmemResize(); + /* * Try to advance postmaster's state machine, if a child requests it. */ diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c index bd68b69ee98..ac844b114bd 100644 --- a/src/backend/storage/buffer/buf_init.c +++ b/src/backend/storage/buffer/buf_init.c @@ -17,6 +17,7 @@ #include "storage/aio.h" #include "storage/buf_internals.h" #include "storage/bufmgr.h" +#include "storage/pg_shmem.h" BufferDescPadded *BufferDescriptors; char *BufferBlocks; @@ -24,7 +25,6 @@ ConditionVariableMinimallyPadded *BufferIOCVArray; WritebackContext BackendWritebackContext; CkptSortItem *CkptBufferIds; - /* * Data Structures: * buffers live in a freelist and a lookup data structure. @@ -62,18 +62,28 @@ CkptSortItem *CkptBufferIds; * Initialize shared buffer pool * * This is called once during shared-memory initialization (either in the - * postmaster, or in a standalone backend). Size of data structures initialized - * here depends on NBuffers, and to be able to change NBuffers without a - * restart we store each structure into a separate shared memory segment, which - * could be resized on demand. + * postmaster, or in a standalone backend) or during shared-memory resize. Size + * of data structures initialized here depends on NBuffers, and to be able to + * change NBuffers without a restart we store each structure into a separate + * shared memory segment, which could be resized on demand. + * + * FirstBufferToInit tells where to start initializing buffers. For + * initialization it always will be zero, but when resizing shared-memory it + * indicates the number of already initialized buffers. + * + * No locks are taking in this function, it is the caller responsibility to + * make sure only one backend can work with new buffers. */ void -BufferManagerShmemInit(void) +BufferManagerShmemInit(int FirstBufferToInit) { bool foundBufs, foundDescs, foundIOCV, foundBufCkpt; + int i; + elog(DEBUG1, "BufferManagerShmemInit from %d to %d", + FirstBufferToInit, NBuffers); /* Align descriptors to a cacheline boundary. */ BufferDescriptors = (BufferDescPadded *) @@ -110,43 +120,44 @@ BufferManagerShmemInit(void) { /* should find all of these, or none of them */ Assert(foundDescs && foundBufs && foundIOCV && foundBufCkpt); - /* note: this path is only taken in EXEC_BACKEND case */ - } - else - { - int i; - /* - * Initialize all the buffer headers. + * note: this path is only taken in EXEC_BACKEND case when initializing + * shared memory, or in all cases when resizing shared memory. */ - for (i = 0; i < NBuffers; i++) - { - BufferDesc *buf = GetBufferDescriptor(i); + } - ClearBufferTag(&buf->tag); +#ifndef EXEC_BACKEND + /* + * Initialize all the buffer headers. + */ + for (i = FirstBufferToInit; i < NBuffers; i++) + { + BufferDesc *buf = GetBufferDescriptor(i); - pg_atomic_init_u32(&buf->state, 0); - buf->wait_backend_pgprocno = INVALID_PROC_NUMBER; + ClearBufferTag(&buf->tag); - buf->buf_id = i; + pg_atomic_init_u32(&buf->state, 0); + buf->wait_backend_pgprocno = INVALID_PROC_NUMBER; - pgaio_wref_clear(&buf->io_wref); + buf->buf_id = i; - /* - * Initially link all the buffers together as unused. Subsequent - * management of this list is done by freelist.c. - */ - buf->freeNext = i + 1; + pgaio_wref_clear(&buf->io_wref); - LWLockInitialize(BufferDescriptorGetContentLock(buf), - LWTRANCHE_BUFFER_CONTENT); + /* + * Initially link all the buffers together as unused. Subsequent + * management of this list is done by freelist.c. + */ + buf->freeNext = i + 1; - ConditionVariableInit(BufferDescriptorGetIOCV(buf)); - } + LWLockInitialize(BufferDescriptorGetContentLock(buf), + LWTRANCHE_BUFFER_CONTENT); - /* Correct last entry of linked list */ - GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST; + ConditionVariableInit(BufferDescriptorGetIOCV(buf)); } +#endif + + /* Correct last entry of linked list */ + GetBufferDescriptor(NBuffers - 1)->freeNext = FREENEXT_END_OF_LIST; /* Init other shared buffer-management stuff */ StrategyInitialize(!foundDescs); diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c index 9d00b80b4f8..abeb91e24fd 100644 --- a/src/backend/storage/ipc/ipci.c +++ b/src/backend/storage/ipc/ipci.c @@ -84,6 +84,9 @@ RequestAddinShmemSpace(Size size) * * If num_semaphores is not NULL, it will be set to the number of semaphores * required. + * + * XXX: Calculation for non main shared memory segments are incorrect, it + * includes more than needed for buffers only. */ Size CalculateShmemSize(int *num_semaphores, int shmem_segment) @@ -151,6 +154,14 @@ CalculateShmemSize(int *num_semaphores, int shmem_segment) size = add_size(size, SlotSyncShmemSize()); size = add_size(size, AioShmemSize()); + /* + * XXX: For some reason slightly more memory is needed for larger + * shared_buffers, but this size is enough for any large value I've tested + * with. Is it a mistake in how slots are split, or there was a hidden + * inconsistency in shmem calculation? + */ + size = add_size(size, 1024 * 1024 * 100); + /* include additional requested shmem from preload libraries */ size = add_size(size, total_addin_request); @@ -298,7 +309,7 @@ CreateOrAttachShmemStructs(void) CommitTsShmemInit(); SUBTRANSShmemInit(); MultiXactShmemInit(); - BufferManagerShmemInit(); + BufferManagerShmemInit(0); /* * Set up lock manager @@ -310,6 +321,11 @@ CreateOrAttachShmemStructs(void) */ PredicateLockShmemInit(); + /* + * Set up shared memory resize manager + */ + ShmemControlInit(); + /* * Set up process table */ diff --git a/src/backend/storage/ipc/procsignal.c b/src/backend/storage/ipc/procsignal.c index 8e313ad9bf8..35c42f260a8 100644 --- a/src/backend/storage/ipc/procsignal.c +++ b/src/backend/storage/ipc/procsignal.c @@ -27,6 +27,7 @@ #include "storage/condition_variable.h" #include "storage/ipc.h" #include "storage/latch.h" +#include "storage/pg_shmem.h" #include "storage/shmem.h" #include "storage/sinval.h" #include "storage/smgr.h" @@ -112,6 +113,10 @@ static bool CheckProcSignal(ProcSignalReason reason); static void CleanupProcSignalState(int status, Datum arg); static void ResetProcSignalBarrierBits(uint32 flags); +#ifdef DEBUG_SHMEM_RESIZE +bool delay_proc_signal_init = false; +#endif + /* * ProcSignalShmemSize * Compute space needed for ProcSignal's shared memory @@ -175,6 +180,43 @@ ProcSignalInit(char *cancel_key, int cancel_key_len) uint32 old_pss_pid; Assert(cancel_key_len >= 0 && cancel_key_len <= MAX_CANCEL_KEY_LENGTH); + +#ifdef DEBUG_SHMEM_RESIZE + /* + * Introduced for debugging purposes. You can change the variable at + * runtime using gdb, then start new backends with delayed ProcSignal + * initialization. Simple pg_usleep wont work here due to SIGHUP interrupt + * needed for testing. Taken from pg_sleep; + */ + if (delay_proc_signal_init) + { +#define GetNowFloat() ((float8) GetCurrentTimestamp() / 1000000.0) + float8 endtime = GetNowFloat() + 5; + + for (;;) + { + float8 delay; + long delay_ms; + + CHECK_FOR_INTERRUPTS(); + + delay = endtime - GetNowFloat(); + if (delay >= 600.0) + delay_ms = 600000; + else if (delay > 0.0) + delay_ms = (long) (delay * 1000.0); + else + break; + + (void) WaitLatch(MyLatch, + WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH, + delay_ms, + WAIT_EVENT_PG_SLEEP); + ResetLatch(MyLatch); + } + } +#endif + if (MyProcNumber < 0) elog(ERROR, "MyProcNumber not set"); if (MyProcNumber >= NumProcSignalSlots) @@ -614,6 +656,10 @@ ProcessProcSignalBarrier(void) case PROCSIGNAL_BARRIER_SMGRRELEASE: processed = ProcessBarrierSmgrRelease(); break; + case PROCSIGNAL_BARRIER_SHMEM_RESIZE: + processed = ProcessBarrierShmemResize( + &ShmemCtrl->Barrier); + break; } /* diff --git a/src/backend/storage/ipc/shmem.c b/src/backend/storage/ipc/shmem.c index 389abc82519..0fd421f004e 100644 --- a/src/backend/storage/ipc/shmem.c +++ b/src/backend/storage/ipc/shmem.c @@ -493,17 +493,26 @@ ShmemInitStructInSegment(const char *name, Size size, bool *foundPtr, { /* * Structure is in the shmem index so someone else has allocated it - * already. The size better be the same as the size we are trying to - * initialize to, or there is a name conflict (or worse). + * already. Verify the structure's size: + * - If it's the same, we've found the expected structure. + * - If it's different, we're resizing the expected structure. + * + * XXX: There is an implicit assumption this can only happen in + * "resizable" segments, where only one shared structure is allowed. + * This has to be implemented more cleanly. */ if (result->size != size) { - LWLockRelease(ShmemIndexLock); - ereport(ERROR, - (errmsg("ShmemIndex entry size is wrong for data structure" - " \"%s\": expected %zu, actual %zu", - name, size, result->size))); + Size delta = size - result->size; + + result->size = size; + + /* Reflect size change in the shared segment */ + SpinLockAcquire(Segments[shmem_segment].ShmemLock); + Segments[shmem_segment].ShmemSegHdr->freeoffset += delta; + SpinLockRelease(Segments[shmem_segment].ShmemLock); } + structPtr = result->location; } else diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c index b1fba850f02..58f1a05fd2a 100644 --- a/src/backend/tcop/postgres.c +++ b/src/backend/tcop/postgres.c @@ -62,6 +62,7 @@ #include "rewrite/rewriteHandler.h" #include "storage/bufmgr.h" #include "storage/ipc.h" +#include "storage/pg_shmem.h" #include "storage/pmsignal.h" #include "storage/proc.h" #include "storage/procsignal.h" @@ -4311,6 +4312,15 @@ PostgresMain(const char *dbname, const char *username) */ BeginReportingGUCOptions(); + /* Verify the shared barrier, if it's still active: join and wait. */ + WaitOnShmemBarrier(); + + /* + * After waiting on the barrier above we guaranteed to have NSharedBuffers + * broadcasted, so we can use it in the function below. + */ + AdjustShmemSize(); + /* * Also set up handler to log session end; we have to wait till now to be * sure Log_disconnections has its final value. diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 8bce14c38fd..e0ba8384fdd 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -155,6 +155,8 @@ REPLICATION_ORIGIN_DROP "Waiting for a replication origin to become inactive so REPLICATION_SLOT_DROP "Waiting for a replication slot to become inactive so it can be dropped." RESTORE_COMMAND "Waiting for to complete." SAFE_SNAPSHOT "Waiting to obtain a valid snapshot for a READ ONLY DEFERRABLE transaction." +SHMEM_RESIZE_START "Waiting for other backends to start resizing shared memory." +SHMEM_RESIZE_DONE "Waiting for other backends to finish resizing shared memory." SYNC_REP "Waiting for confirmation from a remote server during synchronous replication." WAL_BUFFER_INIT "Waiting on WAL buffer to be initialized." WAL_RECEIVER_EXIT "Waiting for the WAL receiver to exit." @@ -351,6 +353,7 @@ DSMRegistry "Waiting to read or update the dynamic shared memory registry." InjectionPoint "Waiting to read or update information related to injection points." SerialControl "Waiting to read or update shared pg_serial state." AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue." +ShmemResize "Waiting to resize shared memory." # # END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE) diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c index dede37f7905..1e70853ccdb 100644 --- a/src/backend/utils/misc/guc_tables.c +++ b/src/backend/utils/misc/guc_tables.c @@ -2354,14 +2354,14 @@ struct config_int ConfigureNamesInt[] = * checking for overflow, so we mustn't allow more than INT_MAX / 2. */ { - {"shared_buffers", PGC_POSTMASTER, RESOURCES_MEM, + {"shared_buffers", PGC_SIGHUP, RESOURCES_MEM, gettext_noop("Sets the number of shared memory buffers used by the server."), NULL, GUC_UNIT_BLOCKS }, &NBuffers, 16384, 16, INT_MAX / 2, - NULL, NULL, NULL + NULL, assign_shared_buffers, NULL }, { diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h index 0d8528b2875..405d0a7e65d 100644 --- a/src/include/miscadmin.h +++ b/src/include/miscadmin.h @@ -173,6 +173,7 @@ extern PGDLLIMPORT char *DataDir; extern PGDLLIMPORT int data_directory_mode; extern PGDLLIMPORT int NBuffers; +extern PGDLLIMPORT int MaxAvailableMemory; extern PGDLLIMPORT int MaxBackends; extern PGDLLIMPORT int MaxConnections; extern PGDLLIMPORT int max_worker_processes; diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h index 1977001e533..52633dd7537 100644 --- a/src/include/storage/bufmgr.h +++ b/src/include/storage/bufmgr.h @@ -307,7 +307,7 @@ extern void LimitAdditionalLocalPins(uint32 *additional_pins); extern bool EvictUnpinnedBuffer(Buffer buf); /* in buf_init.c */ -extern void BufferManagerShmemInit(void); +extern void BufferManagerShmemInit(int); extern Size BufferManagerShmemSize(int); /* in localbuf.c */ diff --git a/src/include/storage/ipc.h b/src/include/storage/ipc.h index 6ebda479ced..bb7ae4d33b3 100644 --- a/src/include/storage/ipc.h +++ b/src/include/storage/ipc.h @@ -64,6 +64,7 @@ typedef void (*shmem_startup_hook_type) (void); /* ipc.c */ extern PGDLLIMPORT bool proc_exit_inprogress; extern PGDLLIMPORT bool shmem_exit_inprogress; +extern PGDLLIMPORT volatile bool pending_pm_shmem_resize; pg_noreturn extern void proc_exit(int code); extern void shmem_exit(int code); @@ -83,5 +84,7 @@ extern void CreateSharedMemoryAndSemaphores(void); extern void AttachSharedMemoryStructs(void); #endif extern void InitializeShmemGUCs(void); +extern void CoordinateShmemResize(void); +extern bool AnonymousShmemResize(void); #endif /* IPC_H */ diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index a9681738146..558da6fdd55 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -84,3 +84,4 @@ PG_LWLOCK(50, DSMRegistry) PG_LWLOCK(51, InjectionPoint) PG_LWLOCK(52, SerialControl) PG_LWLOCK(53, AioWorkerSubmissionQueue) +PG_LWLOCK(54, ShmemResize) diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h index c5009a1cd73..2e47b222cbb 100644 --- a/src/include/storage/pg_shmem.h +++ b/src/include/storage/pg_shmem.h @@ -24,6 +24,7 @@ #ifndef PG_SHMEM_H #define PG_SHMEM_H +#include "storage/barrier.h" #include "storage/dsm_impl.h" #include "storage/spin.h" @@ -56,6 +57,25 @@ typedef struct ShmemSegment extern PGDLLIMPORT ShmemSegment Segments[ANON_MAPPINGS]; +/* + * ShmemControl is shared between backends and helps to coordinate shared + * memory resize. + */ +typedef struct +{ + pg_atomic_uint32 NSharedBuffers; + Barrier Barrier; + pg_atomic_uint64 Generation; + bool Resizable; +} ShmemControl; + +extern PGDLLIMPORT ShmemControl *ShmemCtrl; + +/* The phases for shared memory resizing, used by for ProcSignal barrier. */ +#define SHMEM_RESIZE_REQUESTED 0 +#define SHMEM_RESIZE_START 1 +#define SHMEM_RESIZE_DONE 2 + /* GUC variables */ extern PGDLLIMPORT int shared_memory_type; extern PGDLLIMPORT int huge_pages; @@ -107,6 +127,12 @@ extern void PGSharedMemoryDetach(void); extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags); void *ReserveAnonymousMemory(Size reserve_size); +bool ProcessBarrierShmemResize(Barrier *barrier); +void assign_shared_buffers(int newval, void *extra, bool *pending); +void AdjustShmemSize(void); +extern void WaitOnShmemBarrier(void); +extern void ShmemControlInit(void); + /* * To be able to dynamically resize largest parts of the data stored in shared * memory, we split it into multiple shared memory mappings segments. Each diff --git a/src/include/storage/pmsignal.h b/src/include/storage/pmsignal.h index 67fa9ac06e1..27bc6a81191 100644 --- a/src/include/storage/pmsignal.h +++ b/src/include/storage/pmsignal.h @@ -42,6 +42,7 @@ typedef enum PMSIGNAL_START_WALRECEIVER, /* start a walreceiver */ PMSIGNAL_ADVANCE_STATE_MACHINE, /* advance postmaster's state machine */ PMSIGNAL_XLOG_IS_SHUTDOWN, /* ShutdownXLOG() completed */ + PMSIGNAL_SHMEM_RESIZE, /* resize shared memory */ } PMSignalReason; #define NUM_PMSIGNALS (PMSIGNAL_XLOG_IS_SHUTDOWN+1) diff --git a/src/include/storage/procsignal.h b/src/include/storage/procsignal.h index defd8b66a19..522b8de1e02 100644 --- a/src/include/storage/procsignal.h +++ b/src/include/storage/procsignal.h @@ -54,6 +54,7 @@ typedef enum typedef enum { PROCSIGNAL_BARRIER_SMGRRELEASE, /* ask smgr to close files */ + PROCSIGNAL_BARRIER_SHMEM_RESIZE, /* ask backends to resize shared memory */ } ProcSignalBarrierType; /* diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 1a30437ad96..6755b302858 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2738,6 +2738,7 @@ ShellTypeInfo ShippableCacheEntry ShippableCacheKey ShmemIndexEnt +ShmemControl ShutdownForeignScan_function ShutdownInformation ShutdownMode -- 2.45.1 --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="v4-0007-Use-anonymous-files-to-back-shared-memory-segment.patch" From 0e3c671082743f2826a7e8a96a19a071f5c8aeb3 Mon Sep 17 00:00:00 2001 From: Dmitrii Dolgov <9erthalion6@gmail.com> Date: Sat, 15 Mar 2025 16:39:45 +0100 Subject: [PATCH v4 7/8] Use anonymous files to back shared memory segments Allow to use anonymous files for shared memory, instead of plain anonymous memory. Such an anonymous file is created via memfd_create, it lives in memory, behaves like a regular file and semantically equivalent to an anonymous memory allocated via mmap with MAP_ANONYMOUS. Advantages of using anon files are following: * We've got a file descriptor, which could be used for regular file operations (modification, truncation, you name it). * The file could be given a name, which improves readability when it comes to process maps. Here is how it looks like 7f90cde00000-7f90d5126000 rw-s 00000000 00:01 5463 /memfd:main (deleted) 7f90d5126000-7f914de00000 ---p 00000000 00:00 0 7f914de00000-7f9175128000 rw-s 00000000 00:01 5466 /memfd:buffers (deleted) 7f9175128000-7f944de00000 ---p 00000000 00:00 0 7f944de00000-7f9455528000 rw-s 00000000 00:01 5469 /memfd:descriptors (deleted) 7f9455528000-7f94cde00000 ---p 00000000 00:00 0 7f94cde00000-7f94d5228000 rw-s 00000000 00:01 5472 /memfd:iocv (deleted) 7f94d5228000-7f954de00000 ---p 00000000 00:00 0 7f954de00000-7f9555266000 rw-s 00000000 00:01 5475 /memfd:checkpoint (deleted) 7f9555266000-7f958de00000 ---p 00000000 00:00 0 7f958de00000-7f95954aa000 rw-s 00000000 00:01 5478 /memfd:strategy (deleted) 7f95954aa000-7f95cde00000 ---p 00000000 00:00 0 * By default, Linux will not add file-backed shared mappings into a core dump, making it more convenient to work with them in PostgreSQL: no more huge dumps to process. The downside is that memfd_create is Linux specific. --- src/backend/port/sysv_shmem.c | 73 +++++++++++++++++++++++++++++----- src/backend/port/win32_shmem.c | 2 +- src/backend/storage/ipc/ipci.c | 2 +- src/include/portability/mem.h | 2 +- src/include/storage/pg_shmem.h | 3 +- 5 files changed, 68 insertions(+), 14 deletions(-) diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c index a3437973784..87000a24eea 100644 --- a/src/backend/port/sysv_shmem.c +++ b/src/backend/port/sysv_shmem.c @@ -107,6 +107,7 @@ typedef struct AnonymousMapping Pointer shmem; /* Pointer to the start of the mapped memory */ Pointer seg_addr; /* SysV shared memory for the header */ unsigned long seg_id; /* IPC key */ + int segment_fd; /* fd for the backing anon file */ } AnonymousMapping; static AnonymousMapping Mappings[ANON_MAPPINGS]; @@ -127,7 +128,7 @@ static int next_free_segment = 0; * 00400000-00490000 /path/bin/postgres * ... * 012d9000-0133e000 [heap] - * 7f443a800000-7f470a800000 /dev/zero (deleted) + * 7f443a800000-7f470a800000 /memfd:main (deleted) * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2 * ... @@ -150,9 +151,9 @@ static int next_free_segment = 0; * The result would look like this: * * 012d9000-0133e000 [heap] - * 7f4426f54000-7f442e010000 /dev/zero (deleted) + * 7f4426f54000-7f442e010000 /memfd:main (deleted) * 7f442e010000-7f443a800000 # reserved empty space - * 7f443a800000-7f444196c000 /dev/zero (deleted) + * 7f443a800000-7f444196c000 /memfd:buffers (deleted) * 7f444196c000-7f470a800000 # reserved empty space * 7f470a800000-7f471831d000 /usr/lib/locale/locale-archive * 7f4718400000-7f4718401000 /usr/lib64/libicudata.so.74.2 @@ -643,13 +644,14 @@ PGSharedMemoryAttach(IpcMemoryId shmId, * *hugepagesize and *mmap_flags are set to 0. */ void -GetHugePageSize(Size *hugepagesize, int *mmap_flags) +GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags) { #ifdef MAP_HUGETLB Size default_hugepagesize = 0; Size hugepagesize_local = 0; int mmap_flags_local = 0; + int memfd_flags_local = 0; /* * System-dependent code to find out the default huge page size. @@ -708,6 +710,7 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags) } mmap_flags_local = MAP_HUGETLB; + memfd_flags_local = MFD_HUGETLB; /* * On recent enough Linux, also include the explicit page size, if @@ -718,7 +721,16 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags) { int shift = pg_ceil_log2_64(hugepagesize_local); - mmap_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT; + memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT; + } +#endif + +#if defined(MFD_HUGE_MASK) && defined(MFD_HUGE_SHIFT) + if (hugepagesize_local != default_hugepagesize) + { + int shift = pg_ceil_log2_64(hugepagesize_local); + + memfd_flags_local |= (shift & MAP_HUGE_MASK) << MAP_HUGE_SHIFT; } #endif @@ -727,6 +739,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags) *mmap_flags = mmap_flags_local; if (hugepagesize) *hugepagesize = hugepagesize_local; + if (memfd_flags) + *memfd_flags = memfd_flags_local; #else @@ -734,6 +748,8 @@ GetHugePageSize(Size *hugepagesize, int *mmap_flags) *hugepagesize = 0; if (mmap_flags) *mmap_flags = 0; + if (memfd_flags) + *memfd_flags = 0; #endif /* MAP_HUGETLB */ } @@ -771,7 +787,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base) Size allocsize = mapping->shmem_size; void *ptr = MAP_FAILED; int mmap_errno = 0; - int mmap_flags = PG_MMAP_FLAGS; + int mmap_flags = PG_MMAP_FLAGS, memfd_flags = 0; #ifndef MAP_HUGETLB /* ReserveAnonymousMemory should have dealt with this case */ @@ -785,7 +801,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base) Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY); /* Round up the request size to a suitable large value */ - GetHugePageSize(&hugepagesize, &mmap_flags); + GetHugePageSize(&hugepagesize, &mmap_flags, &memfd_flags); if (allocsize % hugepagesize != 0) allocsize += hugepagesize - (allocsize % hugepagesize); @@ -794,6 +810,29 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base) } #endif + /* + * Prepare an anonymous file backing the segment. Its size will be + * specified later via ftruncate. + * + * The file behaves like a regular file, but lives in memory. Once all + * references to the file are dropped, it is automatically released. + * Anonymous memory is used for all backing pages of the file, thus it has + * the same semantics as anonymous memory allocations using mmap with the + * MAP_ANONYMOUS flag. + */ + mapping->segment_fd = memfd_create(MappingName(mapping->shmem_segment), + memfd_flags); + + /* + * Specify the segment file size using allocsize, which contains + * potentially modified size. + */ + if(ftruncate(mapping->segment_fd, allocsize) == -1) + ereport(FATAL, + (errcode(ERRCODE_SYSTEM_ERROR), + errmsg("could not truncase anonymous file for \"%s\": %m", + MappingName(mapping->shmem_segment)))); + elog(DEBUG1, "segment[%s]: mmap(%zu) at address %p", MappingName(mapping->shmem_segment), allocsize, base + reserved_offset); @@ -807,7 +846,7 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base) * a restart. */ ptr = mmap(base + reserved_offset, allocsize, PROT_READ | PROT_WRITE, - mmap_flags | MAP_FIXED, -1, 0); + mmap_flags | MAP_FIXED, mapping->segment_fd, 0); mmap_errno = errno; if (ptr == MAP_FAILED) @@ -817,8 +856,15 @@ CreateAnonymousSegment(AnonymousMapping *mapping, Pointer base) "fallback to the non-resizable allocation", MappingName(mapping->shmem_segment), allocsize, base + reserved_offset); + /* Specify the segment file size using allocsize. */ + if(ftruncate(mapping->segment_fd, allocsize) == -1) + ereport(FATAL, + (errcode(ERRCODE_SYSTEM_ERROR), + errmsg("could not truncase anonymous file for \"%s\": %m", + MappingName(mapping->shmem_segment)))); + ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE, - PG_MMAP_FLAGS, -1, 0); + PG_MMAP_FLAGS, mapping->segment_fd, 0); mmap_errno = errno; } else @@ -889,7 +935,7 @@ ReserveAnonymousMemory(Size reserve_size) Size hugepagesize, total_size = 0; int mmap_flags; - GetHugePageSize(&hugepagesize, &mmap_flags); + GetHugePageSize(&hugepagesize, &mmap_flags, NULL); /* * Figure out how much memory is needed for all segments, keeping in @@ -1070,6 +1116,13 @@ AnonymousShmemResize(void) if (m->shmem_size == new_size) continue; + /* Resize the backing anon file. */ + if(ftruncate(m->segment_fd, new_size) == -1) + ereport(FATAL, + (errcode(ERRCODE_SYSTEM_ERROR), + errmsg("could not truncase anonymous file for \"%s\": %m", + MappingName(m->shmem_segment)))); + /* Clean up some reserved space to resize into */ if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1) ereport(FATAL, diff --git a/src/backend/port/win32_shmem.c b/src/backend/port/win32_shmem.c index ce719f1b412..ba972106de1 100644 --- a/src/backend/port/win32_shmem.c +++ b/src/backend/port/win32_shmem.c @@ -627,7 +627,7 @@ pgwin32_ReserveSharedMemoryRegion(HANDLE hChild) * use GetLargePageMinimum() instead. */ void -GetHugePageSize(Size *hugepagesize, int *mmap_flags) +GetHugePageSize(Size *hugepagesize, int *mmap_flags, int *memfd_flags) { if (hugepagesize) *hugepagesize = 0; diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c index abeb91e24fd..dc2b4becf4a 100644 --- a/src/backend/storage/ipc/ipci.c +++ b/src/backend/storage/ipc/ipci.c @@ -396,7 +396,7 @@ InitializeShmemGUCs(void) /* * Calculate the number of huge pages required. */ - GetHugePageSize(&hp_size, NULL); + GetHugePageSize(&hp_size, NULL, NULL); if (hp_size != 0) { Size hp_required; diff --git a/src/include/portability/mem.h b/src/include/portability/mem.h index ef9800732d9..40588ff6968 100644 --- a/src/include/portability/mem.h +++ b/src/include/portability/mem.h @@ -38,7 +38,7 @@ #define MAP_NOSYNC 0 #endif -#define PG_MMAP_FLAGS (MAP_SHARED|MAP_ANONYMOUS|MAP_HASSEMAPHORE) +#define PG_MMAP_FLAGS (MAP_SHARED|MAP_HASSEMAPHORE) /* Some really old systems don't define MAP_FAILED. */ #ifndef MAP_FAILED diff --git a/src/include/storage/pg_shmem.h b/src/include/storage/pg_shmem.h index 2e47b222cbb..b9573520d9a 100644 --- a/src/include/storage/pg_shmem.h +++ b/src/include/storage/pg_shmem.h @@ -124,7 +124,8 @@ extern PGShmemHeader *PGSharedMemoryCreate(Size size, PGShmemHeader **shim, Pointer base); extern bool PGSharedMemoryIsInUse(unsigned long id1, unsigned long id2); extern void PGSharedMemoryDetach(void); -extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags); +extern void GetHugePageSize(Size *hugepagesize, int *mmap_flags, + int *memfd_flags); void *ReserveAnonymousMemory(Size reserve_size); bool ProcessBarrierShmemResize(Barrier *barrier); -- 2.45.1 --vninua6xybvzgrci Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="v4-0008-Support-resize-for-hugetlb.patch" From 08476af71724fcb3035fc907dc98a6ff351fe58e Mon Sep 17 00:00:00 2001 From: Dmitrii Dolgov <9erthalion6@gmail.com> Date: Sat, 5 Apr 2025 19:51:33 +0200 Subject: [PATCH v4 8/8] Support resize for hugetlb Linux kernel has a set of limitations on remapping hugetlb segments: it can't increase size of such segment [1], and shrinking it will not release the memory back. In fact support for hugetlb mremap was implemented no so long time ago [2]. As a workaround, avoid mremap for resizing shared memory. Instead unmap the whole segment and map it back at the same address with the new size, relying on the fact that fd for the anon file behind the segment is still open and will keep the memory content. [1]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/mremap.c?id=f4d2ef48250ad057e4f00087967b5ff366da9f39#n1593 [2]: https://web.git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/mm/mremap.c?id=550a7d60bd5e35a56942dba6d8a26752beb26c9f --- src/backend/port/sysv_shmem.c | 60 +++++++++++++++++++++++++---------- 1 file changed, 44 insertions(+), 16 deletions(-) diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c index 87000a24eea..f0b53ce1d7c 100644 --- a/src/backend/port/sysv_shmem.c +++ b/src/backend/port/sysv_shmem.c @@ -1109,6 +1109,7 @@ AnonymousShmemResize(void) /* Note that CalculateShmemSize indirectly depends on NBuffers */ Size new_size = CalculateShmemSize(&numSemas, i); AnonymousMapping *m = &Mappings[i]; + int mmap_flags = PG_MMAP_FLAGS; if (m->shmem == NULL) continue; @@ -1116,6 +1117,44 @@ AnonymousShmemResize(void) if (m->shmem_size == new_size) continue; +#ifndef MAP_HUGETLB + /* ReserveAnonymousMemory should have dealt with this case */ + Assert(huge_pages != HUGE_PAGES_ON && !huge_pages_on); +#else + if (huge_pages_on) + { + Size hugepagesize; + + /* Make sure nothing is messed up */ + Assert(huge_pages == HUGE_PAGES_ON || huge_pages == HUGE_PAGES_TRY); + + /* Round up the new size to a suitable large value */ + GetHugePageSize(&hugepagesize, &mmap_flags, NULL); + + if (new_size % hugepagesize != 0) + new_size += hugepagesize - (new_size % hugepagesize); + + mmap_flags = PG_MMAP_FLAGS | mmap_flags; + } +#endif + + /* + * Linux limitations do not allow us to mremap hugetlb in the way we + * want. E.g. no size increase is allowed, and for shrinking the memory + * will not be released back. To work around this unmap the segment and + * create a new one at the same address. Thanks for the backing anon + * file the content will still be kept in memory. + */ + elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p", + MappingName(m->shmem_segment), m->shmem_size, + new_size, m->shmem); + + if (munmap(m->shmem, m->shmem_size) < 0) + ereport(FATAL, + (errcode(ERRCODE_SYSTEM_ERROR), + errmsg("could not unmap shared memory segment %s [%p]: %m", + MappingName(m->shmem_segment), m->shmem))); + /* Resize the backing anon file. */ if(ftruncate(m->segment_fd, new_size) == -1) ereport(FATAL, @@ -1123,25 +1162,14 @@ AnonymousShmemResize(void) errmsg("could not truncase anonymous file for \"%s\": %m", MappingName(m->shmem_segment)))); - /* Clean up some reserved space to resize into */ - if (munmap(m->shmem + m->shmem_size, new_size - m->shmem_size) == -1) - ereport(FATAL, - (errcode(ERRCODE_SYSTEM_ERROR), - errmsg("could not unmap %zu from reserved shared memory %p: %m", - new_size - m->shmem_size, m->shmem))); - - /* Claim the unused space */ - elog(DEBUG1, "segment[%s]: remap from %zu to %zu at address %p", - MappingName(m->shmem_segment), m->shmem_size, - new_size, m->shmem); - - ptr = mremap(m->shmem, m->shmem_size, new_size, 0); + /* Reclaim the space */ + ptr = mmap(m->shmem, new_size, PROT_READ | PROT_WRITE, + mmap_flags | MAP_FIXED, m->segment_fd, 0); if (ptr == MAP_FAILED) ereport(FATAL, (errcode(ERRCODE_SYSTEM_ERROR), - errmsg("could not resize shared memory segment %s [%p] to %d (%zu): %m", - MappingName(m->shmem_segment), m->shmem, NBuffers, - new_size))); + errmsg("could not map shared memory segment %s [%p] with size %zu: %m", + MappingName(m->shmem_segment), m->shmem, new_size))); reinit = true; m->shmem_size = new_size; -- 2.45.1 --vninua6xybvzgrci--