agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedRe: AIO v2.4
2+ messages / 1 participants
[nested] [flat]
* Re: AIO v2.4
@ 2025-02-19 19:10 Andres Freund <andres@anarazel.de>
2025-02-24 10:50 ` Re: AIO v2.4 Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 2+ messages in thread
From: Andres Freund @ 2025-02-19 19:10 UTC (permalink / raw)
To: pgsql-hackers; +Cc: Thomas Munro <thomas.munro@gmail.com>; pgsql-hackers@postgresql.org, Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi,
Attached is v2.4 of the AIO patchset.
Changes:
- Introduce "batchmode", while not in batchmode, IOs get submitted immediately.
Thomas didn't like how this worked previously, and while this was a
surprisingly large amount of work, I agree that it looks better now.
I vaccilated a bunch on the naming. For now it's
extern void pgaio_enter_batchmode(void);
extern void pgaio_exit_batchmode(void);
I did adjust the README and wrote a reasonably long comment above enter:
https://github.com/anarazel/postgres/blob/a324870186ddff9a31b10472b790eb4e744c40b3/src/backend/stora...
- Batchmode needs to be exited in case of errors, for that
- a new pgaio_after_error() call has been added to all the relevant places
- xact.c calls to aio have been (re-)added to check that there are no
in-progress batches / unsubmitted IOs at the end of a transaction.
Before that I had just removed at-eoxact "callbacks" :)
This checking has holes though:
https://postgr.es/m/upkkyhyuv6ultnejrutqcu657atw22kluh4lt2oidzxxtjqux3%40a4hdzamh4wzo
Because this only means that we will not detect all buggy code, rather
than misbehaving for correct code, I think this may be ok for now.
- Renamed aio_init.h to aio_subsys.h
The newly added pgaio_after_error() calls would have required including
aio.h in a good bit more places that won't themselves issue AIO. That seemed
wrong. There already was a aio_init.h to avoid needing to include aio.h in
places like ipci.c, but it seemed wrong to put pgaio_after_error() in
aio_init.h. So I renamed it to aio_subsys.h - not sure that's the best
name, but I can live with it.
- Now that Thomas submitted the necessary read_stream.c improvements, the
prior big TODO about one StartReadBuffers() call needing to start many IOs
has been addressed.
Thomas' thread: https://postgr.es/m/CA%2BhUKGK_%3D4CVmMHvsHjOVrK6t4F%3DLBpFzsrr3R%2BaJYN8kcTfWg%40mail.gmail.com
For now I've also included Thomas patches in my queue, but they should get
pushed independently. Review comments specific to those patches probably
are better put on the other thread.
Thomas' patches also fix several issues that were addressed in my WIP
adjustments to read_stream.c. There are a few left, but it does look
better.
The included commits are 0003-0008.
- I rewrote the tests into a tap test. That was exceedingly painful. Partially
due to tap infrastructure bugs on windows that would sometimes cause
inscrutable failures, see
https://www.postgresql.org/message-id/wmovm6xcbwh7twdtymxuboaoarbvwj2haasd3sikzlb3dkgz76%40n45rzyclu...
I just pushed that fix earlier today.
- Added docs for new GUCs, moved them to a more appropriate section
See also https://postgr.es/m/x3tlw2jk5gm3r3mv47hwrshffyw7halpczkfbk3peksxds7bvc%40lguk43z3bsyq
- If IO workers fail to reopen the file for an IO, the IO is now marked as
failed. Previously we'd just hang.
To test this I added an injection point that triggers the failure. I don't
know how else this could be tested.
- Added liburing dependency build documentation
- Added check hook to ensure io_max_concurrency = isn't set to 0 (-1 is for
auto-config)
- Fixed that with io_method == sync we'd issue fadvise calls when not
appropriate, that was a consequence of my hacky read_stream.c changes.
- Renamed some the aio<->bufmgr.c interface functions. Don't think they're
quite perfect, but they're in later patches, so I don't want to focus too
much on them rn.
- Comment improvements etc.
- Got rid of an unused wait event and renamed other wait events to make more
sense.
- Previously the injection points were added as part of the test patch, I now
moved them into the commits adding the code being tested. Was too annoying
to edit otherwise.
Todo:
- there's a decent amount of FIXMEs in later commits related to ereport(LOG)s
needing relpath() while in a critical section. I did propose a solution to
that yesterday:
https://postgr.es/m/h3a7ftrxypgxbw6ukcrrkspjon5dlninedwb5udkrase3rgqvn%403cokde6btlrl
- A few more corner case tests for the interaction of multiple backends trying
to do IO on overlapping buffers would be good.
- Our temp table test coverage is atrociously bad
Questions:
- The test module requires StartBufferIO() to be visible outside of bufmgr.c -
I think that's ok, would be good to know if others agree.
I'm planning to push the first two commits soon, I think they're ok on their
own, even if nothing else were to go in.
Greetings,
Andres Freund
Attachments:
[text/x-diff] v2.4-0001-Ensure-a-resowner-exists-for-all-paths-that-may.patch (2.6K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/2-v2.4-0001-Ensure-a-resowner-exists-for-all-paths-that-may.patch)
download | inline diff:
From 27fd90a15ef275451f1112f340e2133ee0cbdb97 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 8 Oct 2024 14:34:38 -0400
Subject: [PATCH v2.4 01/29] Ensure a resowner exists for all paths that may
perform AIO
Reviewed-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/1f6b50a7-38ef-4d87-8246-786d39f46ab9@iki.fi
---
src/backend/bootstrap/bootstrap.c | 7 +++++++
src/backend/replication/logical/logical.c | 6 ++++++
src/backend/utils/init/postinit.c | 6 +++++-
3 files changed, 18 insertions(+), 1 deletion(-)
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index 6db864892d0..e554504e1f0 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -361,8 +361,15 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
BaseInit();
bootstrap_signals();
+
+ /* need a resowner for IO during BootStrapXLOG() */
+ CreateAuxProcessResourceOwner();
+
BootStrapXLOG(bootstrap_data_checksum_version);
+ ReleaseAuxProcessResources(true);
+ CurrentResourceOwner = NULL;
+
/*
* To ensure that src/common/link-canary.c is linked into the backend, we
* must call it from somewhere. Here is as good as anywhere.
diff --git a/src/backend/replication/logical/logical.c b/src/backend/replication/logical/logical.c
index 8ea846bfc3b..2b936714040 100644
--- a/src/backend/replication/logical/logical.c
+++ b/src/backend/replication/logical/logical.c
@@ -386,6 +386,12 @@ CreateInitDecodingContext(const char *plugin,
slot->data.plugin = plugin_name;
SpinLockRelease(&slot->mutex);
+ if (CurrentResourceOwner == NULL)
+ {
+ Assert(am_walsender);
+ CurrentResourceOwner = AuxProcessResourceOwner;
+ }
+
if (XLogRecPtrIsInvalid(restart_lsn))
ReplicationSlotReserveWal();
else
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 01bb6a410cb..b491d04de58 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -755,8 +755,12 @@ InitPostgres(const char *in_dbname, Oid dboid,
* We don't yet have an aux-process resource owner, but StartupXLOG
* and ShutdownXLOG will need one. Hence, create said resource owner
* (and register a callback to clean it up after ShutdownXLOG runs).
+ *
+ * In bootstrap mode CreateAuxProcessResourceOwner() was already
+ * called in BootstrapModeMain().
*/
- CreateAuxProcessResourceOwner();
+ if (!bootstrap)
+ CreateAuxProcessResourceOwner();
StartupXLOG();
/* Release (and warn about) any buffer pins leaked in StartupXLOG */
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0002-Allow-lwlocks-to-be-unowned.patch (4.6K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/3-v2.4-0002-Allow-lwlocks-to-be-unowned.patch)
download | inline diff:
From 18b1f01230fe768b807c8391789a9920b6120e1a Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 5 Jan 2021 10:10:36 -0800
Subject: [PATCH v2.4 02/29] Allow lwlocks to be unowned
This is required for AIO so that the lock hold during a write can be released
in another backend. Which in turn is required to avoid the potential for
deadlocks.
---
src/include/storage/lwlock.h | 2 +
src/backend/storage/lmgr/lwlock.c | 108 +++++++++++++++++++++++-------
2 files changed, 85 insertions(+), 25 deletions(-)
diff --git a/src/include/storage/lwlock.h b/src/include/storage/lwlock.h
index 2aa46fd50da..13a7dc89980 100644
--- a/src/include/storage/lwlock.h
+++ b/src/include/storage/lwlock.h
@@ -129,6 +129,8 @@ extern bool LWLockAcquireOrWait(LWLock *lock, LWLockMode mode);
extern void LWLockRelease(LWLock *lock);
extern void LWLockReleaseClearVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64 val);
extern void LWLockReleaseAll(void);
+extern void LWLockDisown(LWLock *l);
+extern void LWLockReleaseDisowned(LWLock *l, LWLockMode mode);
extern bool LWLockHeldByMe(LWLock *lock);
extern bool LWLockAnyHeldByMe(LWLock *lock, int nlocks, size_t stride);
extern bool LWLockHeldByMeInMode(LWLock *lock, LWLockMode mode);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index f1e74f184f1..b02625194be 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -1773,36 +1773,15 @@ LWLockUpdateVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64 val)
}
}
-
/*
- * LWLockRelease - release a previously acquired lock
+ * Helper function to release lock, shared between LWLockRelease() and
+ * LWLockeleaseDisowned().
*/
-void
-LWLockRelease(LWLock *lock)
+static void
+LWLockReleaseInternal(LWLock *lock, LWLockMode mode)
{
- LWLockMode mode;
uint32 oldstate;
bool check_waiters;
- int i;
-
- /*
- * Remove lock from list of locks held. Usually, but not always, it will
- * be the latest-acquired lock; so search array backwards.
- */
- for (i = num_held_lwlocks; --i >= 0;)
- if (lock == held_lwlocks[i].lock)
- break;
-
- if (i < 0)
- elog(ERROR, "lock %s is not held", T_NAME(lock));
-
- mode = held_lwlocks[i].mode;
-
- num_held_lwlocks--;
- for (; i < num_held_lwlocks; i++)
- held_lwlocks[i] = held_lwlocks[i + 1];
-
- PRINT_LWDEBUG("LWLockRelease", lock, mode);
/*
* Release my hold on lock, after that it can immediately be acquired by
@@ -1840,6 +1819,85 @@ LWLockRelease(LWLock *lock)
LOG_LWDEBUG("LWLockRelease", lock, "releasing waiters");
LWLockWakeup(lock);
}
+}
+
+void
+LWLockReleaseDisowned(LWLock *lock, LWLockMode mode)
+{
+ LWLockReleaseInternal(lock, mode);
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * This is the code that can be shared between actually releasing a lock
+ * (LWLockRelease()) and just not tracking ownership of the lock anymore
+ * without releasing the lock (LWLockDisown()).
+ *
+ * Returns the mode in which the lock was held by the current backend.
+ *
+ * NB: This does not call RESUME_INTERRUPTS(), but leaves that responsibility
+ * of the caller.
+ *
+ * NB: This will leave lock->owner pointing to the current backend (if
+ * LOCK_DEBUG is set). This is somewhat intentional, as it makes it easier to
+ * debug cases of missing wakeups during lock release.
+ */
+static inline LWLockMode
+LWLockDisownInternal(LWLock *lock)
+{
+ LWLockMode mode;
+ int i;
+
+ /*
+ * Remove lock from list of locks held. Usually, but not always, it will
+ * be the latest-acquired lock; so search array backwards.
+ */
+ for (i = num_held_lwlocks; --i >= 0;)
+ if (lock == held_lwlocks[i].lock)
+ break;
+
+ if (i < 0)
+ elog(ERROR, "lock %s is not held", T_NAME(lock));
+
+ mode = held_lwlocks[i].mode;
+
+ num_held_lwlocks--;
+ for (; i < num_held_lwlocks; i++)
+ held_lwlocks[i] = held_lwlocks[i + 1];
+
+ return mode;
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * After calling this function it's the callers responsibility to ensure that
+ * the lock gets released (via LWLockReleaseDisowned()), even in case of an
+ * error. This only is desirable if the lock is going to be released in a
+ * different process than the process that acquired it.
+ */
+void
+LWLockDisown(LWLock *lock)
+{
+ LWLockDisownInternal(lock);
+
+ RESUME_INTERRUPTS();
+}
+
+/*
+ * LWLockRelease - release a previously acquired lock
+ */
+void
+LWLockRelease(LWLock *lock)
+{
+ LWLockMode mode;
+
+ mode = LWLockDisownInternal(lock);
+
+ PRINT_LWDEBUG("LWLockRelease", lock, mode);
+
+ LWLockReleaseInternal(lock, mode);
/*
* Now okay to allow cancel/die interrupts.
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0003-Refactor-read_stream.c-s-circular-arithmetic.patch (4.6K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/4-v2.4-0003-Refactor-read_stream.c-s-circular-arithmetic.patch)
download | inline diff:
From 9cf33cb6c80e867651681216d682ae2505e0e954 Mon Sep 17 00:00:00 2001
From: Thomas Munro <thomas.munro@gmail.com>
Date: Sat, 15 Feb 2025 14:47:25 +1300
Subject: [PATCH v2.4 03/29] Refactor read_stream.c's circular arithmetic.
Several places have open-coded circular index arithmetic. Make some
common functions for better readability and consistent assertion
checking.
This avoids adding yet more open-coded duplication in later patches, and
standardizes on the vocabulary "advance" and "retreat" as used elsewhere
in PostgreSQL.
---
src/backend/storage/aio/read_stream.c | 78 +++++++++++++++++++++------
1 file changed, 61 insertions(+), 17 deletions(-)
diff --git a/src/backend/storage/aio/read_stream.c b/src/backend/storage/aio/read_stream.c
index 99e44ed99fe..1c93fcae19b 100644
--- a/src/backend/storage/aio/read_stream.c
+++ b/src/backend/storage/aio/read_stream.c
@@ -224,6 +224,55 @@ read_stream_unget_block(ReadStream *stream, BlockNumber blocknum)
stream->buffered_blocknum = blocknum;
}
+/*
+ * Increment index, wrapping around at queue size.
+ */
+static inline void
+read_stream_index_advance(ReadStream *stream, int16 *index)
+{
+ Assert(*index >= 0);
+ Assert(*index < stream->queue_size);
+
+ *index += 1;
+ if (*index == stream->queue_size)
+ *index = 0;
+}
+
+/*
+ * Increment index by n, wrapping around at queue size.
+ */
+static inline void
+read_stream_index_advance_n(ReadStream *stream, int16 *index, int16 n)
+{
+ Assert(*index >= 0);
+ Assert(*index < stream->queue_size);
+ Assert(n <= MAX_IO_COMBINE_LIMIT);
+
+ *index += n;
+ if (*index >= stream->queue_size)
+ *index -= stream->queue_size;
+
+ Assert(*index >= 0);
+ Assert(*index < stream->queue_size);
+}
+
+#if defined(CLOBBER_FREED_MEMORY) || defined(USE_VALGRIND)
+/*
+ * Decrement index, wrapping around at queue size.
+ */
+static inline void
+read_stream_index_retreat(ReadStream *stream, int16 *index)
+{
+ Assert(*index >= 0);
+ Assert(*index < stream->queue_size);
+
+ if (*index == 0)
+ *index = stream->queue_size - 1;
+ else
+ *index -= 1;
+}
+#endif
+
static void
read_stream_start_pending_read(ReadStream *stream, bool suppress_advice)
{
@@ -302,11 +351,8 @@ read_stream_start_pending_read(ReadStream *stream, bool suppress_advice)
&stream->buffers[stream->queue_size],
sizeof(stream->buffers[0]) * overflow);
- /* Compute location of start of next read, without using % operator. */
- buffer_index += nblocks;
- if (buffer_index >= stream->queue_size)
- buffer_index -= stream->queue_size;
- Assert(buffer_index >= 0 && buffer_index < stream->queue_size);
+ /* Move to the location of start of next read. */
+ read_stream_index_advance_n(stream, &buffer_index, nblocks);
stream->next_buffer_index = buffer_index;
/* Adjust the pending read to cover the remaining portion, if any. */
@@ -334,12 +380,12 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
/*
* See which block the callback wants next in the stream. We need to
* compute the index of the Nth block of the pending read including
- * wrap-around, but we don't want to use the expensive % operator.
+ * wrap-around.
*/
- buffer_index = stream->next_buffer_index + stream->pending_read_nblocks;
- if (buffer_index >= stream->queue_size)
- buffer_index -= stream->queue_size;
- Assert(buffer_index >= 0 && buffer_index < stream->queue_size);
+ buffer_index = stream->next_buffer_index;
+ read_stream_index_advance_n(stream,
+ &buffer_index,
+ stream->pending_read_nblocks);
per_buffer_data = get_per_buffer_data(stream, buffer_index);
blocknum = read_stream_get_block(stream, per_buffer_data);
if (blocknum == InvalidBlockNumber)
@@ -777,12 +823,12 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
*/
if (stream->per_buffer_data)
{
+ int16 index;
void *per_buffer_data;
- per_buffer_data = get_per_buffer_data(stream,
- oldest_buffer_index == 0 ?
- stream->queue_size - 1 :
- oldest_buffer_index - 1);
+ index = oldest_buffer_index;
+ read_stream_index_retreat(stream, &index);
+ per_buffer_data = get_per_buffer_data(stream, index);
#if defined(CLOBBER_FREED_MEMORY)
/* This also tells Valgrind the memory is "noaccess". */
@@ -800,9 +846,7 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
stream->pinned_buffers--;
/* Advance oldest buffer, with wrap-around. */
- stream->oldest_buffer_index++;
- if (stream->oldest_buffer_index == stream->queue_size)
- stream->oldest_buffer_index = 0;
+ read_stream_index_advance(stream, &stream->oldest_buffer_index);
/* Prepare for the next call. */
read_stream_look_ahead(stream, false);
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0004-Allow-more-buffers-for-sequential-read-streams.patch (2.4K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/5-v2.4-0004-Allow-more-buffers-for-sequential-read-streams.patch)
download | inline diff:
From 7dda209d32625b9b237de0417f8ebc0783d8551c Mon Sep 17 00:00:00 2001
From: Thomas Munro <thomas.munro@gmail.com>
Date: Tue, 21 Jan 2025 08:08:08 +1300
Subject: [PATCH v2.4 04/29] Allow more buffers for sequential read streams.
Read streams currently only start concurrent I/Os (via read-ahead
advice) for random access, with a hard-coded guesstimate that their
average size is likely to be at most 4 blocks when planning the size of
the buffer queue. Sequential streams benefit from kernel readahead when
using buffered I/O, and read-ahead advice doesn't exist for direct I/O
by definition, so we didn't need to look ahead more than
io_combine_limit in that case.
Proposed patches need more buffers to be able start multiple
asynchronous I/O operations even for sequential access. Adjust the
arithmetic in preparation, replacing "4" with io_combine_limit, though
there is no benefit yet, just some wasted queue space.
As of the time of writing, the maximum GUC values for
effective_io_concurrent (1000) and io_combine_limit (32) imply a queue
with around 32K entries (slightly more for technical reasons), though
those numbers are likely to change. That requires a wider type in one
place that has a intermediate value that might overflow before clamping.
---
src/backend/storage/aio/read_stream.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/src/backend/storage/aio/read_stream.c b/src/backend/storage/aio/read_stream.c
index 1c93fcae19b..edeef292f75 100644
--- a/src/backend/storage/aio/read_stream.c
+++ b/src/backend/storage/aio/read_stream.c
@@ -499,7 +499,7 @@ read_stream_begin_impl(int flags,
* overflow (even though that's not possible with the current GUC range
* limits), allowing also for the spare entry and the overflow space.
*/
- max_pinned_buffers = Max(max_ios * 4, io_combine_limit);
+ max_pinned_buffers = Max(max_ios, 1) * io_combine_limit;
max_pinned_buffers = Min(max_pinned_buffers,
PG_INT16_MAX - io_combine_limit - 1);
@@ -771,7 +771,7 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
stream->ios[stream->oldest_io_index].buffer_index == oldest_buffer_index)
{
int16 io_index = stream->oldest_io_index;
- int16 distance;
+ int32 distance; /* wider temporary value, clamped below */
/* Sanity check that we still agree on the buffers. */
Assert(stream->ios[io_index].op.buffers ==
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0005-Improve-buffer-pool-API-for-per-backend-pin-lim.patch (6.4K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/6-v2.4-0005-Improve-buffer-pool-API-for-per-backend-pin-lim.patch)
download | inline diff:
From 47fbdb600a5011d06cf18db2f4791afb6b2d7ab4 Mon Sep 17 00:00:00 2001
From: Thomas Munro <thomas.munro@gmail.com>
Date: Fri, 24 Jan 2025 10:59:39 +1300
Subject: [PATCH v2.4 05/29] Improve buffer pool API for per-backend pin
limits.
Previously the support functions assumed that you needed one additional
pin to make progress, and could optionally use some more. Add a couple
more functions for callers that want to know:
* what the maximum possible number could be, for space planning
purposes, called the "soft pin limit"
* how many additional pins they could acquire right now, without the
special case allowing one pin (ie for users that already hold pins and
can already make progress even if zero extra pins are available now)
These APIs are better suited to read_stream.c, which will be adjusted in
a follow-up patch. Also move the computation of the each backend's fair
share of the buffer pool to backend initialization time, since the
answer doesn't change and we don't want to perform a division operation
every time we compute availability.
---
src/include/storage/bufmgr.h | 4 ++
src/backend/storage/buffer/bufmgr.c | 75 ++++++++++++++++++++-------
src/backend/storage/buffer/localbuf.c | 16 ++++++
3 files changed, 77 insertions(+), 18 deletions(-)
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 7c1e4316dde..597ecb97897 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -290,6 +290,10 @@ extern bool HoldingBufferPinThatDelaysRecovery(void);
extern bool BgBufferSync(struct WritebackContext *wb_context);
+extern uint32 GetSoftPinLimit(void);
+extern uint32 GetSoftLocalPinLimit(void);
+extern uint32 GetAdditionalPinLimit(void);
+extern uint32 GetAdditionalLocalPinLimit(void);
extern void LimitAdditionalPins(uint32 *additional_pins);
extern void LimitAdditionalLocalPins(uint32 *additional_pins);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 75cfc2b6fe9..5ad1e2b18a9 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -211,6 +211,8 @@ static int32 PrivateRefCountOverflowed = 0;
static uint32 PrivateRefCountClock = 0;
static PrivateRefCountEntry *ReservedRefCountEntry = NULL;
+static uint32 MaxProportionalPins;
+
static void ReservePrivateRefCountEntry(void);
static PrivateRefCountEntry *NewPrivateRefCountEntry(Buffer buffer);
static PrivateRefCountEntry *GetPrivateRefCountEntry(Buffer buffer, bool do_move);
@@ -2097,6 +2099,46 @@ again:
return buf;
}
+/*
+ * Return the maximum number of buffer than this backend should try to pin at
+ * once, to avoid pinning more than its fair share. This is the highest value
+ * that GetAdditionalPinLimit() and LimitAdditionalPins() could ever return.
+ *
+ * It's called a soft limit because nothing stops a backend from trying to
+ * acquire more pins than this this with ReadBuffer(), but code that wants more
+ * for I/O optimizations should respect this per-backend limit when it can
+ * still make progress without them.
+ */
+uint32
+GetSoftPinLimit(void)
+{
+ return MaxProportionalPins;
+}
+
+/*
+ * Return the maximum number of additional buffers that this backend should
+ * pin if it wants to stay under the per-backend soft limit, considering the
+ * number of buffers it has already pinned.
+ */
+uint32
+GetAdditionalPinLimit(void)
+{
+ uint32 estimated_pins_held;
+
+ /*
+ * We get the number of "overflowed" pins for free, but don't know the
+ * number of pins in PrivateRefCountArray. The cost of calculating that
+ * exactly doesn't seem worth it, so just assume the max.
+ */
+ estimated_pins_held = PrivateRefCountOverflowed + REFCOUNT_ARRAY_ENTRIES;
+
+ /* Is this backend already holding more than its fair share? */
+ if (estimated_pins_held > MaxProportionalPins)
+ return 0;
+
+ return MaxProportionalPins - estimated_pins_held;
+}
+
/*
* Limit the number of pins a batch operation may additionally acquire, to
* avoid running out of pinnable buffers.
@@ -2112,28 +2154,15 @@ again:
void
LimitAdditionalPins(uint32 *additional_pins)
{
- uint32 max_backends;
- int max_proportional_pins;
+ uint32 limit;
if (*additional_pins <= 1)
return;
- max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
- max_proportional_pins = NBuffers / max_backends;
-
- /*
- * Subtract the approximate number of buffers already pinned by this
- * backend. We get the number of "overflowed" pins for free, but don't
- * know the number of pins in PrivateRefCountArray. The cost of
- * calculating that exactly doesn't seem worth it, so just assume the max.
- */
- max_proportional_pins -= PrivateRefCountOverflowed + REFCOUNT_ARRAY_ENTRIES;
-
- if (max_proportional_pins <= 0)
- max_proportional_pins = 1;
-
- if (*additional_pins > max_proportional_pins)
- *additional_pins = max_proportional_pins;
+ limit = GetAdditionalPinLimit();
+ limit = Max(limit, 1);
+ if (limit < *additional_pins)
+ *additional_pins = limit;
}
/*
@@ -3574,6 +3603,16 @@ InitBufferManagerAccess(void)
{
HASHCTL hash_ctl;
+ /*
+ * The soft limit on the number of pins each backend should respect, bast
+ * on shared_buffers and the maximum number of connections possible.
+ * That's very pessimistic, but outside toy-sized shared_buffers it should
+ * allow plenty of pins. Higher level code that pins non-trivial numbers
+ * of buffers should use LimitAdditionalPins() or GetAdditionalPinLimit()
+ * to stay under this limit.
+ */
+ MaxProportionalPins = NBuffers / (MaxBackends + NUM_AUXILIARY_PROCS);
+
memset(&PrivateRefCountArray, 0, sizeof(PrivateRefCountArray));
hash_ctl.keysize = sizeof(int32);
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 64931efaa75..3c055f6ec8b 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -286,6 +286,22 @@ GetLocalVictimBuffer(void)
return BufferDescriptorGetBuffer(bufHdr);
}
+/* see GetSoftPinLimit() */
+uint32
+GetSoftLocalPinLimit(void)
+{
+ /* Every backend has its own temporary buffers, and can pin them all. */
+ return num_temp_buffers;
+}
+
+/* see GetAdditionalPinLimit() */
+uint32
+GetAdditionalLocalPinLimit(void)
+{
+ Assert(NLocalPinnedBuffers <= num_temp_buffers);
+ return num_temp_buffers - NLocalPinnedBuffers;
+}
+
/* see LimitAdditionalPins() */
void
LimitAdditionalLocalPins(uint32 *additional_pins)
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0006-Respect-pin-limits-accurately-in-read_stream.c.patch (9.9K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/7-v2.4-0006-Respect-pin-limits-accurately-in-read_stream.c.patch)
download | inline diff:
From ccf77f7584d30ceb9997a0d47bd1e3e05cab1cdc Mon Sep 17 00:00:00 2001
From: Thomas Munro <thomas.munro@gmail.com>
Date: Fri, 24 Jan 2025 23:52:53 +1300
Subject: [PATCH v2.4 06/29] Respect pin limits accurately in read_stream.c.
Read streams pin multiple buffers at once as required to combine I/O.
This also avoids having to unpin and repin later when issuing read-ahead
advice, and will be needed for proposed work that starts "real"
asynchronous I/O.
To avoid pinning too much of the buffer pool at once, we previously used
LimitAdditionalBuffers() to avoid pinning more than this backend's fair
share of the pool as a cap. The coding was a naive and only checked the
cap once at stream initialization.
This commit moves the check to the time of use with new bufmgr APIs from
an earlier commit, since the result might change later due to pins
acquired later outside this stream. No extra CPU cycles are added to
the all-buffered fast-path code (it only pins one buffer at a time), but
the I/O-starting path now re-checks the limit every time using simple
arithmetic.
In practice it was difficult to exceed the limit, but you could contrive
a workload to do it using multiple CURSORs and FETCHing from sequential
scans in round-robin fashion, so that each underlying stream computes
its limit before all the others have ramped up to their full look-ahead
distance. Therefore, no back-patch for now.
Per code review from Andres, in the course of his AIO work.
Reported-by: Andres Freund <andres@anarazel.de>
---
src/backend/storage/aio/read_stream.c | 111 ++++++++++++++++++++++----
1 file changed, 95 insertions(+), 16 deletions(-)
diff --git a/src/backend/storage/aio/read_stream.c b/src/backend/storage/aio/read_stream.c
index edeef292f75..1a51e6eed31 100644
--- a/src/backend/storage/aio/read_stream.c
+++ b/src/backend/storage/aio/read_stream.c
@@ -115,6 +115,7 @@ struct ReadStream
int16 pinned_buffers;
int16 distance;
bool advice_enabled;
+ bool temporary;
/*
* One-block buffer to support 'ungetting' a block number, to resolve flow
@@ -274,7 +275,9 @@ read_stream_index_retreat(ReadStream *stream, int16 *index)
#endif
static void
-read_stream_start_pending_read(ReadStream *stream, bool suppress_advice)
+read_stream_start_pending_read(ReadStream *stream,
+ int16 buffer_limit,
+ bool suppress_advice)
{
bool need_wait;
int nblocks;
@@ -308,10 +311,14 @@ read_stream_start_pending_read(ReadStream *stream, bool suppress_advice)
else
flags = 0;
- /* We say how many blocks we want to read, but may be smaller on return. */
+ /*
+ * We say how many blocks we want to read, but may be smaller on return.
+ * On memory-constrained systems we may be also have to ask for a smaller
+ * read ourselves.
+ */
buffer_index = stream->next_buffer_index;
io_index = stream->next_io_index;
- nblocks = stream->pending_read_nblocks;
+ nblocks = Min(buffer_limit, stream->pending_read_nblocks);
need_wait = StartReadBuffers(&stream->ios[io_index].op,
&stream->buffers[buffer_index],
stream->pending_read_blocknum,
@@ -360,11 +367,60 @@ read_stream_start_pending_read(ReadStream *stream, bool suppress_advice)
stream->pending_read_nblocks -= nblocks;
}
+/*
+ * How many more buffers could we use, while respecting the soft limit?
+ */
+static int16
+read_stream_get_buffer_limit(ReadStream *stream)
+{
+ uint32 buffers;
+
+ /* Check how many local or shared pins we could acquire. */
+ if (stream->temporary)
+ buffers = GetAdditionalLocalPinLimit();
+ else
+ buffers = GetAdditionalPinLimit();
+
+ /*
+ * Each stream is always allowed to try to acquire one pin if it doesn't
+ * hold one already. This is needed to guarantee progress, and just like
+ * the simple ReadBuffer() operation in code that is not using this stream
+ * API, if a buffer can't be pinned we'll raise an error when trying to
+ * pin, ie the buffer pool is simply too small for the workload.
+ */
+ if (buffers == 0 && stream->pinned_buffers == 0)
+ return 1;
+
+ /*
+ * Otherwise, see how many additional pins the backend can currently pin,
+ * which may be zero. As above, this only guarantees that this backend
+ * won't use more than its fair share if all backends can respect the soft
+ * limit, not that a pin can actually be acquired without error.
+ */
+ return Min(buffers, INT16_MAX);
+}
+
static void
read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
{
+ int16 buffer_limit;
+
+ /*
+ * Check how many pins we could acquire now. We do this here rather than
+ * pushing it down into read_stream_start_pending_read(), because it
+ * allows more flexibility in behavior when we run out of allowed pins.
+ * Currently the policy is to start an I/O when we've run out of allowed
+ * pins only if we have to to make progress, and otherwise to stop looking
+ * ahead until more pins become available, so that we don't start issuing
+ * a lot of smaller I/Os, prefering to build the largest ones we can. This
+ * choice is debatable, but it should only really come up with the buffer
+ * pool/connection ratio is very constrained.
+ */
+ buffer_limit = read_stream_get_buffer_limit(stream);
+
while (stream->ios_in_progress < stream->max_ios &&
- stream->pinned_buffers + stream->pending_read_nblocks < stream->distance)
+ stream->pinned_buffers + stream->pending_read_nblocks <
+ Min(stream->distance, buffer_limit))
{
BlockNumber blocknum;
int16 buffer_index;
@@ -372,7 +428,9 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
if (stream->pending_read_nblocks == io_combine_limit)
{
- read_stream_start_pending_read(stream, suppress_advice);
+ read_stream_start_pending_read(stream, buffer_limit,
+ suppress_advice);
+ buffer_limit = read_stream_get_buffer_limit(stream);
suppress_advice = false;
continue;
}
@@ -406,11 +464,12 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
/* We have to start the pending read before we can build another. */
while (stream->pending_read_nblocks > 0)
{
- read_stream_start_pending_read(stream, suppress_advice);
+ read_stream_start_pending_read(stream, buffer_limit, suppress_advice);
+ buffer_limit = read_stream_get_buffer_limit(stream);
suppress_advice = false;
- if (stream->ios_in_progress == stream->max_ios)
+ if (stream->ios_in_progress == stream->max_ios || buffer_limit == 0)
{
- /* And we've hit the limit. Rewind, and stop here. */
+ /* And we've hit a limit. Rewind, and stop here. */
read_stream_unget_block(stream, blocknum);
return;
}
@@ -426,16 +485,17 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
* limit, preferring to give it another chance to grow to full
* io_combine_limit size once more buffers have been consumed. However,
* if we've already reached io_combine_limit, or we've reached the
- * distance limit and there isn't anything pinned yet, or the callback has
- * signaled end-of-stream, we start the read immediately.
+ * distance limit or buffer limit and there isn't anything pinned yet, or
+ * the callback has signaled end-of-stream, we start the read immediately.
*/
if (stream->pending_read_nblocks > 0 &&
(stream->pending_read_nblocks == io_combine_limit ||
- (stream->pending_read_nblocks == stream->distance &&
+ ((stream->pending_read_nblocks == stream->distance ||
+ stream->pending_read_nblocks == buffer_limit) &&
stream->pinned_buffers == 0) ||
stream->distance == 0) &&
stream->ios_in_progress < stream->max_ios)
- read_stream_start_pending_read(stream, suppress_advice);
+ read_stream_start_pending_read(stream, buffer_limit, suppress_advice);
}
/*
@@ -464,6 +524,7 @@ read_stream_begin_impl(int flags,
int max_ios;
int strategy_pin_limit;
uint32 max_pinned_buffers;
+ uint32 max_possible_buffer_limit;
Oid tablespace_id;
/*
@@ -507,12 +568,23 @@ read_stream_begin_impl(int flags,
strategy_pin_limit = GetAccessStrategyPinLimit(strategy);
max_pinned_buffers = Min(strategy_pin_limit, max_pinned_buffers);
- /* Don't allow this backend to pin more than its share of buffers. */
+ /*
+ * Also limit by the maximum possible number of pins we could be allowed
+ * to acquire according to bufmgr. We may not be able to use them all due
+ * to other pins held by this backend, but we'll enforce the dynamic limit
+ * later when starting I/O.
+ */
if (SmgrIsTemp(smgr))
- LimitAdditionalLocalPins(&max_pinned_buffers);
+ max_possible_buffer_limit = GetSoftLocalPinLimit();
else
- LimitAdditionalPins(&max_pinned_buffers);
- Assert(max_pinned_buffers > 0);
+ max_possible_buffer_limit = GetSoftPinLimit();
+ max_pinned_buffers = Min(max_pinned_buffers, max_possible_buffer_limit);
+
+ /*
+ * The soft limit might be zero on a system configured with more
+ * connections than buffers. We need at least one.
+ */
+ max_pinned_buffers = Max(1, max_pinned_buffers);
/*
* We need one extra entry for buffers and per-buffer data, because users
@@ -572,6 +644,7 @@ read_stream_begin_impl(int flags,
stream->callback = callback;
stream->callback_private_data = callback_private_data;
stream->buffered_blocknum = InvalidBlockNumber;
+ stream->temporary = SmgrIsTemp(smgr);
/*
* Skip the initial ramp-up phase if the caller says we're going to be
@@ -700,6 +773,12 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
* arbitrary I/O entry (they're all free). We don't have to
* adjust pinned_buffers because we're transferring one to caller
* but pinning one more.
+ *
+ * In the fast path we don't need to check the pin limit. We're
+ * always allowed at least one pin so that progress can be made,
+ * and that's all we need here. Although two pins are momentarily
+ * held at the same time, the model used here is that the stream
+ * holds only one, and the other now belongs to the caller.
*/
if (likely(!StartReadBuffer(&stream->ios[0].op,
&stream->buffers[oldest_buffer_index],
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0007-Support-buffer-forwarding-in-read_stream.c.patch (8.9K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/8-v2.4-0007-Support-buffer-forwarding-in-read_stream.c.patch)
download | inline diff:
From 2dc7addc4f3965ec958c3cd288ec278f4d14a602 Mon Sep 17 00:00:00 2001
From: Thomas Munro <thomas.munro@gmail.com>
Date: Thu, 30 Jan 2025 11:42:03 +1300
Subject: [PATCH v2.4 07/29] Support buffer forwarding in read_stream.c.
In preparation for a following change to the buffer manager, teach read
stream to keep track of buffers that were "forwarded" from one call to
StartReadBuffers() to the next.
Since StartReadBuffers() buffers argument will become an in/out
argument, we need to initialize the buffer queue entries with
InvalidBuffer. We don't want to do that up front, because we try to
keep stream initialization cheap and code that uses the fast path stays
in one single buffer queue element. Satisfy both goals by initializing
the queue incrementally on the first cycle.
---
src/backend/storage/aio/read_stream.c | 108 ++++++++++++++++++++++----
1 file changed, 94 insertions(+), 14 deletions(-)
diff --git a/src/backend/storage/aio/read_stream.c b/src/backend/storage/aio/read_stream.c
index 1a51e6eed31..32e5def29f8 100644
--- a/src/backend/storage/aio/read_stream.c
+++ b/src/backend/storage/aio/read_stream.c
@@ -112,8 +112,10 @@ struct ReadStream
int16 ios_in_progress;
int16 queue_size;
int16 max_pinned_buffers;
+ int16 forwarded_buffers;
int16 pinned_buffers;
int16 distance;
+ int16 initialized_buffers;
bool advice_enabled;
bool temporary;
@@ -280,7 +282,9 @@ read_stream_start_pending_read(ReadStream *stream,
bool suppress_advice)
{
bool need_wait;
+ int requested_nblocks;
int nblocks;
+ int forwarded;
int flags;
int16 io_index;
int16 overflow;
@@ -312,13 +316,34 @@ read_stream_start_pending_read(ReadStream *stream,
flags = 0;
/*
- * We say how many blocks we want to read, but may be smaller on return.
- * On memory-constrained systems we may be also have to ask for a smaller
- * read ourselves.
+ * On buffer-constrained systems we may need to limit the I/O size by the
+ * available pin count.
*/
+ requested_nblocks = Min(buffer_limit, stream->pending_read_nblocks);
+ nblocks = requested_nblocks;
buffer_index = stream->next_buffer_index;
io_index = stream->next_io_index;
- nblocks = Min(buffer_limit, stream->pending_read_nblocks);
+
+ /*
+ * The first time around the queue we initialize it as we go, including
+ * the overflow zone, because otherwise the entries would appear as
+ * forwarded buffers. This avoids initializing the whole queue up front
+ * in cases where it is large but we don't ever use it due to the
+ * all-cached fast path or small scans.
+ */
+ while (stream->initialized_buffers < buffer_index + nblocks)
+ stream->buffers[stream->initialized_buffers++] = InvalidBuffer;
+
+ /*
+ * Start the I/O. Any buffers that are not InvalidBuffer will be
+ * interpreted as already pinned, forwarded by an earlier call to
+ * StartReadBuffers(), and must map to the expected blocks. The nblocks
+ * value may be smaller on return indicating the size of the I/O that
+ * could be started. Buffers beyond the output nblocks number may also
+ * have been pinned without starting I/O due to various edge cases. In
+ * that case we'll just leave them in the queue ahead of us, "forwarded"
+ * to the next call, avoiding the need to unpin/repin.
+ */
need_wait = StartReadBuffers(&stream->ios[io_index].op,
&stream->buffers[buffer_index],
stream->pending_read_blocknum,
@@ -347,16 +372,35 @@ read_stream_start_pending_read(ReadStream *stream,
stream->seq_blocknum = stream->pending_read_blocknum + nblocks;
}
+ /*
+ * How many pins were acquired but forwarded to the next call? These need
+ * to be passed to the next StartReadBuffers() call, or released if the
+ * stream ends early. We need the number for accounting purposes, since
+ * they are not counted in stream->pinned_buffers but we already hold
+ * them.
+ */
+ forwarded = 0;
+ while (nblocks + forwarded < requested_nblocks &&
+ stream->buffers[buffer_index + nblocks + forwarded] != InvalidBuffer)
+ forwarded++;
+ stream->forwarded_buffers = forwarded;
+
/*
* We gave a contiguous range of buffer space to StartReadBuffers(), but
- * we want it to wrap around at queue_size. Slide overflowing buffers to
- * the front of the array.
+ * we want it to wrap around at queue_size. Copy overflowing buffers to
+ * the front of the array where they'll be consumed, but also leave a copy
+ * in the overflow zone which the I/O operation has a pointer to (it needs
+ * a contiguous array). Both copies will be cleared when the buffers are
+ * handed to the consumer.
*/
- overflow = (buffer_index + nblocks) - stream->queue_size;
+ overflow = (buffer_index + nblocks + forwarded) - stream->queue_size;
if (overflow > 0)
- memmove(&stream->buffers[0],
- &stream->buffers[stream->queue_size],
- sizeof(stream->buffers[0]) * overflow);
+ {
+ Assert(overflow < stream->queue_size); /* can't overlap */
+ memcpy(&stream->buffers[0],
+ &stream->buffers[stream->queue_size],
+ sizeof(stream->buffers[0]) * overflow);
+ }
/* Move to the location of start of next read. */
read_stream_index_advance_n(stream, &buffer_index, nblocks);
@@ -381,6 +425,15 @@ read_stream_get_buffer_limit(ReadStream *stream)
else
buffers = GetAdditionalPinLimit();
+ /*
+ * If we already have some forwarded buffers, we can certainly use those.
+ * They are already pinned, and are mapped to the starting blocks of the
+ * pending read, they just don't have any I/O started yet and are not
+ * counted in stream->pinned_buffers.
+ */
+ Assert(stream->forwarded_buffers <= stream->pending_read_nblocks);
+ buffers += stream->forwarded_buffers;
+
/*
* Each stream is always allowed to try to acquire one pin if it doesn't
* hold one already. This is needed to guarantee progress, and just like
@@ -389,7 +442,7 @@ read_stream_get_buffer_limit(ReadStream *stream)
* pin, ie the buffer pool is simply too small for the workload.
*/
if (buffers == 0 && stream->pinned_buffers == 0)
- return 1;
+ buffers = 1;
/*
* Otherwise, see how many additional pins the backend can currently pin,
@@ -751,10 +804,12 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
/* Fast path assumptions. */
Assert(stream->ios_in_progress == 0);
+ Assert(stream->forwarded_buffers == 0);
Assert(stream->pinned_buffers == 1);
Assert(stream->distance == 1);
Assert(stream->pending_read_nblocks == 0);
Assert(stream->per_buffer_data_size == 0);
+ Assert(stream->initialized_buffers > stream->oldest_buffer_index);
/* We're going to return the buffer we pinned last time. */
oldest_buffer_index = stream->oldest_buffer_index;
@@ -803,6 +858,7 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
stream->distance = 0;
stream->oldest_buffer_index = stream->next_buffer_index;
stream->pinned_buffers = 0;
+ stream->buffers[oldest_buffer_index] = InvalidBuffer;
}
stream->fast_path = false;
@@ -887,10 +943,15 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
}
}
-#ifdef CLOBBER_FREED_MEMORY
- /* Clobber old buffer for debugging purposes. */
+ /*
+ * We must zap this queue entry, or else it would appear as a forwarded
+ * buffer. If it's potentially in the overflow zone (ie it wrapped around
+ * the queue), also zap that copy.
+ */
stream->buffers[oldest_buffer_index] = InvalidBuffer;
-#endif
+ if (oldest_buffer_index < io_combine_limit - 1)
+ stream->buffers[stream->queue_size + oldest_buffer_index] =
+ InvalidBuffer;
#if defined(CLOBBER_FREED_MEMORY) || defined(USE_VALGRIND)
@@ -933,6 +994,7 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
#ifndef READ_STREAM_DISABLE_FAST_PATH
/* See if we can take the fast path for all-cached scans next time. */
if (stream->ios_in_progress == 0 &&
+ stream->forwarded_buffers == 0 &&
stream->pinned_buffers == 1 &&
stream->distance == 1 &&
stream->pending_read_nblocks == 0 &&
@@ -968,6 +1030,7 @@ read_stream_next_block(ReadStream *stream, BufferAccessStrategy *strategy)
void
read_stream_reset(ReadStream *stream)
{
+ int16 index;
Buffer buffer;
/* Stop looking ahead. */
@@ -981,6 +1044,23 @@ read_stream_reset(ReadStream *stream)
while ((buffer = read_stream_next_buffer(stream, NULL)) != InvalidBuffer)
ReleaseBuffer(buffer);
+ /* Unpin any unused forwarded buffers. */
+ index = stream->next_buffer_index;
+ while (index < stream->initialized_buffers &&
+ (buffer = stream->buffers[index]) != InvalidBuffer)
+ {
+ Assert(stream->forwarded_buffers > 0);
+ stream->forwarded_buffers--;
+ ReleaseBuffer(buffer);
+
+ stream->buffers[index] = InvalidBuffer;
+ if (index < io_combine_limit - 1)
+ stream->buffers[stream->queue_size + index] = InvalidBuffer;
+
+ read_stream_index_advance(stream, &index);
+ }
+
+ Assert(stream->forwarded_buffers == 0);
Assert(stream->pinned_buffers == 0);
Assert(stream->ios_in_progress == 0);
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0008-Support-buffer-forwarding-in-StartReadBuffers.patch (10.5K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/9-v2.4-0008-Support-buffer-forwarding-in-StartReadBuffers.patch)
download | inline diff:
From 6c9f6faffb23c81250298975112128b6521da429 Mon Sep 17 00:00:00 2001
From: Thomas Munro <thomas.munro@gmail.com>
Date: Mon, 10 Feb 2025 21:55:40 +1300
Subject: [PATCH v2.4 08/29] Support buffer forwarding in StartReadBuffers().
Sometimes we have to perform a short read because we hit a cached block
that ends a contiguous run of blocks requiring I/O. We don't want
StartReadBuffers() to have to start more than one I/O, so we stop there.
We also don't want to have to unpin the cached block (and repin it
later), so previously we'd silently pretend the hit was part of the I/O,
and just leave it out of the read from disk. Now, we'll "forward" it to
the next call. We still write it to the buffers[] array for the caller
to pass back to us later, but it's not included in *nblocks.
This policy means that we no longer mix hits and misses in a single
operation's results, so we avoid the requirement to call
WaitReadBuffers(), which might stall, before the caller can make use of
the hits. The caller will get the hit in the next call instead, and
know that it doesn't have to wait. That's important for later work on
out-of-order read streams that minimize I/O stalls.
This also makes life easier for proposed work on true AIO, which
occasionally needs to split a large I/O after pinning all the buffers,
while the current coding only ever forwards a single bookending hit.
This API is natural for read_stream.c: it just leaves forwarded buffers
where they are in its circular queue, where the next call will pick them
up and continue, minimizing pin churn.
If we ever think of a good reason to disable this feature, i.e. for
other users of StartReadBuffers() that don't want to deal with forwarded
buffers, then we could add a flag for that. For now read_steam.c is the
only user.
---
src/include/storage/bufmgr.h | 1 -
src/backend/storage/buffer/bufmgr.c | 128 ++++++++++++++++++++--------
2 files changed, 91 insertions(+), 38 deletions(-)
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 597ecb97897..4a035f59a7d 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -130,7 +130,6 @@ struct ReadBuffersOperation
BlockNumber blocknum;
int flags;
int16 nblocks;
- int16 io_buffers_len;
};
typedef struct ReadBuffersOperation ReadBuffersOperation;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 5ad1e2b18a9..47e1c3442b4 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -1257,10 +1257,10 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
Buffer *buffers,
BlockNumber blockNum,
int *nblocks,
- int flags)
+ int flags,
+ bool allow_forwarding)
{
int actual_nblocks = *nblocks;
- int io_buffers_len = 0;
int maxcombine = 0;
Assert(*nblocks > 0);
@@ -1270,30 +1270,80 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
{
bool found;
- buffers[i] = PinBufferForBlock(operation->rel,
- operation->smgr,
- operation->persistence,
- operation->forknum,
- blockNum + i,
- operation->strategy,
- &found);
+ if (allow_forwarding && buffers[i] != InvalidBuffer)
+ {
+ BufferDesc *bufHdr;
+
+ /*
+ * This is a buffer that was pinned by an earlier call to
+ * StartReadBuffers(), but couldn't be handled in one operation at
+ * that time. The operation was split, and the caller has passed
+ * an already pinned buffer back to us to handle the rest of the
+ * operation. It must continue at the expected block number.
+ */
+ Assert(BufferGetBlockNumber(buffers[i]) == blockNum + i);
+
+ /*
+ * It might be an already valid buffer (a hit) that followed the
+ * final contiguous block of an earlier I/O (a miss) marking the
+ * end of it, or a buffer that some other backend has since made
+ * valid by performing the I/O for us, in which case we can handle
+ * it as a hit now. It is safe to check for a BM_VALID flag with
+ * a relaxed load, because we got a fresh view of it while pinning
+ * it in the previous call.
+ *
+ * On the other hand if we don't see BM_VALID yet, it must be an
+ * I/O that was split by the previous call and we need to try to
+ * start a new I/O from this block. We're also racing against any
+ * other backend that might start the I/O or even manage to mark
+ * it BM_VALID after this check, BM_VALID after this check, but
+ * StartBufferIO() will handle those cases.
+ */
+ if (BufferIsLocal(buffers[i]))
+ bufHdr = GetLocalBufferDescriptor(-buffers[i] - 1);
+ else
+ bufHdr = GetBufferDescriptor(buffers[i] - 1);
+ found = pg_atomic_read_u32(&bufHdr->state) & BM_VALID;
+ }
+ else
+ {
+ buffers[i] = PinBufferForBlock(operation->rel,
+ operation->smgr,
+ operation->persistence,
+ operation->forknum,
+ blockNum + i,
+ operation->strategy,
+ &found);
+ }
if (found)
{
/*
- * Terminate the read as soon as we get a hit. It could be a
- * single buffer hit, or it could be a hit that follows a readable
- * range. We don't want to create more than one readable range,
- * so we stop here.
+ * We have a hit. If it's the first block in the requested range,
+ * we can return it immediately and report that WaitReadBuffers()
+ * does not need to be called. If the initial value of *nblocks
+ * was larger, the caller will have to call again for the rest.
*/
- actual_nblocks = i + 1;
+ if (i == 0)
+ {
+ *nblocks = 1;
+ return false;
+ }
+
+ /*
+ * Otherwise we already have an I/O to perform, but this block
+ * can't be included as it is already valid. Split the I/O here.
+ * There may or may not be more blocks requiring I/O after this
+ * one, we haven't checked, but it can't be contiguous with this
+ * hit in the way. We'll leave this buffer pinned, forwarding it
+ * to the next call, avoiding the need to unpin it here and re-pin
+ * it in the next call.
+ */
+ actual_nblocks = i;
break;
}
else
{
- /* Extend the readable range to cover this block. */
- io_buffers_len++;
-
/*
* Check how many blocks we can cover with the same IO. The smgr
* implementation might e.g. be limited due to a segment boundary.
@@ -1314,15 +1364,11 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
}
*nblocks = actual_nblocks;
- if (likely(io_buffers_len == 0))
- return false;
-
/* Populate information needed for I/O. */
operation->buffers = buffers;
operation->blocknum = blockNum;
operation->flags = flags;
operation->nblocks = actual_nblocks;
- operation->io_buffers_len = io_buffers_len;
if (flags & READ_BUFFERS_ISSUE_ADVICE)
{
@@ -1337,7 +1383,7 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
smgrprefetch(operation->smgr,
operation->forknum,
blockNum,
- operation->io_buffers_len);
+ actual_nblocks);
}
/* Indicate that WaitReadBuffers() should be called. */
@@ -1351,11 +1397,21 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
* actual number, which may be fewer than requested. Caller sets some of the
* members of operation; see struct definition.
*
+ * The initial contents of the elements of buffers up to *nblocks should
+ * either be InvalidBuffer or an already-pinned buffer that was left by an
+ * preceding call to StartReadBuffers() that had to be split. On return, some
+ * elements of buffers may hold pinned buffers beyond the number indicated by
+ * the updated value of *nblocks. Operations are split on boundaries known to
+ * smgr (eg md.c segment boundaries that require crossing into a different
+ * underlying file), or when already cached blocks are found in the buffer
+ * that prevent the formation of a contiguous read.
+ *
* If false is returned, no I/O is necessary. If true is returned, one I/O
* has been started, and WaitReadBuffers() must be called with the same
* operation object before the buffers are accessed. Along with the operation
* object, the caller-supplied array of buffers must remain valid until
- * WaitReadBuffers() is called.
+ * WaitReadBuffers() is called, and any forwarded buffers must also be
+ * preserved for a future call unless explicitly released.
*
* Currently the I/O is only started with optional operating system advice if
* requested by the caller with READ_BUFFERS_ISSUE_ADVICE, and the real I/O
@@ -1369,13 +1425,18 @@ StartReadBuffers(ReadBuffersOperation *operation,
int *nblocks,
int flags)
{
- return StartReadBuffersImpl(operation, buffers, blockNum, nblocks, flags);
+ return StartReadBuffersImpl(operation, buffers, blockNum, nblocks, flags,
+ true /* expect forwarded buffers */ );
}
/*
* Single block version of the StartReadBuffers(). This might save a few
* instructions when called from another translation unit, because it is
* specialized for nblocks == 1.
+ *
+ * This version does not support "forwarded" buffers: they cannot be created
+ * by reading only one block, and the current contents of *buffer is ignored
+ * on entry.
*/
bool
StartReadBuffer(ReadBuffersOperation *operation,
@@ -1386,7 +1447,8 @@ StartReadBuffer(ReadBuffersOperation *operation,
int nblocks = 1;
bool result;
- result = StartReadBuffersImpl(operation, buffer, blocknum, &nblocks, flags);
+ result = StartReadBuffersImpl(operation, buffer, blocknum, &nblocks, flags,
+ false /* single block, no forwarding */ );
Assert(nblocks == 1); /* single block can't be short */
return result;
@@ -1416,24 +1478,16 @@ WaitReadBuffers(ReadBuffersOperation *operation)
IOObject io_object;
char persistence;
- /*
- * Currently operations are only allowed to include a read of some range,
- * with an optional extra buffer that is already pinned at the end. So
- * nblocks can be at most one more than io_buffers_len.
- */
- Assert((operation->nblocks == operation->io_buffers_len) ||
- (operation->nblocks == operation->io_buffers_len + 1));
-
/* Find the range of the physical read we need to perform. */
- nblocks = operation->io_buffers_len;
- if (nblocks == 0)
- return; /* nothing to do */
-
+ nblocks = operation->nblocks;
buffers = &operation->buffers[0];
blocknum = operation->blocknum;
forknum = operation->forknum;
persistence = operation->persistence;
+ Assert(nblocks > 0);
+ Assert(nblocks <= MAX_IO_COMBINE_LIMIT);
+
if (persistence == RELPERSISTENCE_TEMP)
{
io_context = IOCONTEXT_NORMAL;
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0009-aio-Basic-subsystem-initialization.patch (16.7K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/10-v2.4-0009-aio-Basic-subsystem-initialization.patch)
download | inline diff:
From fbee247f9ad419ec275a8ebb754a21c0d1cbf4e7 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 11 Feb 2025 14:24:31 -0500
Subject: [PATCH v2.4 09/29] aio: Basic subsystem initialization
This is just separate to make it easier to review the tendrils into various
places.
---
src/include/storage/aio.h | 38 +++++++++++
src/include/storage/aio_subsys.h | 28 ++++++++
src/include/utils/guc.h | 1 +
src/include/utils/guc_hooks.h | 2 +
src/include/utils/resowner.h | 5 ++
src/backend/storage/aio/Makefile | 2 +
src/backend/storage/aio/aio.c | 67 +++++++++++++++++++
src/backend/storage/aio/aio_init.c | 37 ++++++++++
src/backend/storage/aio/meson.build | 2 +
src/backend/storage/ipc/ipci.c | 3 +
src/backend/utils/init/postinit.c | 7 ++
src/backend/utils/misc/guc_tables.c | 23 +++++++
src/backend/utils/misc/postgresql.conf.sample | 6 ++
src/backend/utils/resowner/resowner.c | 29 ++++++++
doc/src/sgml/config.sgml | 51 ++++++++++++++
src/tools/pgindent/typedefs.list | 1 +
16 files changed, 302 insertions(+)
create mode 100644 src/include/storage/aio.h
create mode 100644 src/include/storage/aio_subsys.h
create mode 100644 src/backend/storage/aio/aio.c
create mode 100644 src/backend/storage/aio/aio_init.c
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
new file mode 100644
index 00000000000..e79d5343038
--- /dev/null
+++ b/src/include/storage/aio.h
@@ -0,0 +1,38 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio.h
+ * Main AIO interface
+ *
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/aio.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef AIO_H
+#define AIO_H
+
+
+
+/* Enum for io_method GUC. */
+typedef enum IoMethod
+{
+ IOMETHOD_SYNC = 0,
+} IoMethod;
+
+/* We'll default to synchronous execution. */
+#define DEFAULT_IO_METHOD IOMETHOD_SYNC
+
+
+struct dlist_node;
+extern void pgaio_io_release_resowner(struct dlist_node *ioh_node, bool on_error);
+
+
+/* GUCs */
+extern PGDLLIMPORT int io_method;
+extern PGDLLIMPORT int io_max_concurrency;
+
+
+#endif /* AIO_H */
diff --git a/src/include/storage/aio_subsys.h b/src/include/storage/aio_subsys.h
new file mode 100644
index 00000000000..0b8aaaf00d3
--- /dev/null
+++ b/src/include/storage/aio_subsys.h
@@ -0,0 +1,28 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_subsys.h
+ * Interaction with AIO as a subsystem, rather than actually issuing AIO
+ *
+ * This header is for AIO related functionality that's being called by files
+ * that don't perform AIO but interact with it in some form. E.g. postmaster.c
+ * and shared memory initialization need to initialize AIO but don't perform
+ * AIO.
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/aio_subsys.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef AIO_SUBSYS_H
+#define AIO_SUBSYS_H
+
+
+/* aio_init.c */
+extern Size AioShmemSize(void);
+extern void AioShmemInit(void);
+
+extern void pgaio_init_backend(void);
+
+#endif /* AIO_SUBSYS_H */
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 1233e07d7da..de3bc37264f 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -318,6 +318,7 @@ extern PGDLLIMPORT bool optimize_bounded_sort;
*/
extern PGDLLIMPORT const struct config_enum_entry archive_mode_options[];
extern PGDLLIMPORT const struct config_enum_entry dynamic_shared_memory_options[];
+extern PGDLLIMPORT const struct config_enum_entry io_method_options[];
extern PGDLLIMPORT const struct config_enum_entry recovery_target_action_options[];
extern PGDLLIMPORT const struct config_enum_entry wal_level_options[];
extern PGDLLIMPORT const struct config_enum_entry wal_sync_method_options[];
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 951451a9765..d1b49adc547 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -62,6 +62,8 @@ extern bool check_default_with_oids(bool *newval, void **extra,
extern bool check_effective_io_concurrency(int *newval, void **extra,
GucSource source);
extern bool check_huge_page_size(int *newval, void **extra, GucSource source);
+extern void assign_io_method(int newval, void *extra);
+extern bool check_io_max_concurrency(int *newval, void **extra, GucSource source);
extern const char *show_in_hot_standby(void);
extern bool check_locale_messages(char **newval, void **extra, GucSource source);
extern void assign_locale_messages(const char *newval, void *extra);
diff --git a/src/include/utils/resowner.h b/src/include/utils/resowner.h
index e8d452ca7ee..aede4bfc820 100644
--- a/src/include/utils/resowner.h
+++ b/src/include/utils/resowner.h
@@ -164,4 +164,9 @@ struct LOCALLOCK;
extern void ResourceOwnerRememberLock(ResourceOwner owner, struct LOCALLOCK *locallock);
extern void ResourceOwnerForgetLock(ResourceOwner owner, struct LOCALLOCK *locallock);
+/* special support for AIO */
+struct dlist_node;
+extern void ResourceOwnerRememberAioHandle(ResourceOwner owner, struct dlist_node *ioh_node);
+extern void ResourceOwnerForgetAioHandle(ResourceOwner owner, struct dlist_node *ioh_node);
+
#endif /* RESOWNER_H */
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index 2f29a9ec4d1..eaeaeeee8e3 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -9,6 +9,8 @@ top_builddir = ../../../..
include $(top_builddir)/src/Makefile.global
OBJS = \
+ aio.o \
+ aio_init.o \
read_stream.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
new file mode 100644
index 00000000000..8eb6a5f1292
--- /dev/null
+++ b/src/backend/storage/aio/aio.c
@@ -0,0 +1,67 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio.c
+ * AIO - Core Logic
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "lib/ilist.h"
+#include "storage/aio.h"
+#include "utils/guc.h"
+#include "utils/guc_hooks.h"
+
+
+/* Options for io_method. */
+const struct config_enum_entry io_method_options[] = {
+ {"sync", IOMETHOD_SYNC, false},
+ {NULL, 0, false}
+};
+
+/* GUCs */
+int io_method = DEFAULT_IO_METHOD;
+int io_max_concurrency = -1;
+
+
+
+void
+assign_io_method(int newval, void *extra)
+{
+}
+
+bool
+check_io_max_concurrency(int *newval, void **extra, GucSource source)
+{
+ if (*newval == -1)
+ {
+ /*
+ * Auto-tuning will be applied later during startup, as auto-tuning
+ * depends on the value of various GUCs.
+ */
+ return true;
+ }
+ else if (*newval == 0)
+ {
+ GUC_check_errdetail("Only -1 or values bigger than 0 are valid.");
+ return false;
+ }
+
+ return true;
+}
+
+/*
+ * Release IO handle during resource owner cleanup.
+ */
+void
+pgaio_io_release_resowner(dlist_node *ioh_node, bool on_error)
+{
+ /* placeholder for later */
+}
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
new file mode 100644
index 00000000000..aeacc144149
--- /dev/null
+++ b/src/backend/storage/aio/aio_init.c
@@ -0,0 +1,37 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_init.c
+ * AIO - Subsystem Initialization
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio_init.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio_subsys.h"
+
+
+
+Size
+AioShmemSize(void)
+{
+ Size sz = 0;
+
+ return sz;
+}
+
+void
+AioShmemInit(void)
+{
+}
+
+void
+pgaio_init_backend(void)
+{
+}
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 8abe0eb4863..c822fd4ddf7 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -1,5 +1,7 @@
# Copyright (c) 2024-2025, PostgreSQL Global Development Group
backend_sources += files(
+ 'aio.c',
+ 'aio_init.c',
'read_stream.c',
)
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 174eed70367..2fa045e6b0f 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -37,6 +37,7 @@
#include "replication/slotsync.h"
#include "replication/walreceiver.h"
#include "replication/walsender.h"
+#include "storage/aio_subsys.h"
#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/dsm_registry.h"
@@ -148,6 +149,7 @@ CalculateShmemSize(int *num_semaphores)
size = add_size(size, WaitEventCustomShmemSize());
size = add_size(size, InjectionPointShmemSize());
size = add_size(size, SlotSyncShmemSize());
+ size = add_size(size, AioShmemSize());
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -340,6 +342,7 @@ CreateOrAttachShmemStructs(void)
StatsShmemInit();
WaitEventCustomShmemInit();
InjectionPointShmemInit();
+ AioShmemInit();
}
/*
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index b491d04de58..4410bbe602b 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -43,6 +43,7 @@
#include "replication/slot.h"
#include "replication/slotsync.h"
#include "replication/walsender.h"
+#include "storage/aio_subsys.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
#include "storage/ipc.h"
@@ -626,6 +627,12 @@ BaseInit(void)
*/
pgstat_initialize();
+ /*
+ * Initialize AIO before infrastructure that might need to actually
+ * execute AIO.
+ */
+ pgaio_init_backend();
+
/* Do local initialization of storage and buffer managers */
InitSync();
smgrinit();
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 3cde94a1759..b7c84d061e2 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -71,6 +71,7 @@
#include "replication/slot.h"
#include "replication/slotsync.h"
#include "replication/syncrep.h"
+#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/bufpage.h"
#include "storage/large_object.h"
@@ -3252,6 +3253,18 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"io_max_concurrency",
+ PGC_POSTMASTER,
+ RESOURCES_IO,
+ gettext_noop("Max number of IOs that one process can execute simultaneously."),
+ NULL,
+ },
+ &io_max_concurrency,
+ -1, -1, 1024,
+ check_io_max_concurrency, NULL, NULL
+ },
+
{
{"backend_flush_after", PGC_USERSET, RESOURCES_IO,
gettext_noop("Number of pages after which previously performed writes are flushed to disk."),
@@ -5286,6 +5299,16 @@ struct config_enum ConfigureNamesEnum[] =
NULL, NULL, NULL
},
+ {
+ {"io_method", PGC_POSTMASTER, RESOURCES_IO,
+ gettext_noop("Selects the method for executing asynchronous I/O."),
+ NULL
+ },
+ &io_method,
+ DEFAULT_IO_METHOD, io_method_options,
+ NULL, assign_io_method, NULL
+ },
+
/* End-of-list marker */
{
{NULL, 0, 0, NULL, NULL}, NULL, 0, NULL, NULL, NULL, NULL
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index 415f253096c..186bc47b700 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -199,6 +199,12 @@
#maintenance_io_concurrency = 10 # 1-1000; 0 disables prefetching
#io_combine_limit = 128kB # usually 1-32 blocks (depends on OS)
+#io_method = sync # sync (change requires restart)
+#io_max_concurrency = -1 # Max number of IOs that one process
+ # can execute simultaneously
+ # -1 sets based on shared_buffers
+ # (change requires restart)
+
# - Worker Processes -
#max_worker_processes = 8 # (change requires restart)
diff --git a/src/backend/utils/resowner/resowner.c b/src/backend/utils/resowner/resowner.c
index ac5ca4a765e..76b9cec1e26 100644
--- a/src/backend/utils/resowner/resowner.c
+++ b/src/backend/utils/resowner/resowner.c
@@ -47,6 +47,8 @@
#include "common/hashfn.h"
#include "common/int.h"
+#include "lib/ilist.h"
+#include "storage/aio.h"
#include "storage/ipc.h"
#include "storage/predicate.h"
#include "storage/proc.h"
@@ -155,6 +157,12 @@ struct ResourceOwnerData
/* The local locks cache. */
LOCALLOCK *locks[MAX_RESOWNER_LOCKS]; /* list of owned locks */
+
+ /*
+ * AIO handles need be registered in critical sections and therefore
+ * cannot use the normal ResoureElem mechanism.
+ */
+ dlist_head aio_handles;
};
@@ -425,6 +433,8 @@ ResourceOwnerCreate(ResourceOwner parent, const char *name)
parent->firstchild = owner;
}
+ dlist_init(&owner->aio_handles);
+
return owner;
}
@@ -725,6 +735,13 @@ ResourceOwnerReleaseInternal(ResourceOwner owner,
* so issue warnings. In the abort case, just clean up quietly.
*/
ResourceOwnerReleaseAll(owner, phase, isCommit);
+
+ while (!dlist_is_empty(&owner->aio_handles))
+ {
+ dlist_node *node = dlist_head_node(&owner->aio_handles);
+
+ pgaio_io_release_resowner(node, !isCommit);
+ }
}
else if (phase == RESOURCE_RELEASE_LOCKS)
{
@@ -1082,3 +1099,15 @@ ResourceOwnerForgetLock(ResourceOwner owner, LOCALLOCK *locallock)
elog(ERROR, "lock reference %p is not owned by resource owner %s",
locallock, owner->name);
}
+
+void
+ResourceOwnerRememberAioHandle(ResourceOwner owner, struct dlist_node *ioh_node)
+{
+ dlist_push_tail(&owner->aio_handles, ioh_node);
+}
+
+void
+ResourceOwnerForgetAioHandle(ResourceOwner owner, struct dlist_node *ioh_node)
+{
+ dlist_delete_from(&owner->aio_handles, ioh_node);
+}
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 9eedcf6f0f4..0306827afbd 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -2618,6 +2618,57 @@ include_dir 'conf.d'
</para>
</listitem>
</varlistentry>
+
+ <varlistentry id="guc-io-max-concurrency" xreflabel="io_max_concurrency">
+ <term><varname>io_max_concurrency</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>io_max_concurrency</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Controls the maximum number of I/O operations that one process can
+ execute simultaneously.
+ </para>
+ <para>
+ The default setting of <literal>-1</literal> selects a number based
+ on <xref linkend="guc-shared-buffers"/> and the maximum number of
+ processes (<xref linkend="guc-max-connections"/>, <xref
+ linkend="guc-autovacuum-worker-slots"/>, <xref
+ linkend="guc-max-worker-processes"/> and <xref
+ linkend="guc-max-wal-senders"/>), but not more than
+ <literal>64</literal>.
+ </para>
+ <para>
+ This parameter can only be set at server start.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry id="guc-io-method" xreflabel="io_method">
+ <term><varname>io_method</varname> (<type>enum</type>)
+ <indexterm>
+ <primary><varname>io_method</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Selects the method for executing asynchronous I/O.
+ Possible values are:
+ <itemizedlist>
+ <listitem>
+ <para>
+ <literal>sync</literal> (execute asynchronous I/O synchronously)
+ </para>
+ </listitem>
+ </itemizedlist>
+ </para>
+ <para>
+ This parameter can only be set at server start.
+ </para>
+ </listitem>
+ </varlistentry>
+
</variablelist>
</sect2>
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index fb39c915d76..406c0893440 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1267,6 +1267,7 @@ IntoClause
InvalMessageArray
InvalidationInfo
InvalidationMsgsGroup
+IoMethod
IpcMemoryId
IpcMemoryKey
IpcMemoryState
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0010-aio-Core-AIO-implementation.patch (88.5K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/11-v2.4-0010-aio-Core-AIO-implementation.patch)
download | inline diff:
From fa8a603bd1f8017bae0edcdac28b12b81ab10671 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 11 Feb 2025 14:57:59 -0500
Subject: [PATCH v2.4 10/29] aio: Core AIO implementation
This commit is not sufficient to actually perform AIO, it just contains the
infrastructure for using AIO. Subsequent commits will introduce different
methods of executing AIO and support for performing AIO on different targets.
---
src/include/storage/aio.h | 299 +++++
src/include/storage/aio_internal.h | 348 +++++
src/include/storage/aio_subsys.h | 5 +
src/include/storage/aio_types.h | 115 ++
src/backend/access/transam/xact.c | 12 +
src/backend/postmaster/autovacuum.c | 2 +
src/backend/postmaster/bgwriter.c | 2 +
src/backend/postmaster/checkpointer.c | 2 +
src/backend/postmaster/pgarch.c | 2 +
src/backend/postmaster/walsummarizer.c | 2 +
src/backend/postmaster/walwriter.c | 2 +
src/backend/replication/walsender.c | 2 +
src/backend/storage/aio/Makefile | 4 +
src/backend/storage/aio/aio.c | 1121 ++++++++++++++++-
src/backend/storage/aio/aio_callback.c | 288 +++++
src/backend/storage/aio/aio_init.c | 186 +++
src/backend/storage/aio/aio_io.c | 180 +++
src/backend/storage/aio/aio_target.c | 108 ++
src/backend/storage/aio/meson.build | 4 +
src/backend/storage/aio/method_sync.c | 47 +
.../utils/activity/wait_event_names.txt | 1 +
src/tools/pgindent/typedefs.list | 21 +
22 files changed, 2749 insertions(+), 4 deletions(-)
create mode 100644 src/include/storage/aio_internal.h
create mode 100644 src/include/storage/aio_types.h
create mode 100644 src/backend/storage/aio/aio_callback.c
create mode 100644 src/backend/storage/aio/aio_io.c
create mode 100644 src/backend/storage/aio/aio_target.c
create mode 100644 src/backend/storage/aio/method_sync.c
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index e79d5343038..d87cfe96b20 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -14,6 +14,9 @@
#ifndef AIO_H
#define AIO_H
+#include "storage/aio_types.h"
+#include "storage/procnumber.h"
+
/* Enum for io_method GUC. */
@@ -26,9 +29,305 @@ typedef enum IoMethod
#define DEFAULT_IO_METHOD IOMETHOD_SYNC
+/*
+ * Flags for an IO that can be set with pgaio_io_set_flag().
+ */
+typedef enum PgAioHandleFlags
+{
+ /*
+ * The IO references backend local memory.
+ *
+ * This needs to be set on an IO whenever the IO references process-local
+ * memory. Some IO methods do not support executing IO that references
+ * process local memory and thus need to fall back to executing IO
+ * synchronously for IOs with this flag set.
+ *
+ * Required for correctness.
+ */
+ PGAIO_HF_REFERENCES_LOCAL = 1 << 1,
+
+ /*
+ * Hint that IO will be executed synchronously.
+ *
+ * This can make it a bit cheaper to execute synchronous IO via the AIO
+ * interface, to avoid needing an AIO and non-AIO version of code.
+ *
+ * Advantageous to set, if applicable, but not required for correctness.
+ */
+ PGAIO_HF_SYNCHRONOUS = 1 << 0,
+
+ /*
+ * IO is using buffered IO, used to control heuristic in some IO methods.
+ *
+ * Advantageous to set, if applicable, but not required for correctness.
+ */
+ PGAIO_HF_BUFFERED = 1 << 2,
+} PgAioHandleFlags;
+
+/*
+ * The IO operations supported by the AIO subsystem.
+ *
+ * This could be in aio_internal.h, as it is not publically referenced, but
+ * PgAioOpData currently *does* need to be public, therefore keeping this
+ * public seems to make sense.
+ */
+typedef enum PgAioOp
+{
+ /* intentionally the zero value, to help catch zeroed memory etc */
+ PGAIO_OP_INVALID = 0,
+
+ PGAIO_OP_READV,
+ PGAIO_OP_WRITEV,
+
+ /**
+ * In the near term we'll need at least:
+ * - fsync / fdatasync
+ * - flush_range
+ *
+ * Eventually we'll additionally want at least:
+ * - send
+ * - recv
+ * - accept
+ **/
+} PgAioOp;
+
+#define PGAIO_OP_COUNT (PGAIO_OP_WRITEV + 1)
+
+
+/*
+ * On what is IO being performed.
+ *
+ * PgAioTargetID specific behaviour should be implemented in
+ * aio_target.c.
+ */
+typedef enum PgAioTargetID
+{
+ /* intentionally the zero value, to help catch zeroed memory etc */
+ PGAIO_TID_INVALID = 0,
+} PgAioTargetID;
+
+#define PGAIO_TID_COUNT (PGAIO_TID_INVALID + 1)
+
+
+/*
+ * Data necessary for support IO operations (see PgAioOp).
+ *
+ * NB: Note that the FDs in here may *not* be relied upon for re-issuing
+ * requests (e.g. for partial reads/writes) - the FD might be from another
+ * process, or closed since. That's not a problem for IOs waiting to be issued
+ * only because the queue is flushed when closing an FD.
+ */
+typedef union
+{
+ struct
+ {
+ int fd;
+ uint16 iov_length;
+ uint64 offset;
+ } read;
+
+ struct
+ {
+ int fd;
+ uint16 iov_length;
+ uint64 offset;
+ } write;
+} PgAioOpData;
+
+
+/*
+ * Information the object that IO is executed on. Mostly callbacks that
+ * operate on PgAioTargetData.
+ */
+typedef struct PgAioTargetInfo
+{
+ void (*reopen) (PgAioHandle *ioh);
+
+ char *(*describe_identity) (const PgAioTargetData *sd);
+
+ const char *name;
+} PgAioTargetInfo;
+
+
+/*
+ * IDs for callbacks that can be registered on an IO.
+ *
+ * Callbacks are identified by an ID rather than a function pointer. There are
+ * two main reasons:
+ *
+ * 1) Memory within PgAioHandle is precious, due to the number of PgAioHandle
+ * structs in pre-allocated shared memory.
+ *
+ * 2) Due to EXEC_BACKEND function pointers are not necessarily stable between
+ * different backends, therefore function pointers cannot directly be in
+ * shared memory.
+ *
+ * Without 2), we could fairly easily allow to add new callbacks, by filling a
+ * ID->pointer mapping table on demand. In the presence of 2 that's still
+ * doable, but harder, because every process has to re-register the pointers
+ * so that a local ID->"backend local pointer" mapping can be maintained.
+ */
+typedef enum PgAioHandleCallbackID
+{
+ PGAIO_HCB_INVALID,
+} PgAioHandleCallbackID;
+
+
+typedef void (*PgAioHandleCallbackStage) (PgAioHandle *ioh);
+typedef PgAioResult (*PgAioHandleCallbackComplete) (PgAioHandle *ioh, PgAioResult prior_result);
+typedef void (*PgAioHandleCallbackReport) (PgAioResult result, const PgAioTargetData *target_data, int elevel);
+
+typedef struct PgAioHandleCallbacks
+{
+ /*
+ * Prepare resources affected by the IO for execution. This could e.g.
+ * include moving ownership of buffer pins to the AIO subsystem.
+ */
+ PgAioHandleCallbackStage stage;
+
+ /*
+ * Update the state of resources affected by the IO to reflect completion
+ * of the IO. This could e.g. include updating shared buffer state to
+ * signal the IO has finished.
+ *
+ * The _shared suffix indicates that this is executed by the backend that
+ * completed the IO, which may or may not be the backend that issued the
+ * IO. Obviously the callback thus can only modify resources in shared
+ * memory.
+ *
+ * The latest registered callback is called first. This allows
+ * higher-level code to register callbacks that can rely on callbacks
+ * registered by lower-level code to already have been executed.
+ *
+ * NB: This is called in a critical section. Errors can be signalled by
+ * the callback's return value, it's the responsibility of the IO's issuer
+ * to react appropriately.
+ */
+ PgAioHandleCallbackComplete complete_shared;
+
+ /*
+ * Like complete_shared, except called in the issuing backend.
+ *
+ * This variant of the completion callback is useful when backend-local
+ * state has to be updated to reflect the IO's completion. E.g. a
+ * temporary buffer's BufferDesc isn't accessible in complete_shared.
+ *
+ * Local callbacks are only called after complete_shared for all
+ * registered callbacks has been called.
+ */
+ PgAioHandleCallbackComplete complete_local;
+
+ /*
+ * Report the result of an IO operation. This is e.g. used to raise an
+ * error after an IO failed at the appropriate time (i.e. not when the IO
+ * failed, but under control of the code that issued the IO).
+ */
+ PgAioHandleCallbackReport report;
+} PgAioHandleCallbacks;
+
+
+
+/*
+ * How many callbacks can be registered for one IO handle. Currently we only
+ * need two, but it's not hard to imagine needing a few more.
+ */
+#define PGAIO_HANDLE_MAX_CALLBACKS 4
+
+
+
+/* AIO API */
+
+
+/* --------------------------------------------------------------------------------
+ * IO Handles
+ * --------------------------------------------------------------------------------
+ */
+
+/* functions in aio.c */
+struct ResourceOwnerData;
+extern PgAioHandle *pgaio_io_acquire(struct ResourceOwnerData *resowner, PgAioReturn *ret);
+extern PgAioHandle *pgaio_io_acquire_nb(struct ResourceOwnerData *resowner, PgAioReturn *ret);
+
+extern void pgaio_io_release(PgAioHandle *ioh);
struct dlist_node;
extern void pgaio_io_release_resowner(struct dlist_node *ioh_node, bool on_error);
+extern void pgaio_io_set_flag(PgAioHandle *ioh, PgAioHandleFlags flag);
+
+extern int pgaio_io_get_id(PgAioHandle *ioh);
+extern ProcNumber pgaio_io_get_owner(PgAioHandle *ioh);
+
+extern void pgaio_io_get_wref(PgAioHandle *ioh, PgAioWaitRef *iow);
+
+/* functions in aio_io.c */
+struct iovec;
+extern int pgaio_io_get_iovec(PgAioHandle *ioh, struct iovec **iov);
+
+extern PgAioOpData *pgaio_io_get_op_data(PgAioHandle *ioh);
+
+extern void pgaio_io_prep_readv(PgAioHandle *ioh,
+ int fd, int iovcnt, uint64 offset);
+extern void pgaio_io_prep_writev(PgAioHandle *ioh,
+ int fd, int iovcnt, uint64 offset);
+
+/* functions in aio_target.c */
+extern void pgaio_io_set_target(PgAioHandle *ioh, PgAioTargetID targetid);
+extern bool pgaio_io_has_target(PgAioHandle *ioh);
+extern PgAioTargetData *pgaio_io_get_target_data(PgAioHandle *ioh);
+extern char *pgaio_io_get_target_description(PgAioHandle *ioh);
+
+/* functions in aio_callback.c */
+extern void pgaio_io_register_callbacks(PgAioHandle *ioh, PgAioHandleCallbackID cbid);
+extern void pgaio_io_set_handle_data_64(PgAioHandle *ioh, uint64 *data, uint8 len);
+extern void pgaio_io_set_handle_data_32(PgAioHandle *ioh, uint32 *data, uint8 len);
+extern uint64 *pgaio_io_get_handle_data(PgAioHandle *ioh, uint8 *len);
+
+
+
+/* --------------------------------------------------------------------------------
+ * IO Wait References
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_wref_clear(PgAioWaitRef *iow);
+extern bool pgaio_wref_valid(PgAioWaitRef *iow);
+extern int pgaio_wref_get_id(PgAioWaitRef *iow);
+
+extern void pgaio_wref_wait(PgAioWaitRef *iow);
+extern bool pgaio_wref_check_done(PgAioWaitRef *iow);
+
+
+
+/* --------------------------------------------------------------------------------
+ * IO Result
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_result_report(PgAioResult result, const PgAioTargetData *target_data,
+ int elevel);
+
+
+
+/* --------------------------------------------------------------------------------
+ * Actions on multiple IOs.
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_enter_batchmode(void);
+extern void pgaio_exit_batchmode(void);
+extern void pgaio_submit_staged(void);
+extern bool pgaio_have_staged(void);
+
+
+
+/* --------------------------------------------------------------------------------
+ * Other
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_closing_fd(int fd);
+
+
/* GUCs */
extern PGDLLIMPORT int io_method;
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
new file mode 100644
index 00000000000..e980b06c1f3
--- /dev/null
+++ b/src/include/storage/aio_internal.h
@@ -0,0 +1,348 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_internal.h
+ * AIO related declarations that shoul only be used by the AIO subsystem
+ * internally.
+ *
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/aio_internal.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef AIO_INTERNAL_H
+#define AIO_INTERNAL_H
+
+
+#include "lib/ilist.h"
+#include "port/pg_iovec.h"
+#include "storage/aio.h"
+#include "storage/condition_variable.h"
+
+
+/*
+ * The maximum number of IOs that can be batch submitted at once.
+ */
+#define PGAIO_SUBMIT_BATCH_SIZE 32
+
+
+
+typedef enum PgAioHandleState
+{
+ /* not in use */
+ PGAIO_HS_IDLE = 0,
+
+ /* returned by pgaio_io_acquire() */
+ PGAIO_HS_HANDED_OUT,
+
+ /* pgaio_io_prep_*() has been called, but IO hasn't been submitted yet */
+ PGAIO_HS_DEFINED,
+
+ /* target's stage() callback has been called, ready to be submitted */
+ PGAIO_HS_STAGED,
+
+ /* IO has been submitted and is being executed */
+ PGAIO_HS_SUBMITTED,
+
+ /* IO finished, but result has not yet been processed */
+ PGAIO_HS_COMPLETED_IO,
+
+ /* IO completed, shared completion has been called */
+ PGAIO_HS_COMPLETED_SHARED,
+
+ /* IO completed, local completion has been called */
+ PGAIO_HS_COMPLETED_LOCAL,
+} PgAioHandleState;
+
+
+struct ResourceOwnerData;
+
+/* typedef is in public header */
+struct PgAioHandle
+{
+ /* all state updates should go through pgaio_io_update_state() */
+ PgAioHandleState state:8;
+
+ /* what are we operating on */
+ PgAioTargetID target:8;
+
+ /* which IO operation */
+ PgAioOp op:8;
+
+ /* bitfield of PgAioHandleFlags */
+ uint8 flags;
+
+ uint8 num_shared_callbacks;
+
+ /* using the proper type here would use more space */
+ uint8 shared_callbacks[PGAIO_HANDLE_MAX_CALLBACKS];
+
+ /*
+ * Length of data associated with handle using
+ * pgaio_io_set_handle_data_*().
+ */
+ uint8 handle_data_len;
+
+ /* XXX: could be optimized out with some pointer math */
+ int32 owner_procno;
+
+ /* raw result of the IO operation */
+ int32 result;
+
+ /*
+ * Index into PgAioCtl->iovecs and PgAioCtl->handle_data.
+ *
+ * At the moment there's no need to differentiate between the two, but
+ * that won't necessarily stay that way.
+ */
+ uint32 iovec_off;
+
+ /**
+ * In which list the handle is registered, depends on the state:
+ * - IDLE, in per-backend list
+ * - HANDED_OUT - not in a list
+ * - DEFINED - in per-backend staged list
+ * - STAGED - in per-backend staged list
+ * - SUBMITTED - in issuer's in_flight list
+ * - COMPLETED_IO - in issuer's in_flight list
+ * - COMPLETED_SHARED - in issuer's in_flight list
+ **/
+ dlist_node node;
+
+ struct ResourceOwnerData *resowner;
+ dlist_node resowner_node;
+
+ /* incremented every time the IO handle is reused */
+ uint64 generation;
+
+ ConditionVariable cv;
+
+ /* result of shared callback, passed to issuer callback */
+ PgAioResult distilled_result;
+
+ PgAioReturn *report_return;
+
+ PgAioOpData op_data;
+
+ /*
+ * Data necessary to identify the object undergoing IO to higher-level
+ * code. Needs to be sufficient to allow another backend to reopen the
+ * file.
+ */
+ PgAioTargetData target_data;
+};
+
+
+typedef struct PgAioBackend
+{
+ /* index into PgAioCtl->io_handles */
+ uint32 io_handle_off;
+
+ /* IO Handles that currently are not used */
+ dclist_head idle_ios;
+
+ /*
+ * Only one IO may be returned by pgaio_io_acquire()/pgaio_io_acquire()
+ * without having been either defined (by actually associating it with IO)
+ * or by released (with pgaio_io_release()). This restriction is necessary
+ * to guarantee that we always can acquire an IO. ->handed_out_io is used
+ * to enforce that rule.
+ */
+ PgAioHandle *handed_out_io;
+
+ /* Are we currently in batchmode? See pgaio_enter_batchmode(). */
+ bool in_batchmode;
+
+ /*
+ * IOs that are defined, but not yet submitted.
+ */
+ uint16 num_staged_ios;
+ PgAioHandle *staged_ios[PGAIO_SUBMIT_BATCH_SIZE];
+
+ /*
+ * List of in-flight IOs. Also contains IOs that aren't strict speaking
+ * in-flight anymore, but have been waited-for and completed by another
+ * backend. Once this backend sees such an IO it'll be reclaimed.
+ *
+ * The list is ordered by submission time, with more recently submitted
+ * IOs being appended at the end.
+ */
+ dclist_head in_flight_ios;
+} PgAioBackend;
+
+
+typedef struct PgAioCtl
+{
+ int backend_state_count;
+ PgAioBackend *backend_state;
+
+ /*
+ * Array of iovec structs. Each iovec is owned by a specific backend. The
+ * allocation is in PgAioCtl to allow the maximum number of iovecs for
+ * individual IOs to be configurable with PGC_POSTMASTER GUC.
+ */
+ uint64 iovec_count;
+ struct iovec *iovecs;
+
+ /*
+ * For, e.g., an IO covering multiple buffers in shared / temp buffers, we
+ * need to get Buffer IDs during completion to be able to change the
+ * BufferDesc state accordingly. This space can be used to store e.g.
+ * Buffer IDs. Note that the actual iovec might be shorter than this,
+ * because we combine neighboring pages into one larger iovec entry.
+ */
+ uint64 *handle_data;
+
+ uint64 io_handle_count;
+ PgAioHandle *io_handles;
+} PgAioCtl;
+
+
+
+/*
+ * Callbacks used to implement an IO method.
+ */
+typedef struct IoMethodOps
+{
+ /* global initialization */
+
+ /*
+ * Amount of additional shared memory to reserve for the io_method. Called
+ * just like a normal ipci.c style *Size() function. Optional.
+ */
+ size_t (*shmem_size) (void);
+
+ /*
+ * Initialize shared memory. First time is true if AIO's shared memory was
+ * just initialized, false otherwise. Optional.
+ */
+ void (*shmem_init) (bool first_time);
+
+ /*
+ * Per-backend initialization. Optional.
+ */
+ void (*init_backend) (void);
+
+
+ /* handling of IOs */
+
+ /* optional */
+ bool (*needs_synchronous_execution) (PgAioHandle *ioh);
+
+ /*
+ * Start executing passed in IOs.
+ *
+ * Will not be called if ->needs_synchronous_execution() returned true.
+ *
+ * num_staged_ios is <= PGAIO_SUBMIT_BATCH_SIZE.
+ *
+ */
+ int (*submit) (uint16 num_staged_ios, PgAioHandle **staged_ios);
+
+ /*
+ * Wait for the IO to complete. Optional.
+ *
+ * If not provided, it needs to be guaranteed that the IO method calls
+ * pgaio_io_process_completion() without further interaction by the
+ * issuing backend.
+ */
+ void (*wait_one) (PgAioHandle *ioh,
+ uint64 ref_generation);
+} IoMethodOps;
+
+
+/* aio.c */
+extern bool pgaio_io_was_recycled(PgAioHandle *ioh, uint64 ref_generation, PgAioHandleState *state);
+extern void pgaio_io_stage(PgAioHandle *ioh, PgAioOp op);
+extern void pgaio_io_process_completion(PgAioHandle *ioh, int result);
+extern void pgaio_io_prepare_submit(PgAioHandle *ioh);
+extern bool pgaio_io_needs_synchronous_execution(PgAioHandle *ioh);
+extern const char *pgaio_io_get_state_name(PgAioHandle *ioh);
+const char *pgaio_result_status_string(PgAioResultStatus rs);
+extern void pgaio_shutdown(int code, Datum arg);
+
+/* aio_callback.c */
+extern void pgaio_io_call_stage(PgAioHandle *ioh);
+extern void pgaio_io_call_complete_shared(PgAioHandle *ioh);
+extern void pgaio_io_call_complete_local(PgAioHandle *ioh);
+
+/* aio_io.c */
+extern void pgaio_io_perform_synchronously(PgAioHandle *ioh);
+extern const char *pgaio_io_get_op_name(PgAioHandle *ioh);
+
+/* aio_target.c */
+extern bool pgaio_io_can_reopen(PgAioHandle *ioh);
+extern void pgaio_io_reopen(PgAioHandle *ioh);
+extern const char *pgaio_io_get_target_name(PgAioHandle *ioh);
+
+
+/*
+ * The AIO subsystem has fairly verbose debug logging support. This can be
+ * enabled/disabled at buildtime. The reason for this is that
+ * a) the verbosity can make debugging things on higher levels hard
+ * b) even if logging can be skipped due to elevel checks, it still causes a
+ * measurable slowdown
+ */
+#define PGAIO_VERBOSE 1
+
+/*
+ * Simple ereport() wrapper that only logs if PGAIO_VERBOSE is defined.
+ *
+ * This intentionally still compiles the code, guarded by a constant if (0),
+ * if verbose logging is disabled, to make it less likely that debug logging
+ * is silently broken.
+ *
+ * The current definition requires passing at least one argument.
+ */
+#define pgaio_debug(elevel, msg, ...) \
+ do { \
+ if (PGAIO_VERBOSE) \
+ ereport(elevel, \
+ errhidestmt(true), errhidecontext(true), \
+ errmsg_internal(msg, \
+ __VA_ARGS__)); \
+ } while(0)
+
+/*
+ * Simple ereport() wrapper. Note that the definition requires passing at
+ * least one argument.
+ */
+#define pgaio_debug_io(elevel, ioh, msg, ...) \
+ pgaio_debug(elevel, "io %-10d|op %-5s|target %-4s|state %-16s: " msg, \
+ pgaio_io_get_id(ioh), \
+ pgaio_io_get_op_name(ioh), \
+ pgaio_io_get_target_name(ioh), \
+ pgaio_io_get_state_name(ioh), \
+ __VA_ARGS__)
+
+
+#ifdef USE_INJECTION_POINTS
+
+extern void pgaio_io_call_inj(PgAioHandle *ioh, const char *injection_point);
+
+/* just for use in tests, from within injection points */
+extern PgAioHandle *pgaio_inj_io_get(void);
+
+#else
+
+#define pgaio_io_call_inj(ioh, injection_point) (void) 0
+/*
+ * no fallback for pgaio_inj_io_get, all code using injection points better be
+ * guarded by USE_INJECTION_POINTS.
+ */
+
+#endif
+
+
+/* Declarations for the tables of function pointers exposed by each IO method. */
+extern PGDLLIMPORT const IoMethodOps pgaio_sync_ops;
+
+extern PGDLLIMPORT const IoMethodOps *pgaio_method_ops;
+extern PGDLLIMPORT PgAioCtl *pgaio_ctl;
+extern PGDLLIMPORT PgAioBackend *pgaio_my_backend;
+
+
+
+#endif /* AIO_INTERNAL_H */
diff --git a/src/include/storage/aio_subsys.h b/src/include/storage/aio_subsys.h
index 0b8aaaf00d3..e4faf692a38 100644
--- a/src/include/storage/aio_subsys.h
+++ b/src/include/storage/aio_subsys.h
@@ -25,4 +25,9 @@ extern void AioShmemInit(void);
extern void pgaio_init_backend(void);
+
+/* aio.c */
+extern void pgaio_error_cleanup(void);
+extern void AtEOXact_Aio(bool is_commit);
+
#endif /* AIO_SUBSYS_H */
diff --git a/src/include/storage/aio_types.h b/src/include/storage/aio_types.h
new file mode 100644
index 00000000000..d2617139a25
--- /dev/null
+++ b/src/include/storage/aio_types.h
@@ -0,0 +1,115 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_types.h
+ * AIO related types that are useful to include separately, to reduce the
+ * "include burden".
+ *
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/aio_types.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef AIO_TYPES_H
+#define AIO_TYPES_H
+
+#include "storage/block.h"
+#include "storage/relfilelocator.h"
+
+
+typedef struct PgAioHandle PgAioHandle;
+
+/*
+ * A reference to an IO that can be used to wait for the IO (using
+ * pgaio_wref_wait()) to complete.
+ *
+ * These can be passed across process boundaries.
+ */
+typedef struct PgAioWaitRef
+{
+ /* internal ID identifying the specific PgAioHandle */
+ uint32 aio_index;
+
+ /*
+ * IO handles are reused. To detect if a handle was reused, and thereby
+ * avoid unnecessarily waiting for a newer IO, each time the handle is
+ * reused a generation number is increased.
+ *
+ * To avoid requiring alignment sufficient for an int64, split the
+ * generation into two.
+ */
+ uint32 generation_upper;
+ uint32 generation_lower;
+} PgAioWaitRef;
+
+
+/*
+ * Information identifying what the IO is being performed on.
+ *
+ * This needs sufficient information to
+ *
+ * a) Reopen the file for the IO if the IO is executed in a context that
+ * cannot use the FD provided initially (e.g. because the IO is executed in
+ * a worker process).
+ *
+ * b) Describe the object the IO is performed on in log / error messages.
+ */
+typedef union PgAioTargetData
+{
+ /* just as an example placeholder for later */
+ struct
+ {
+ uint32 queue_id;
+ } wal;
+} PgAioTargetData;
+
+
+/*
+ * The status of an AIO operation.
+ */
+typedef enum PgAioResultStatus
+{
+ ARS_UNKNOWN, /* not yet completed / uninitialized */
+ ARS_OK,
+ ARS_PARTIAL, /* did not fully succeed, but no error */
+ ARS_ERROR,
+} PgAioResultStatus;
+
+
+/*
+ * Result of IO operation, visible only to the initiator of IO.
+ */
+typedef struct PgAioResult
+{
+ /*
+ * This is of type PgAioHandleCallbackID, but can't use a bitfield of an
+ * enum, because some compilers treat enums as signed.
+ */
+ uint32 id:8;
+
+ /* of type PgAioResultStatus, see above */
+ uint32 status:2;
+
+ /* meaning defined by callback->error */
+ uint32 error_data:22;
+
+ int32 result;
+} PgAioResult;
+
+
+/*
+ * Combination of PgAioResult with minimal metadata about the IO.
+ *
+ * Contains sufficient information to be able, in case the IO [partially]
+ * fails, to log/raise an error under control of the IO issuing code.
+ */
+typedef struct PgAioReturn
+{
+ PgAioResult result;
+ PgAioTargetData target_data;
+} PgAioReturn;
+
+
+#endif /* AIO_TYPES_H */
diff --git a/src/backend/access/transam/xact.c b/src/backend/access/transam/xact.c
index 1b4f21a88d3..b885513f765 100644
--- a/src/backend/access/transam/xact.c
+++ b/src/backend/access/transam/xact.c
@@ -51,6 +51,7 @@
#include "replication/origin.h"
#include "replication/snapbuild.h"
#include "replication/syncrep.h"
+#include "storage/aio_subsys.h"
#include "storage/condition_variable.h"
#include "storage/fd.h"
#include "storage/lmgr.h"
@@ -2411,6 +2412,8 @@ CommitTransaction(void)
RESOURCE_RELEASE_BEFORE_LOCKS,
true, true);
+ AtEOXact_Aio(true);
+
/* Check we've released all buffer pins */
AtEOXact_Buffers(true);
@@ -2716,6 +2719,8 @@ PrepareTransaction(void)
RESOURCE_RELEASE_BEFORE_LOCKS,
true, true);
+ AtEOXact_Aio(true);
+
/* Check we've released all buffer pins */
AtEOXact_Buffers(true);
@@ -2830,6 +2835,8 @@ AbortTransaction(void)
pgstat_report_wait_end();
pgstat_progress_end_command();
+ pgaio_error_cleanup();
+
/* Clean up buffer content locks, too */
UnlockBuffers();
@@ -2960,6 +2967,7 @@ AbortTransaction(void)
ResourceOwnerRelease(TopTransactionResourceOwner,
RESOURCE_RELEASE_BEFORE_LOCKS,
false, true);
+ AtEOXact_Aio(false);
AtEOXact_Buffers(false);
AtEOXact_RelationCache(false);
AtEOXact_TypeCache();
@@ -5232,6 +5240,9 @@ AbortSubTransaction(void)
pgstat_report_wait_end();
pgstat_progress_end_command();
+
+ pgaio_error_cleanup();
+
UnlockBuffers();
/* Reset WAL record construction state */
@@ -5326,6 +5337,7 @@ AbortSubTransaction(void)
RESOURCE_RELEASE_BEFORE_LOCKS,
false, false);
+ AtEOXact_Aio(false);
AtEOSubXact_RelationCache(false, s->subTransactionId,
s->parent->subTransactionId);
AtEOSubXact_TypeCache();
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index ade2708b59e..88effd259f7 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -88,6 +88,7 @@
#include "postmaster/autovacuum.h"
#include "postmaster/interrupt.h"
#include "postmaster/postmaster.h"
+#include "storage/aio_subsys.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
#include "storage/latch.h"
@@ -465,6 +466,7 @@ AutoVacLauncherMain(char *startup_data, size_t startup_data_len)
*/
LWLockReleaseAll();
pgstat_report_wait_end();
+ pgaio_error_cleanup();
UnlockBuffers();
/* this is probably dead code, but let's be safe: */
if (AuxProcessResourceOwner)
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index 3eff5dc6f0e..ec1225c433f 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -38,6 +38,7 @@
#include "postmaster/auxprocess.h"
#include "postmaster/bgwriter.h"
#include "postmaster/interrupt.h"
+#include "storage/aio_subsys.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
@@ -168,6 +169,7 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
*/
LWLockReleaseAll();
ConditionVariableCancelSleep();
+ pgaio_error_cleanup();
UnlockBuffers();
ReleaseAuxProcessResources(false);
AtEOXact_Buffers(false);
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index b94f9cdff21..d254b2d1587 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -49,6 +49,7 @@
#include "postmaster/bgwriter.h"
#include "postmaster/interrupt.h"
#include "replication/syncrep.h"
+#include "storage/aio_subsys.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/fd.h"
@@ -276,6 +277,7 @@ CheckpointerMain(char *startup_data, size_t startup_data_len)
LWLockReleaseAll();
ConditionVariableCancelSleep();
pgstat_report_wait_end();
+ pgaio_error_cleanup();
UnlockBuffers();
ReleaseAuxProcessResources(false);
AtEOXact_Buffers(false);
diff --git a/src/backend/postmaster/pgarch.c b/src/backend/postmaster/pgarch.c
index 12ee815a626..6302e2a8314 100644
--- a/src/backend/postmaster/pgarch.c
+++ b/src/backend/postmaster/pgarch.c
@@ -40,6 +40,7 @@
#include "postmaster/interrupt.h"
#include "postmaster/pgarch.h"
#include "storage/condition_variable.h"
+#include "storage/aio_subsys.h"
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/latch.h"
@@ -568,6 +569,7 @@ pgarch_archiveXlog(char *xlog)
LWLockReleaseAll();
ConditionVariableCancelSleep();
pgstat_report_wait_end();
+ pgaio_error_cleanup();
ReleaseAuxProcessResources(false);
AtEOXact_Files(false);
AtEOXact_HashTables(false);
diff --git a/src/backend/postmaster/walsummarizer.c b/src/backend/postmaster/walsummarizer.c
index ffbf0439358..f38e181cf5e 100644
--- a/src/backend/postmaster/walsummarizer.c
+++ b/src/backend/postmaster/walsummarizer.c
@@ -37,6 +37,7 @@
#include "postmaster/interrupt.h"
#include "postmaster/walsummarizer.h"
#include "replication/walreceiver.h"
+#include "storage/aio_subsys.h"
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/latch.h"
@@ -288,6 +289,7 @@ WalSummarizerMain(char *startup_data, size_t startup_data_len)
LWLockReleaseAll();
ConditionVariableCancelSleep();
pgstat_report_wait_end();
+ pgaio_error_cleanup();
ReleaseAuxProcessResources(false);
AtEOXact_Files(false);
AtEOXact_HashTables(false);
diff --git a/src/backend/postmaster/walwriter.c b/src/backend/postmaster/walwriter.c
index df4f7634969..3015a744b5f 100644
--- a/src/backend/postmaster/walwriter.c
+++ b/src/backend/postmaster/walwriter.c
@@ -51,6 +51,7 @@
#include "postmaster/auxprocess.h"
#include "postmaster/interrupt.h"
#include "postmaster/walwriter.h"
+#include "storage/aio_subsys.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/fd.h"
@@ -164,6 +165,7 @@ WalWriterMain(char *startup_data, size_t startup_data_len)
LWLockReleaseAll();
ConditionVariableCancelSleep();
pgstat_report_wait_end();
+ pgaio_error_cleanup();
UnlockBuffers();
ReleaseAuxProcessResources(false);
AtEOXact_Buffers(false);
diff --git a/src/backend/replication/walsender.c b/src/backend/replication/walsender.c
index 446d10c1a7d..69097125606 100644
--- a/src/backend/replication/walsender.c
+++ b/src/backend/replication/walsender.c
@@ -79,6 +79,7 @@
#include "replication/walsender.h"
#include "replication/walsender_private.h"
#include "storage/condition_variable.h"
+#include "storage/aio_subsys.h"
#include "storage/fd.h"
#include "storage/ipc.h"
#include "storage/pmsignal.h"
@@ -327,6 +328,7 @@ WalSndErrorCleanup(void)
LWLockReleaseAll();
ConditionVariableCancelSleep();
pgstat_report_wait_end();
+ pgaio_error_cleanup();
if (xlogreader != NULL && xlogreader->seg.ws_file >= 0)
wal_segment_close(xlogreader);
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index eaeaeeee8e3..89f821ea7e1 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -10,7 +10,11 @@ include $(top_builddir)/src/Makefile.global
OBJS = \
aio.o \
+ aio_callback.o \
aio_init.o \
+ aio_io.o \
+ aio_target.o \
+ method_sync.o \
read_stream.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 8eb6a5f1292..0483580d644 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -3,6 +3,28 @@
* aio.c
* AIO - Core Logic
*
+ * For documentation about how AIO works on a higher level, including a
+ * schematic example, see README.md.
+ *
+ *
+ * AIO is a complicated subsystem. To keep things navigable it is split across
+ * a number of files:
+ *
+ * - method_*.c - different ways of executing AIO (e.g. worker process)
+ *
+ * - aio_target.c - IO on different kinds of targets
+ *
+ * - aio_io.c - method-independent code for specific IO ops (e.g. readv)
+ *
+ * - aio_callback.c - callbacks at IO operation lifecycle events
+ *
+ * - aio_init.c - per-server and per-backend initialization
+ *
+ * - aio.c - all other topics
+ *
+ * - read_stream.c - helper for reading buffered relation data
+ *
+ *
* Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
* Portions Copyright (c) 1994, Regents of the University of California
*
@@ -14,10 +36,28 @@
#include "postgres.h"
-#include "lib/ilist.h"
+#include "miscadmin.h"
+#include "port/atomics.h"
#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "storage/aio_subsys.h"
#include "utils/guc.h"
#include "utils/guc_hooks.h"
+#include "utils/resowner.h"
+#include "utils/wait_event_types.h"
+
+#ifdef USE_INJECTION_POINTS
+#include "utils/injection_point.h"
+#endif
+
+
+static inline void pgaio_io_update_state(PgAioHandle *ioh, PgAioHandleState new_state);
+static void pgaio_io_reclaim(PgAioHandle *ioh);
+static void pgaio_io_resowner_register(PgAioHandle *ioh);
+static void pgaio_io_wait_for_free(void);
+static PgAioHandle *pgaio_io_from_wref(PgAioWaitRef *iow, uint64 *ref_generation);
+static const char *pgaio_io_state_get_name(PgAioHandleState s);
+static void pgaio_io_wait(PgAioHandle *ioh, uint64 ref_generation);
/* Options for io_method. */
@@ -30,11 +70,1054 @@ const struct config_enum_entry io_method_options[] = {
int io_method = DEFAULT_IO_METHOD;
int io_max_concurrency = -1;
+/* global control for AIO */
+PgAioCtl *pgaio_ctl;
+/* current backend's per-backend state */
+PgAioBackend *pgaio_my_backend;
+
+
+static const IoMethodOps *const pgaio_method_ops_table[] = {
+ [IOMETHOD_SYNC] = &pgaio_sync_ops,
+};
+
+/* callbacks for the configured io_method, set by assign_io_method */
+const IoMethodOps *pgaio_method_ops;
+
+
+/*
+ * Currently there's no infrastructure to pass arguments to injection points,
+ * so we instead set this up for the duration of the injection point
+ * invocation.
+ */
+#ifdef USE_INJECTION_POINTS
+static PgAioHandle *pgaio_inj_cur_handle;
+#endif
+
+
+
+/* --------------------------------------------------------------------------------
+ * Public Functions related to PgAioHandle
+ * --------------------------------------------------------------------------------
+ */
+
+/*
+ * Acquire an AioHandle, waiting for IO completion if necessary.
+ *
+ * Each backend can only have one AIO handle that that has been "handed out"
+ * to code, but not yet submitted or released. This restriction is necessary
+ * to ensure that it is possible for code to wait for an unused handle by
+ * waiting for in-flight IO to complete. There is a limited number of handles
+ * in each backend, if multiple handles could be handed out without being
+ * submitted, waiting for all in-flight IO to complete would not guarantee
+ * that handles free up.
+ *
+ * It is cheap to acquire an IO handle, unless all handles are in use. In that
+ * case this function waits for the oldest IO to complete. In case that is not
+ * desirable, see pgaio_io_acquire_nb().
+ *
+ * If a handle was acquired but then does not turn out to be needed,
+ * e.g. because pgaio_io_acquire() is called before starting an IO in a
+ * critical section, the handle needs to be released with pgaio_io_release().
+ *
+ *
+ * To react to the completion of the IO as soon as it is known to have
+ * completed, callbacks can be registered with pgaio_io_register_callbacks().
+ *
+ * To actually execute IO using the returned handle, the pgaio_io_prep_*()
+ * family of functions is used. In many cases the pgaio_io_prep_*() call will
+ * not be done directly by code that acquired the handle, but by lower level
+ * code that gets passed the handle. E.g. if code in bufmgr.c wants to perform
+ * AIO, it typically will pass the handle to smgr., which will pass it on to
+ * md.c, on to fd.c, which then finally calls pgaio_io_prep_*(). This
+ * forwarding allows the various layers to react to the IO's completion by
+ * registering callbacks. These callbacks in turn can translate a lower
+ * layer's result into a result understandable by a higher layer.
+ *
+ * During pgaio_io_prep_*() the IO is staged (i.e. prepared for execution but
+ * not submitted to the kernel). Unless in batchmode
+ * (c.f. pgaio_enter_batchmode()) the IO will also have been submitted for
+ * execution. Note that, whether in batchmode or not, the IO might even
+ * complete before the functions return.
+ *
+ * After pgaio_io_prep_*() the AioHandle is "consumed" and may not be
+ * referenced by the IO issuing code. To e.g. wait for IO, references to the
+ * IO can be established with pgaio_io_get_wref() *before* pgaio_io_prep_*()
+ * is called. pgaio_wref_wait() can be used to wait for the IO to complete.
+ *
+ *
+ * To know if the IO [partially] succeeded or failed, a PgAioReturn * can be
+ * passed to pgaio_io_acquire(). Once the issuing backend has called
+ * pgaio_wref_wait(), the PgAioReturn contains information about whether the
+ * operation succeeded and details about the first failure, if any. The error
+ * can be raised / logged with pgaio_result_report().
+ *
+ * The lifetime of the memory pointed to be *ret needs to be at least as long
+ * as the passed in resowner. If the resowner releases resources before the IO
+ * completes (typically due to an error), the reference to *ret will be
+ * cleared. In case of resowner cleanup *ret will not be updated with the
+ * results of the IO operation.
+ */
+PgAioHandle *
+pgaio_io_acquire(struct ResourceOwnerData *resowner, PgAioReturn *ret)
+{
+ PgAioHandle *h;
+
+ while (true)
+ {
+ h = pgaio_io_acquire_nb(resowner, ret);
+
+ if (h != NULL)
+ return h;
+
+ /*
+ * Evidently all handles by this backend are in use. Just wait for
+ * some to complete.
+ */
+ pgaio_io_wait_for_free();
+ }
+}
+
+/*
+ * Acquire an AioHandle, returning NULL if no handles are free.
+ *
+ * See pgaio_io_acquire(). The only difference is that this function will return
+ * NULL if there are no idle handles, instead of blocking.
+ */
+PgAioHandle *
+pgaio_io_acquire_nb(struct ResourceOwnerData *resowner, PgAioReturn *ret)
+{
+ if (pgaio_my_backend->num_staged_ios >= PGAIO_SUBMIT_BATCH_SIZE)
+ {
+ Assert(pgaio_my_backend->num_staged_ios == PGAIO_SUBMIT_BATCH_SIZE);
+ pgaio_submit_staged();
+ }
+
+ if (pgaio_my_backend->handed_out_io)
+ {
+ ereport(ERROR,
+ errmsg("API violation: Only one IO can be handed out"));
+ }
+
+ if (!dclist_is_empty(&pgaio_my_backend->idle_ios))
+ {
+ dlist_node *ion = dclist_pop_head_node(&pgaio_my_backend->idle_ios);
+ PgAioHandle *ioh = dclist_container(PgAioHandle, node, ion);
+
+ Assert(ioh->state == PGAIO_HS_IDLE);
+ Assert(ioh->owner_procno == MyProcNumber);
+
+ pgaio_io_update_state(ioh, PGAIO_HS_HANDED_OUT);
+ pgaio_my_backend->handed_out_io = ioh;
+
+ if (resowner)
+ pgaio_io_resowner_register(ioh);
+
+ if (ret)
+ {
+ ioh->report_return = ret;
+ ret->result.status = ARS_UNKNOWN;
+ }
+
+ return ioh;
+ }
+
+ return NULL;
+}
+
+/*
+ * Release IO handle that turned out to not be required.
+ *
+ * See pgaio_io_acquire() for more details.
+ */
+void
+pgaio_io_release(PgAioHandle *ioh)
+{
+ if (ioh == pgaio_my_backend->handed_out_io)
+ {
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT);
+ Assert(ioh->resowner);
+
+ pgaio_my_backend->handed_out_io = NULL;
+ pgaio_io_reclaim(ioh);
+ }
+ else
+ {
+ elog(ERROR, "release in unexpected state");
+ }
+}
+
+/*
+ * Release IO handle during resource owner cleanup.
+ */
+void
+pgaio_io_release_resowner(dlist_node *ioh_node, bool on_error)
+{
+ PgAioHandle *ioh = dlist_container(PgAioHandle, resowner_node, ioh_node);
+
+ Assert(ioh->resowner);
+
+ ResourceOwnerForgetAioHandle(ioh->resowner, &ioh->resowner_node);
+ ioh->resowner = NULL;
+
+ switch (ioh->state)
+ {
+ case PGAIO_HS_IDLE:
+ elog(ERROR, "unexpected");
+ break;
+ case PGAIO_HS_HANDED_OUT:
+ Assert(ioh == pgaio_my_backend->handed_out_io || pgaio_my_backend->handed_out_io == NULL);
+
+ if (ioh == pgaio_my_backend->handed_out_io)
+ {
+ pgaio_my_backend->handed_out_io = NULL;
+ if (!on_error)
+ elog(WARNING, "leaked AIO handle");
+ }
+
+ pgaio_io_reclaim(ioh);
+ break;
+ case PGAIO_HS_DEFINED:
+ case PGAIO_HS_STAGED:
+ /* XXX: Should we warn about this when is_commit? */
+ pgaio_submit_staged();
+ break;
+ case PGAIO_HS_SUBMITTED:
+ case PGAIO_HS_COMPLETED_IO:
+ case PGAIO_HS_COMPLETED_SHARED:
+ case PGAIO_HS_COMPLETED_LOCAL:
+ /* this is expected to happen */
+ break;
+ }
+
+ /*
+ * Need to unregister the reporting of the IO's result, the memory it's
+ * referencing likely has gone away.
+ */
+ if (ioh->report_return)
+ ioh->report_return = NULL;
+}
+
+/*
+ * Add a [set of] flags to the IO.
+ *
+ * Note that this combines flags with already set flags, rather than set flags
+ * to explicitly the passed in parameters. This is to allow multiple callsites
+ * to set flags.
+ */
+void
+pgaio_io_set_flag(PgAioHandle *ioh, PgAioHandleFlags flag)
+{
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT);
+
+ ioh->flags |= flag;
+}
+
+int
+pgaio_io_get_id(PgAioHandle *ioh)
+{
+ Assert(ioh >= pgaio_ctl->io_handles &&
+ ioh < (pgaio_ctl->io_handles + pgaio_ctl->io_handle_count));
+ return ioh - pgaio_ctl->io_handles;
+}
+
+ProcNumber
+pgaio_io_get_owner(PgAioHandle *ioh)
+{
+ return ioh->owner_procno;
+}
+
+void
+pgaio_io_get_wref(PgAioHandle *ioh, PgAioWaitRef *iow)
+{
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT ||
+ ioh->state == PGAIO_HS_DEFINED ||
+ ioh->state == PGAIO_HS_STAGED);
+ Assert(ioh->generation != 0);
+
+ iow->aio_index = ioh - pgaio_ctl->io_handles;
+ iow->generation_upper = (uint32) (ioh->generation >> 32);
+ iow->generation_lower = (uint32) ioh->generation;
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Internal Functions related to PgAioHandle
+ * --------------------------------------------------------------------------------
+ */
+
+static inline void
+pgaio_io_update_state(PgAioHandle *ioh, PgAioHandleState new_state)
+{
+ pgaio_debug_io(DEBUG5, ioh,
+ "updating state to %s",
+ pgaio_io_state_get_name(new_state));
+
+ /*
+ * Ensure the changes signified by the new state are visible before the
+ * new state becomes visible.
+ */
+ pg_write_barrier();
+
+ ioh->state = new_state;
+}
+
+static void
+pgaio_io_resowner_register(PgAioHandle *ioh)
+{
+ Assert(!ioh->resowner);
+ Assert(CurrentResourceOwner);
+
+ ResourceOwnerRememberAioHandle(CurrentResourceOwner, &ioh->resowner_node);
+ ioh->resowner = CurrentResourceOwner;
+}
+
+/*
+ * Stage IO for execution and, if necessary, submit it immediately.
+ *
+ * Should only be called from pgaio_io_prep_*().
+ */
+void
+pgaio_io_stage(PgAioHandle *ioh, PgAioOp op)
+{
+ bool needs_synchronous;
+
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT);
+ Assert(pgaio_io_has_target(ioh));
+
+ ioh->op = op;
+ ioh->result = 0;
+
+ pgaio_io_update_state(ioh, PGAIO_HS_DEFINED);
+
+ /* allow a new IO to be staged */
+ pgaio_my_backend->handed_out_io = NULL;
+
+ pgaio_io_call_stage(ioh);
+
+ pgaio_io_update_state(ioh, PGAIO_HS_STAGED);
+
+ /*
+ * Synchronous execution has to be executed, well, synchronously, so check
+ * that first.
+ */
+ needs_synchronous = pgaio_io_needs_synchronous_execution(ioh);
+
+ pgaio_debug_io(DEBUG3, ioh,
+ "prepared (synchronous: %d, in_batch: %d)",
+ needs_synchronous, pgaio_my_backend->in_batchmode);
+
+ if (!needs_synchronous)
+ {
+ pgaio_my_backend->staged_ios[pgaio_my_backend->num_staged_ios++] = ioh;
+ Assert(pgaio_my_backend->num_staged_ios <= PGAIO_SUBMIT_BATCH_SIZE);
+
+ /*
+ * Unless code explicitly opted into batching IOs, submit the IO
+ * immediately.
+ */
+ if (!pgaio_my_backend->in_batchmode)
+ pgaio_submit_staged();
+ }
+ else
+ {
+ pgaio_io_prepare_submit(ioh);
+ pgaio_io_perform_synchronously(ioh);
+ }
+}
+
+bool
+pgaio_io_needs_synchronous_execution(PgAioHandle *ioh)
+{
+ /*
+ * If the caller said to execute the IO synchronously, do so.
+ *
+ * XXX: We could optimize the logic when to execute synchronously by first
+ * checking if there are other IOs in flight and only synchronously
+ * executing if not. Unclear whether that'll be sufficiently common to be
+ * worth worrying about.
+ */
+ if (ioh->flags & PGAIO_HF_SYNCHRONOUS)
+ return true;
+
+ /* Check if the IO method requires synchronous execution of IO */
+ if (pgaio_method_ops->needs_synchronous_execution)
+ return pgaio_method_ops->needs_synchronous_execution(ioh);
+
+ return false;
+}
+
+/*
+ * Handle IO being processed by IO method.
+ *
+ * Should be called by IO methods / synchronous IO execution, just before the
+ * IO is performed.
+ */
+void
+pgaio_io_prepare_submit(PgAioHandle *ioh)
+{
+ pgaio_io_update_state(ioh, PGAIO_HS_SUBMITTED);
+
+ dclist_push_tail(&pgaio_my_backend->in_flight_ios, &ioh->node);
+}
+
+/*
+ * Handle IO getting completed by a method.
+ *
+ * Should be called by IO methods / synchronous IO execution, just after the
+ * IO has been performed.
+ *
+ * Expects to be called in a critical section. We expect IOs to be usable for
+ * WAL etc, which requires being able to execute completion callbacks in a
+ * critical section.
+ */
+void
+pgaio_io_process_completion(PgAioHandle *ioh, int result)
+{
+ Assert(ioh->state == PGAIO_HS_SUBMITTED);
+
+ Assert(CritSectionCount > 0);
+
+ ioh->result = result;
+
+ pgaio_io_update_state(ioh, PGAIO_HS_COMPLETED_IO);
+
+ pgaio_io_call_inj(ioh, "AIO_PROCESS_COMPLETION_BEFORE_SHARED");
+
+ pgaio_io_call_complete_shared(ioh);
+
+ pgaio_io_update_state(ioh, PGAIO_HS_COMPLETED_SHARED);
+
+ /* condition variable broadcast ensures state is visible before wakeup */
+ ConditionVariableBroadcast(&ioh->cv);
+
+ /* contains call to pgaio_io_call_complete_local() */
+ if (ioh->owner_procno == MyProcNumber)
+ pgaio_io_reclaim(ioh);
+}
+
+bool
+pgaio_io_was_recycled(PgAioHandle *ioh, uint64 ref_generation, PgAioHandleState *state)
+{
+ *state = ioh->state;
+ pg_read_barrier();
+
+ return ioh->generation != ref_generation;
+}
+
+/*
+ * Wait for IO to complete. External code should never use this, outside of
+ * the AIO subsystem waits are only allowed via pgaio_wref_wait().
+ */
+static void
+pgaio_io_wait(PgAioHandle *ioh, uint64 ref_generation)
+{
+ PgAioHandleState state;
+ bool am_owner;
+
+ am_owner = ioh->owner_procno == MyProcNumber;
+
+ if (pgaio_io_was_recycled(ioh, ref_generation, &state))
+ return;
+
+ if (am_owner)
+ {
+ if (state == PGAIO_HS_STAGED)
+ {
+ /* XXX: Arguably this should be prevented by callers? */
+ pgaio_submit_staged();
+ }
+ else if (state != PGAIO_HS_SUBMITTED
+ && state != PGAIO_HS_COMPLETED_IO
+ && state != PGAIO_HS_COMPLETED_SHARED
+ && state != PGAIO_HS_COMPLETED_LOCAL)
+ {
+ elog(PANIC, "waiting for own IO in wrong state: %d",
+ state);
+ }
+ }
+
+ while (true)
+ {
+ if (pgaio_io_was_recycled(ioh, ref_generation, &state))
+ return;
+
+ switch (state)
+ {
+ case PGAIO_HS_IDLE:
+ case PGAIO_HS_HANDED_OUT:
+ elog(ERROR, "IO in wrong state: %d", state);
+ break;
+
+ case PGAIO_HS_SUBMITTED:
+
+ /*
+ * If we need to wait via the IO method, do so now. Don't
+ * check via the IO method if the issuing backend is executing
+ * the IO synchronously.
+ */
+ if (pgaio_method_ops->wait_one && !(ioh->flags & PGAIO_HF_SYNCHRONOUS))
+ {
+ pgaio_method_ops->wait_one(ioh, ref_generation);
+ continue;
+ }
+ /* fallthrough */
+
+ /* waiting for owner to submit */
+ case PGAIO_HS_DEFINED:
+ case PGAIO_HS_STAGED:
+ /* waiting for reaper to complete */
+ /* fallthrough */
+ case PGAIO_HS_COMPLETED_IO:
+ /* shouldn't be able to hit this otherwise */
+ Assert(IsUnderPostmaster);
+ /* ensure we're going to get woken up */
+ ConditionVariablePrepareToSleep(&ioh->cv);
+
+ while (!pgaio_io_was_recycled(ioh, ref_generation, &state))
+ {
+ if (state == PGAIO_HS_COMPLETED_SHARED ||
+ state == PGAIO_HS_COMPLETED_LOCAL)
+ break;
+ ConditionVariableSleep(&ioh->cv, WAIT_EVENT_AIO_IO_COMPLETION);
+ }
+
+ ConditionVariableCancelSleep();
+ break;
+
+ case PGAIO_HS_COMPLETED_SHARED:
+ case PGAIO_HS_COMPLETED_LOCAL:
+ /* see above */
+ if (am_owner)
+ pgaio_io_reclaim(ioh);
+ return;
+ }
+ }
+}
+
+static void
+pgaio_io_reclaim(PgAioHandle *ioh)
+{
+ /* This is only ok if it's our IO */
+ Assert(ioh->owner_procno == MyProcNumber);
+
+ /*
+ * It's a bit ugly, but right now the easiest place to put the execution
+ * of shared completion callbacks is this function, as we need to execute
+ * locallbacks just before reclaiming at multiple callsites.
+ */
+ if (ioh->state == PGAIO_HS_COMPLETED_SHARED)
+ {
+ pgaio_io_call_complete_local(ioh);
+ pgaio_io_update_state(ioh, PGAIO_HS_COMPLETED_LOCAL);
+ }
+
+ pgaio_debug_io(DEBUG4, ioh,
+ "reclaiming: distilled_result: (status %s, id %u, error_data %d), raw_result: %d",
+ pgaio_result_status_string(ioh->distilled_result.status),
+ ioh->distilled_result.id,
+ ioh->distilled_result.error_data,
+ ioh->result);
+
+ /* if the IO has been defined, we might need to do more work */
+ if (ioh->state != PGAIO_HS_HANDED_OUT)
+ {
+ dclist_delete_from(&pgaio_my_backend->in_flight_ios, &ioh->node);
+
+ if (ioh->report_return)
+ {
+ ioh->report_return->result = ioh->distilled_result;
+ ioh->report_return->target_data = ioh->target_data;
+ }
+ }
+
+ if (ioh->resowner)
+ {
+ ResourceOwnerForgetAioHandle(ioh->resowner, &ioh->resowner_node);
+ ioh->resowner = NULL;
+ }
+
+ Assert(!ioh->resowner);
+
+ ioh->op = PGAIO_OP_INVALID;
+ ioh->target = PGAIO_TID_INVALID;
+ ioh->flags = 0;
+ ioh->num_shared_callbacks = 0;
+ ioh->handle_data_len = 0;
+ ioh->report_return = NULL;
+ ioh->result = 0;
+ ioh->distilled_result.status = ARS_UNKNOWN;
+
+ /* XXX: the barrier is probably superfluous */
+ pg_write_barrier();
+ ioh->generation++;
+
+ pgaio_io_update_state(ioh, PGAIO_HS_IDLE);
+
+ /*
+ * We push the IO to the head of the idle IO list, that seems more cache
+ * efficient in cases where only a few IOs are used.
+ */
+ dclist_push_head(&pgaio_my_backend->idle_ios, &ioh->node);
+}
+
+static void
+pgaio_io_wait_for_free(void)
+{
+ int reclaimed = 0;
+
+ pgaio_debug(DEBUG2, "waiting for self with %d pending",
+ pgaio_my_backend->num_staged_ios);
+
+ /*
+ * First check if any of our IOs actually have completed - when using
+ * worker, that'll often be the case. We could do so as part of the loop
+ * below, but that'd potentially lead us to wait for some IO submitted
+ * before.
+ */
+ for (int i = 0; i < io_max_concurrency; i++)
+ {
+ PgAioHandle *ioh = &pgaio_ctl->io_handles[pgaio_my_backend->io_handle_off + i];
+
+ if (ioh->state == PGAIO_HS_COMPLETED_SHARED)
+ {
+ pgaio_io_reclaim(ioh);
+ reclaimed++;
+ }
+ }
+
+ if (reclaimed > 0)
+ return;
+
+ /*
+ * If we have any unsubmitted IOs, submit them now. We'll start waiting in
+ * a second, so it's better they're in flight. This also addresses the
+ * edge-case that all IOs are unsubmitted.
+ */
+ if (pgaio_my_backend->num_staged_ios > 0)
+ {
+ pgaio_submit_staged();
+ }
+
+ /*
+ * It's possible that we recognized there were free IOs while submitting.
+ */
+ if (dclist_count(&pgaio_my_backend->in_flight_ios) == 0)
+ {
+ elog(ERROR, "no free IOs despite no in-flight IOs");
+ }
+
+ /*
+ * Wait for the oldest in-flight IO to complete.
+ *
+ * XXX: Reusing the general IO wait is suboptimal, we don't need to wait
+ * for that specific IO to complete, we just need *any* IO to complete.
+ */
+ {
+ PgAioHandle *ioh = dclist_head_element(PgAioHandle, node, &pgaio_my_backend->in_flight_ios);
+
+ switch (ioh->state)
+ {
+ /* should not be in in-flight list */
+ case PGAIO_HS_IDLE:
+ case PGAIO_HS_DEFINED:
+ case PGAIO_HS_HANDED_OUT:
+ case PGAIO_HS_STAGED:
+ case PGAIO_HS_COMPLETED_LOCAL:
+ elog(ERROR, "shouldn't get here with io:%d in state %d",
+ pgaio_io_get_id(ioh), ioh->state);
+ break;
+
+ case PGAIO_HS_COMPLETED_IO:
+ case PGAIO_HS_SUBMITTED:
+ pgaio_debug_io(DEBUG2, ioh,
+ "waiting for free io with %d in flight",
+ dclist_count(&pgaio_my_backend->in_flight_ios));
+
+ /*
+ * In a more general case this would be racy, because the
+ * generation could increase after we read ioh->state above.
+ * But we are only looking at IOs by the current backend and
+ * the IO can only be recycled by this backend.
+ */
+ pgaio_io_wait(ioh, ioh->generation);
+ break;
+
+ case PGAIO_HS_COMPLETED_SHARED:
+ /* it's possible that another backend just finished this IO */
+ pgaio_io_reclaim(ioh);
+ break;
+ }
+
+ if (dclist_count(&pgaio_my_backend->idle_ios) == 0)
+ elog(PANIC, "no idle IOs after waiting");
+ return;
+ }
+}
+
+/*
+ * Internal - code outside of AIO should never need this and it'd be hard for
+ * such code to be safe.
+ */
+static PgAioHandle *
+pgaio_io_from_wref(PgAioWaitRef *iow, uint64 *ref_generation)
+{
+ PgAioHandle *ioh;
+
+ Assert(iow->aio_index < pgaio_ctl->io_handle_count);
+
+ ioh = &pgaio_ctl->io_handles[iow->aio_index];
+
+ *ref_generation = ((uint64) iow->generation_upper) << 32 |
+ iow->generation_lower;
+
+ Assert(*ref_generation != 0);
+
+ return ioh;
+}
+
+static const char *
+pgaio_io_state_get_name(PgAioHandleState s)
+{
+#define PGAIO_HS_TOSTR_CASE(sym) case PGAIO_HS_##sym: return #sym
+ switch (s)
+ {
+ PGAIO_HS_TOSTR_CASE(IDLE);
+ PGAIO_HS_TOSTR_CASE(HANDED_OUT);
+ PGAIO_HS_TOSTR_CASE(DEFINED);
+ PGAIO_HS_TOSTR_CASE(STAGED);
+ PGAIO_HS_TOSTR_CASE(SUBMITTED);
+ PGAIO_HS_TOSTR_CASE(COMPLETED_IO);
+ PGAIO_HS_TOSTR_CASE(COMPLETED_SHARED);
+ PGAIO_HS_TOSTR_CASE(COMPLETED_LOCAL);
+ }
+#undef PGAIO_HS_TOSTR_CASE
+
+ return NULL; /* silence compiler */
+}
+
+const char *
+pgaio_io_get_state_name(PgAioHandle *ioh)
+{
+ return pgaio_io_state_get_name(ioh->state);
+}
+
+const char *
+pgaio_result_status_string(PgAioResultStatus rs)
+{
+ switch (rs)
+ {
+ case ARS_UNKNOWN:
+ return "UNKNOWN";
+ case ARS_OK:
+ return "OK";
+ case ARS_PARTIAL:
+ return "PARTIAL";
+ case ARS_ERROR:
+ return "ERROR";
+ }
+
+ return NULL; /* silence compiler */
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Functions primarily related to IO Wait References
+ * --------------------------------------------------------------------------------
+ */
+
+void
+pgaio_wref_clear(PgAioWaitRef *iow)
+{
+ iow->aio_index = PG_UINT32_MAX;
+}
+
+bool
+pgaio_wref_valid(PgAioWaitRef *iow)
+{
+ return iow->aio_index != PG_UINT32_MAX;
+}
+
+int
+pgaio_wref_get_id(PgAioWaitRef *iow)
+{
+ Assert(pgaio_wref_valid(iow));
+ return iow->aio_index;
+}
+
+/*
+ * Wait for the IO to have completed.
+ */
+void
+pgaio_wref_wait(PgAioWaitRef *iow)
+{
+ uint64 ref_generation;
+ PgAioHandle *ioh;
+
+ ioh = pgaio_io_from_wref(iow, &ref_generation);
+
+ pgaio_io_wait(ioh, ref_generation);
+}
+
+/*
+ * Check if the the referenced IO completed, without blocking.
+ */
+bool
+pgaio_wref_check_done(PgAioWaitRef *iow)
+{
+ uint64 ref_generation;
+ PgAioHandleState state;
+ bool am_owner;
+ PgAioHandle *ioh;
+
+ ioh = pgaio_io_from_wref(iow, &ref_generation);
+
+ if (pgaio_io_was_recycled(ioh, ref_generation, &state))
+ return true;
+
+ if (state == PGAIO_HS_IDLE)
+ return true;
+
+ am_owner = ioh->owner_procno == MyProcNumber;
+
+ if (state == PGAIO_HS_COMPLETED_SHARED ||
+ state == PGAIO_HS_COMPLETED_LOCAL)
+ {
+ if (am_owner)
+ pgaio_io_reclaim(ioh);
+ return true;
+ }
+
+ return false;
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Actions on multiple IOs.
+ * --------------------------------------------------------------------------------
+ */
+
+/*
+ * Submit IOs in batches going forward.
+ *
+ * Submitting multiple IOs at once can be substantially faster than doing so
+ * one-by-one. At the same time submitting multiple IOs at once requires more
+ * care to avoid deadlocks.
+ *
+ * Consider backend A staging an IO for buffer 1 and then trying to start IO
+ * on buffer 2, while backend B does the inverse. If A submitted the IO before
+ * moving on to buffer 2, this works just fine, B will wait for the IO to
+ * complete. But if batching were used, each backend will wait for IO that has
+ * not yet been submitted to complete, i.e. forever.
+ *
+ * Batch submission mode needs to explicitly ended with
+ * pgaio_exit_batchmode(), but it is allowed to throw errors, in which case
+ * error recovery will end the batch.
+ *
+ * To avoid deadlocks, code needs to ensure that it will not wait for another
+ * backend while there is unsubmitted IO. E.g. by using conditional lock
+ * acquisition when acquiring buffer locks. To check if there currently are
+ * staged IOs, call pgaio_have_staged() and to submit all staged IOs call
+ * pgaio_submit_staged().
+ *
+ * It is not allowed to enter batchmode while already in batchmode, it's
+ * unlikely to ever be needed, as code needs to be explicitly aware of being
+ * called in batchmode, to avoid the deadlock risks explained above.
+ *
+ * Note that IOs may get submitted before pgaio_exit_batchmode() is called,
+ * e.g. because too many IOs have been staged or because pgaio_submit_staged()
+ * was called.
+ */
+void
+pgaio_enter_batchmode(void)
+{
+ if (pgaio_my_backend->in_batchmode)
+ elog(ERROR, "starting batch while batch already in progress");
+ pgaio_my_backend->in_batchmode = true;
+}
+
+/*
+ * Stop submitting IOs in batches.
+ */
+void
+pgaio_exit_batchmode(void)
+{
+ Assert(pgaio_my_backend->in_batchmode);
+
+ pgaio_submit_staged();
+ pgaio_my_backend->in_batchmode = false;
+}
+
+/*
+ * Are there staged but unsubmitted IOs?
+ *
+ * See comment above pgaio_enter_batchmode() for why code may need to check if
+ * there is IO in that state.
+ */
+bool
+pgaio_have_staged(void)
+{
+ Assert(pgaio_my_backend->in_batchmode ||
+ pgaio_my_backend->num_staged_ios == 0);
+ return pgaio_my_backend->num_staged_ios > 0;
+}
+
+/*
+ * Submit all staged but not yet submitted IOs.
+ *
+ * Unless in batch mode, this never needs to be called, as IOs get submitted
+ * as soon as possible. While in batchmode pgaio_submit_staged() can be called
+ * before waiting on another backend, to avoid the risk of deadlocks. See
+ * pgaio_enter_batchmode().
+ */
+void
+pgaio_submit_staged(void)
+{
+ int total_submitted = 0;
+ int did_submit;
+
+ if (pgaio_my_backend->num_staged_ios == 0)
+ return;
+
+
+ START_CRIT_SECTION();
+
+ did_submit = pgaio_method_ops->submit(pgaio_my_backend->num_staged_ios,
+ pgaio_my_backend->staged_ios);
+
+ END_CRIT_SECTION();
+
+ total_submitted += did_submit;
+
+ Assert(total_submitted == did_submit);
+
+ pgaio_my_backend->num_staged_ios = 0;
+
+ pgaio_debug(DEBUG4,
+ "aio: submitted %d IOs",
+ total_submitted);
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Other
+ * --------------------------------------------------------------------------------
+ */
+
+/*
+ * Need to submit staged but not yet submitted IOs using the fd, otherwise
+ * the IO would end up targeting something bogus.
+ */
+void
+pgaio_closing_fd(int fd)
+{
+ /*
+ * Might be called before AIO is initialized or in a subprocess that
+ * doesn't use AIO.
+ */
+ if (!pgaio_my_backend)
+ return;
+
+ /*
+ * For now just submit all staged IOs - we could be more selective, but
+ * it's probably not worth it.
+ */
+ pgaio_submit_staged();
+}
+
+/*
+ * Perform AIO related cleanup after an error.
+ *
+ * This should be called early in the error recovery paths, as later steps may
+ * need to issue AIO (e.g. to record a transaction abort WAL record).
+ */
+void
+pgaio_error_cleanup(void)
+{
+ /*
+ * It is possible that code errored out after pgaio_enter_batchmode() but
+ * before pgaio_exit_batchmode() was called. In that case we need to
+ * submit the IO now.
+ */
+ if (pgaio_my_backend->in_batchmode)
+ {
+ pgaio_my_backend->in_batchmode = false;
+
+ pgaio_submit_staged();
+ }
+
+ /*
+ * As we aren't in batchmode, there shouldn't be any unsubmitted IOs.
+ */
+ Assert(pgaio_my_backend->num_staged_ios == 0);
+}
+
+/*
+ * Perform AIO related checks at (sub-)transactional boundaries.
+ *
+ * This should be called late during (sub-)transactional commit/abort, after
+ * all steps that might need to perform AIO, so that we can verify that the
+ * AIO subsystem is in a valid state at the end of a transaction.
+ */
+void
+AtEOXact_Aio(bool is_commit)
+{
+ /*
+ * We should never be in batch mode at transactional boundaries. In case
+ * an error was thrown while in batch mode, pgaio_error_cleanup() should
+ * have exited batchmode.
+ *
+ * In case we are in batchmode somehow, make sure to submit all staged
+ * IOs, other backends may need them to complete to continue.
+ */
+ if (pgaio_my_backend->in_batchmode)
+ {
+ pgaio_error_cleanup();
+ elog(WARNING, "open AIO batch at end of (sub-)transaction");
+ }
+
+ /*
+ * As we aren't in batchmode, there shouldn't be any unsubmitted IOs.
+ */
+ Assert(pgaio_my_backend->num_staged_ios == 0);
+}
+
+void
+pgaio_shutdown(int code, Datum arg)
+{
+ AtEOXact_Aio(code == 0);
+ Assert(pgaio_my_backend);
+ Assert(!pgaio_my_backend->handed_out_io);
+
+ /*
+ * Before exiting, make sure that all IOs are finished. That has two main
+ * purposes: - it's somewhat annoying to see partially finished IOs in
+ * stats views etc - it's rumored that some kernel-level AIO mechanisms
+ * don't deal well with the issuer of an AIO exiting
+ */
+
+ while (!dclist_is_empty(&pgaio_my_backend->in_flight_ios))
+ {
+ PgAioHandle *ioh = dclist_head_element(PgAioHandle, node, &pgaio_my_backend->in_flight_ios);
+
+ /* see comment in pgaio_io_wait_for_free() about raciness */
+ pgaio_io_wait(ioh, ioh->generation);
+ }
+
+ pgaio_my_backend = NULL;
+}
void
assign_io_method(int newval, void *extra)
{
+ Assert(pgaio_method_ops_table[newval] != NULL);
+ Assert(newval < lengthof(io_method_options));
+
+ pgaio_method_ops = pgaio_method_ops_table[newval];
}
bool
@@ -57,11 +1140,41 @@ check_io_max_concurrency(int *newval, void **extra, GucSource source)
return true;
}
+
+/* --------------------------------------------------------------------------------
+ * Injection point support
+ * --------------------------------------------------------------------------------
+ */
+
+#ifdef USE_INJECTION_POINTS
+
/*
- * Release IO handle during resource owner cleanup.
+ * Call injection point with support for pgaio_inj_io_get().
*/
void
-pgaio_io_release_resowner(dlist_node *ioh_node, bool on_error)
+pgaio_io_call_inj(PgAioHandle *ioh, const char *injection_point)
{
- /* placeholder for later */
+ pgaio_inj_cur_handle = ioh;
+
+ PG_TRY();
+ {
+ InjectionPointCached(injection_point);
+ }
+ PG_FINALLY();
+ {
+ pgaio_inj_cur_handle = NULL;
+ }
+ PG_END_TRY();
}
+
+/*
+ * Return IO associated with injection point invocation. This is only needed
+ * as injection points currently don't support arguments.
+ */
+PgAioHandle *
+pgaio_inj_io_get(void)
+{
+ return pgaio_inj_cur_handle;
+}
+
+#endif
diff --git a/src/backend/storage/aio/aio_callback.c b/src/backend/storage/aio/aio_callback.c
new file mode 100644
index 00000000000..5629dc4cc94
--- /dev/null
+++ b/src/backend/storage/aio/aio_callback.c
@@ -0,0 +1,288 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_callback.c
+ * AIO - Functionality related to callbacks that can be registered on IO
+ * Handles
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio_callback.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "miscadmin.h"
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "utils/memutils.h"
+
+
+/* just to have something to put into the aio_handle_cbs */
+static const struct PgAioHandleCallbacks aio_invalid_cb = {0};
+
+typedef struct PgAioHandleCallbacksEntry
+{
+ const PgAioHandleCallbacks *const cb;
+ const char *const name;
+} PgAioHandleCallbacksEntry;
+
+/*
+ * Callback definition for the callbacks that can be registered on an IO
+ * handle. See PgAioHandleCallbackID's definition for an explanation for why
+ * callbacks are not identified by a pointer.
+ */
+static const PgAioHandleCallbacksEntry aio_handle_cbs[] = {
+#define CALLBACK_ENTRY(id, callback) [id] = {.cb = &callback, .name = #callback}
+ CALLBACK_ENTRY(PGAIO_HCB_INVALID, aio_invalid_cb),
+#undef CALLBACK_ENTRY
+};
+
+
+
+/*
+ * Register callback for the IO handle.
+ *
+ * Only a limited number (PGAIO_HANDLE_MAX_CALLBACKS) of callbacks can be
+ * registered for each IO.
+ *
+ * Callbacks need to be registered before [indirectly] calling
+ * pgaio_io_prep_*(), as the IO may be executed immediately.
+ *
+ *
+ * Note that callbacks are executed in critical sections. This is necessary
+ * to be able to execute IO in critical sections (consider e.g. WAL
+ * logging). To perform AIO we first need to acquire a handle, which, if there
+ * are no free handles, requires waiting for IOs to complete and to execute
+ * their completion callbacks.
+ *
+ * Callbacks may be executed in the issuing backend but also in another
+ * backend (because that backend is waiting for the IO) or in IO workers (if
+ * io_method=worker is used).
+ *
+ *
+ * See PgAioHandleCallbackID's definition for an explanation for why
+ * callbacks are not identified by a pointer.
+ */
+void
+pgaio_io_register_callbacks(PgAioHandle *ioh, PgAioHandleCallbackID cbid)
+{
+ const PgAioHandleCallbacksEntry *ce = &aio_handle_cbs[cbid];
+
+ if (cbid >= lengthof(aio_handle_cbs))
+ elog(ERROR, "callback %d is out of range", cbid);
+ if (aio_handle_cbs[cbid].cb->complete_shared == NULL &&
+ aio_handle_cbs[cbid].cb->complete_local == NULL)
+ elog(ERROR, "callback %d does not have completion callback", cbid);
+ if (ioh->num_shared_callbacks >= PGAIO_HANDLE_MAX_CALLBACKS)
+ elog(PANIC, "too many callbacks, the max is %d", PGAIO_HANDLE_MAX_CALLBACKS);
+ ioh->shared_callbacks[ioh->num_shared_callbacks] = cbid;
+
+ pgaio_debug_io(DEBUG3, ioh,
+ "adding cb #%d, id %d/%s",
+ ioh->num_shared_callbacks + 1,
+ cbid, ce->name);
+
+ ioh->num_shared_callbacks++;
+}
+
+/*
+ * Associate an array of data with the Handle. This is e.g. useful to the
+ * transport knowledge about which buffers a multi-block IO affects to
+ * completion callbacks.
+ *
+ * Right now this can be done only once for each IO, even though multiple
+ * callbacks can be registered. There aren't any known usecases requiring more
+ * and the required amount of shared memory does add up, so it doesn't seem
+ * worth multiplying memory usage by PGAIO_HANDLE_MAX_CALLBACKS.
+ */
+void
+pgaio_io_set_handle_data_64(PgAioHandle *ioh, uint64 *data, uint8 len)
+{
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT);
+ Assert(ioh->handle_data_len == 0);
+ Assert(len <= PG_IOV_MAX);
+
+ for (int i = 0; i < len; i++)
+ pgaio_ctl->handle_data[ioh->iovec_off + i] = data[i];
+ ioh->handle_data_len = len;
+}
+
+/*
+ * Convenience version of pgaio_io_set_handle_data_64() that converts a 32bit
+ * array to a 64bit array. Without it callers would end up needing to
+ * open-code equivalent code.
+ */
+void
+pgaio_io_set_handle_data_32(PgAioHandle *ioh, uint32 *data, uint8 len)
+{
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT);
+ Assert(ioh->handle_data_len == 0);
+ Assert(len <= PG_IOV_MAX);
+
+ for (int i = 0; i < len; i++)
+ pgaio_ctl->handle_data[ioh->iovec_off + i] = data[i];
+ ioh->handle_data_len = len;
+}
+
+/*
+ * Return data set with pgaio_io_set_handle_data_*().
+ */
+uint64 *
+pgaio_io_get_handle_data(PgAioHandle *ioh, uint8 *len)
+{
+ Assert(ioh->handle_data_len > 0);
+
+ *len = ioh->handle_data_len;
+
+ return &pgaio_ctl->handle_data[ioh->iovec_off];
+}
+
+/*
+ * Internal function which invokes ->stage for all the registered callbacks.
+ */
+void
+pgaio_io_call_stage(PgAioHandle *ioh)
+{
+ Assert(ioh->target > PGAIO_TID_INVALID && ioh->target < PGAIO_TID_COUNT);
+ Assert(ioh->op > PGAIO_OP_INVALID && ioh->op < PGAIO_OP_COUNT);
+
+ for (int i = ioh->num_shared_callbacks; i > 0; i--)
+ {
+ PgAioHandleCallbackID cbid = ioh->shared_callbacks[i - 1];
+ const PgAioHandleCallbacksEntry *ce = &aio_handle_cbs[cbid];
+
+ if (!ce->cb->stage)
+ continue;
+
+ pgaio_debug_io(DEBUG3, ioh,
+ "calling cb #%d %d/%s->stage",
+ i, cbid, ce->name);
+ ce->cb->stage(ioh);
+ }
+}
+
+/*
+ * Internal function which invokes ->complete_shared for all the registered
+ * callbacks.
+ */
+void
+pgaio_io_call_complete_shared(PgAioHandle *ioh)
+{
+ PgAioResult result;
+
+ START_CRIT_SECTION();
+
+ Assert(ioh->target > PGAIO_TID_INVALID && ioh->target < PGAIO_TID_COUNT);
+ Assert(ioh->op > PGAIO_OP_INVALID && ioh->op < PGAIO_OP_COUNT);
+
+ result.status = ARS_OK; /* low level IO is always considered OK */
+ result.result = ioh->result;
+ result.id = PGAIO_HCB_INVALID;
+ result.error_data = 0;
+
+ /*
+ * Call callbacks with the last registered (innermost) callback first.
+ * Each callback can modify the result forwarded to the next callback.
+ */
+ for (int i = ioh->num_shared_callbacks; i > 0; i--)
+ {
+ PgAioHandleCallbackID cbid = ioh->shared_callbacks[i - 1];
+ const PgAioHandleCallbacksEntry *ce = &aio_handle_cbs[cbid];
+
+ if (!ce->cb->complete_shared)
+ continue;
+
+ pgaio_debug_io(DEBUG4, ioh,
+ "calling cb #%d, id %d/%s->complete_shared with distilled result: (status %s, id %u, error_data %d, result %d)",
+ i, cbid, ce->name,
+ pgaio_result_status_string(result.status),
+ result.id, result.error_data, result.result);
+ result = ce->cb->complete_shared(ioh, result);
+ }
+
+ ioh->distilled_result = result;
+
+ pgaio_debug_io(DEBUG3, ioh,
+ "after shared completion: distilled result: (status %s, id %u, error_data: %d, result %d), raw_result: %d",
+ pgaio_result_status_string(result.status),
+ result.id, result.error_data, result.result,
+ ioh->result);
+
+ END_CRIT_SECTION();
+}
+
+
+/*
+ * Internal function which invokes ->complete_local for all the registered
+ * callbacks.
+ *
+ * XXX: It'd be nice to deduplicate with pgaio_io_call_complete_shared().
+ */
+void
+pgaio_io_call_complete_local(PgAioHandle *ioh)
+{
+ PgAioResult result;
+
+ START_CRIT_SECTION();
+
+ Assert(ioh->target > PGAIO_TID_INVALID && ioh->target < PGAIO_TID_COUNT);
+ Assert(ioh->op > PGAIO_OP_INVALID && ioh->op < PGAIO_OP_COUNT);
+
+ /* start with distilled result from shared callback */
+ result = ioh->distilled_result;
+
+ for (int i = ioh->num_shared_callbacks; i > 0; i--)
+ {
+ PgAioHandleCallbackID cbid = ioh->shared_callbacks[i - 1];
+ const PgAioHandleCallbacksEntry *ce = &aio_handle_cbs[cbid];
+
+ if (!ce->cb->complete_local)
+ continue;
+
+ pgaio_debug_io(DEBUG4, ioh,
+ "calling cb #%d, id %d/%s->complete_local with distilled result: status %s, id %u, error_data %d, result %d",
+ i, cbid, ce->name,
+ pgaio_result_status_string(result.status),
+ result.id, result.error_data, result.result);
+ result = ce->cb->complete_local(ioh, result);
+ }
+
+ /*
+ * Note that we don't save the result in ioh->distilled_result, the local
+ * callback's result should not ever matter to other waiters.
+ */
+ pgaio_debug_io(DEBUG3, ioh,
+ "after local completion: distilled result: (status %s, id %u, error_data %d, result %d), raw_result: %d",
+ pgaio_result_status_string(result.status),
+ result.id, result.error_data, result.result,
+ ioh->result);
+
+ END_CRIT_SECTION();
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * IO Result
+ * --------------------------------------------------------------------------------
+ */
+
+void
+pgaio_result_report(PgAioResult result, const PgAioTargetData *target_data, int elevel)
+{
+ PgAioHandleCallbackID cbid = result.id;
+ const PgAioHandleCallbacksEntry *ce = &aio_handle_cbs[cbid];
+
+ Assert(result.status != ARS_UNKNOWN);
+ Assert(result.status != ARS_OK);
+
+ if (ce->cb->report == NULL)
+ elog(ERROR, "callback %d/%s does not have report callback",
+ result.id, ce->name);
+
+ ce->cb->report(result, target_data, elevel);
+}
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index aeacc144149..4223cd1bfd6 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -14,24 +14,210 @@
#include "postgres.h"
+#include "miscadmin.h"
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
#include "storage/aio_subsys.h"
+#include "storage/ipc.h"
+#include "storage/proc.h"
+#include "storage/shmem.h"
+#include "utils/guc.h"
+static Size
+AioCtlShmemSize(void)
+{
+ Size sz;
+
+ /* pgaio_ctl itself */
+ sz = offsetof(PgAioCtl, io_handles);
+
+ return sz;
+}
+
+static uint32
+AioProcs(void)
+{
+ return MaxBackends + NUM_AUXILIARY_PROCS;
+}
+
+static Size
+AioBackendShmemSize(void)
+{
+ return mul_size(AioProcs(), sizeof(PgAioBackend));
+}
+
+static Size
+AioHandleShmemSize(void)
+{
+ Size sz;
+
+ /* ios */
+ sz = mul_size(AioProcs(),
+ mul_size(io_max_concurrency, sizeof(PgAioHandle)));
+
+ return sz;
+}
+
+static Size
+AioHandleIOVShmemSize(void)
+{
+ return mul_size(sizeof(struct iovec),
+ mul_size(mul_size(PG_IOV_MAX, AioProcs()),
+ io_max_concurrency));
+}
+
+static Size
+AioHandleDataShmemSize(void)
+{
+ return mul_size(sizeof(uint64),
+ mul_size(mul_size(PG_IOV_MAX, AioProcs()),
+ io_max_concurrency));
+}
+
+/*
+ * Choose a suitable value for io_max_concurrency.
+ *
+ * It's unlikely that we could have more IOs in flight than buffers that we
+ * would be allowed to pin.
+ *
+ * On the upper end, apply a cap too - just because shared_buffers is large,
+ * it doesn't make sense have millions of buffers undergo IO concurrently.
+ */
+static int
+AioChooseMaxConcurrency(void)
+{
+ uint32 max_backends;
+ int max_proportional_pins;
+
+ /* Similar logic to LimitAdditionalPins() */
+ max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
+ max_proportional_pins = NBuffers / max_backends;
+
+ max_proportional_pins = Max(max_proportional_pins, 1);
+
+ /* apply upper limit */
+ return Min(max_proportional_pins, 64);
+}
+
Size
AioShmemSize(void)
{
Size sz = 0;
+ /*
+ * We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
+ * However, if the DBA explicitly set wal_buffers = -1 in the config file,
+ * then PGC_S_DYNAMIC_DEFAULT will fail to override that and we must force
+ *
+ */
+ if (io_max_concurrency == -1)
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", AioChooseMaxConcurrency());
+ SetConfigOption("io_max_concurrency", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+ if (io_max_concurrency == -1) /* failed to apply it? */
+ SetConfigOption("io_max_concurrency", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+
+ sz = add_size(sz, AioCtlShmemSize());
+ sz = add_size(sz, AioBackendShmemSize());
+ sz = add_size(sz, AioHandleShmemSize());
+ sz = add_size(sz, AioHandleIOVShmemSize());
+ sz = add_size(sz, AioHandleDataShmemSize());
+
+ if (pgaio_method_ops->shmem_size)
+ sz = add_size(sz, pgaio_method_ops->shmem_size());
+
return sz;
}
void
AioShmemInit(void)
{
+ bool found;
+ uint32 io_handle_off = 0;
+ uint32 iovec_off = 0;
+ uint32 per_backend_iovecs = io_max_concurrency * PG_IOV_MAX;
+
+ pgaio_ctl = (PgAioCtl *)
+ ShmemInitStruct("AioCtl", AioCtlShmemSize(), &found);
+
+ if (found)
+ goto out;
+
+ memset(pgaio_ctl, 0, AioCtlShmemSize());
+
+ pgaio_ctl->io_handle_count = AioProcs() * io_max_concurrency;
+ pgaio_ctl->iovec_count = AioProcs() * per_backend_iovecs;
+
+ pgaio_ctl->backend_state = (PgAioBackend *)
+ ShmemInitStruct("AioBackend", AioBackendShmemSize(), &found);
+
+ pgaio_ctl->io_handles = (PgAioHandle *)
+ ShmemInitStruct("AioHandle", AioHandleShmemSize(), &found);
+
+ pgaio_ctl->iovecs = (struct iovec *)
+ ShmemInitStruct("AioHandleIOV", AioHandleIOVShmemSize(), &found);
+ pgaio_ctl->handle_data = (uint64 *)
+ ShmemInitStruct("AioHandleData", AioHandleDataShmemSize(), &found);
+
+ for (int procno = 0; procno < AioProcs(); procno++)
+ {
+ PgAioBackend *bs = &pgaio_ctl->backend_state[procno];
+
+ bs->io_handle_off = io_handle_off;
+ io_handle_off += io_max_concurrency;
+
+ dclist_init(&bs->idle_ios);
+ memset(bs->staged_ios, 0, sizeof(PgAioHandle *) * PGAIO_SUBMIT_BATCH_SIZE);
+ dclist_init(&bs->in_flight_ios);
+
+ /* initialize per-backend IOs */
+ for (int i = 0; i < io_max_concurrency; i++)
+ {
+ PgAioHandle *ioh = &pgaio_ctl->io_handles[bs->io_handle_off + i];
+
+ ioh->generation = 1;
+ ioh->owner_procno = procno;
+ ioh->iovec_off = iovec_off;
+ ioh->handle_data_len = 0;
+ ioh->report_return = NULL;
+ ioh->resowner = NULL;
+ ioh->num_shared_callbacks = 0;
+ ioh->distilled_result.status = ARS_UNKNOWN;
+ ioh->flags = 0;
+
+ ConditionVariableInit(&ioh->cv);
+
+ dclist_push_tail(&bs->idle_ios, &ioh->node);
+ iovec_off += PG_IOV_MAX;
+ }
+ }
+
+out:
+ /* Initialize IO method specific resources. */
+ if (pgaio_method_ops->shmem_init)
+ pgaio_method_ops->shmem_init(!found);
}
void
pgaio_init_backend(void)
{
+ /* shouldn't be initialized twice */
+ Assert(!pgaio_my_backend);
+
+ if (MyProc == NULL || MyProcNumber >= AioProcs())
+ elog(ERROR, "aio requires a normal PGPROC");
+
+ pgaio_my_backend = &pgaio_ctl->backend_state[MyProcNumber];
+
+ if (pgaio_method_ops->init_backend)
+ pgaio_method_ops->init_backend();
+
+ before_shmem_exit(pgaio_shutdown, 0);
}
diff --git a/src/backend/storage/aio/aio_io.c b/src/backend/storage/aio/aio_io.c
new file mode 100644
index 00000000000..89376ff4040
--- /dev/null
+++ b/src/backend/storage/aio/aio_io.c
@@ -0,0 +1,180 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_io.c
+ * AIO - Low Level IO Handling
+ *
+ * Functions related to associating IO operations to IO Handles and IO-method
+ * independent support functions for actually performing IO.
+ *
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio_io.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "miscadmin.h"
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "storage/fd.h"
+#include "utils/wait_event.h"
+
+
+static void pgaio_io_before_prep(PgAioHandle *ioh);
+
+
+
+/* --------------------------------------------------------------------------------
+ * Public IO related functions operating on IO Handles
+ * --------------------------------------------------------------------------------
+ */
+
+/*
+ * Scatter/gather IO needs to associate an iovec with the Handle. To support
+ * worker mode this data needs to be in shared memory.
+ *
+ * XXX: Right now the amount of space available for each IO is
+ * PG_IOV_MAX. While it's tempting to use the io_combine_limit GUC, that's
+ * PGC_USERSET, so we can't allocate shared memory based on that.
+ */
+int
+pgaio_io_get_iovec(PgAioHandle *ioh, struct iovec **iov)
+{
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT);
+
+ *iov = &pgaio_ctl->iovecs[ioh->iovec_off];
+
+ return PG_IOV_MAX;
+}
+
+PgAioOpData *
+pgaio_io_get_op_data(PgAioHandle *ioh)
+{
+ return &ioh->op_data;
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * "Preparation" routines for individual IO operations
+ *
+ * These are called by the code actually initiating an IO, to associate the IO
+ * specific data with an AIO handle.
+ *
+ * Each of the preparation routines first needs to call
+ * pgaio_io_before_prep(), then fill IO specific fields in the handle and then
+ * finally call pgaio_io_stage().
+ * --------------------------------------------------------------------------------
+ */
+
+void
+pgaio_io_prep_readv(PgAioHandle *ioh,
+ int fd, int iovcnt, uint64 offset)
+{
+ pgaio_io_before_prep(ioh);
+
+ ioh->op_data.read.fd = fd;
+ ioh->op_data.read.offset = offset;
+ ioh->op_data.read.iov_length = iovcnt;
+
+ pgaio_io_stage(ioh, PGAIO_OP_READV);
+}
+
+void
+pgaio_io_prep_writev(PgAioHandle *ioh,
+ int fd, int iovcnt, uint64 offset)
+{
+ pgaio_io_before_prep(ioh);
+
+ ioh->op_data.write.fd = fd;
+ ioh->op_data.write.offset = offset;
+ ioh->op_data.write.iov_length = iovcnt;
+
+ pgaio_io_stage(ioh, PGAIO_OP_WRITEV);
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Internal IO related functions operating on IO Handles
+ * --------------------------------------------------------------------------------
+ */
+
+/*
+ * Execute IO operation synchronously. This is implemented here, not in
+ * method_sync.c, because other IO methods lso might use it / fall back to it.
+ */
+void
+pgaio_io_perform_synchronously(PgAioHandle *ioh)
+{
+ ssize_t result = 0;
+ struct iovec *iov = &pgaio_ctl->iovecs[ioh->iovec_off];
+
+ START_CRIT_SECTION();
+
+ /* Perform IO. */
+ switch (ioh->op)
+ {
+ case PGAIO_OP_READV:
+ pgstat_report_wait_start(WAIT_EVENT_DATA_FILE_READ);
+ result = pg_preadv(ioh->op_data.read.fd, iov,
+ ioh->op_data.read.iov_length,
+ ioh->op_data.read.offset);
+ pgstat_report_wait_end();
+ break;
+ case PGAIO_OP_WRITEV:
+ pgstat_report_wait_start(WAIT_EVENT_DATA_FILE_WRITE);
+ result = pg_pwritev(ioh->op_data.write.fd, iov,
+ ioh->op_data.write.iov_length,
+ ioh->op_data.write.offset);
+ pgstat_report_wait_end();
+ break;
+ case PGAIO_OP_INVALID:
+ elog(ERROR, "trying to execute invalid IO operation");
+ }
+
+ ioh->result = result < 0 ? -errno : result;
+
+ pgaio_io_process_completion(ioh, ioh->result);
+
+ END_CRIT_SECTION();
+}
+
+/*
+ * Helper function to be called by IO operation preparation functions, before
+ * any data in the handle is set. Mostly to centralize assertions.
+ */
+static void
+pgaio_io_before_prep(PgAioHandle *ioh)
+{
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT);
+ Assert(pgaio_io_has_target(ioh));
+ Assert(ioh->op == PGAIO_OP_INVALID);
+}
+
+/*
+ * Could be made part of the public interface, but it's not clear there's
+ * really a use case for that.
+ */
+const char *
+pgaio_io_get_op_name(PgAioHandle *ioh)
+{
+ Assert(ioh->op >= 0 && ioh->op < PGAIO_OP_COUNT);
+
+ switch (ioh->op)
+ {
+ case PGAIO_OP_INVALID:
+ return "invalid";
+ case PGAIO_OP_READV:
+ return "read";
+ case PGAIO_OP_WRITEV:
+ return "write";
+ }
+
+ return NULL; /* silence compiler */
+}
diff --git a/src/backend/storage/aio/aio_target.c b/src/backend/storage/aio/aio_target.c
new file mode 100644
index 00000000000..15428968e58
--- /dev/null
+++ b/src/backend/storage/aio/aio_target.c
@@ -0,0 +1,108 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_target.c
+ * AIO - Functionality related to executing IO for different targets
+ *
+ * XXX Write me
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio_target.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+
+
+/*
+ * Registry for entities that can be the target of AIO.
+ *
+ * To support executing using worker processes, the file descriptor for an IO
+ * may need to be be reopened in a different process. This is done via the
+ * PgAioTargetInfo.reopen callback.
+ */
+static const PgAioTargetInfo *pgaio_target_info[] = {
+ [PGAIO_TID_INVALID] = &(PgAioTargetInfo) {
+ .name = "invalid",
+ },
+};
+
+
+
+bool
+pgaio_io_has_target(PgAioHandle *ioh)
+{
+ return ioh->target != PGAIO_TID_INVALID;
+}
+
+/*
+ * Return the name for the target associated with the IO. Mostly useful for
+ * debugging/logging.
+ */
+const char *
+pgaio_io_get_target_name(PgAioHandle *ioh)
+{
+ Assert(ioh->target >= 0 && ioh->target < PGAIO_TID_COUNT);
+
+ return pgaio_target_info[ioh->target]->name;
+}
+
+/*
+ * Assign a target to the IO.
+ *
+ * This has to be called exactly once before pgaio_io_prep_*() is called.
+ */
+void
+pgaio_io_set_target(PgAioHandle *ioh, PgAioTargetID targetid)
+{
+ Assert(ioh->state == PGAIO_HS_HANDED_OUT);
+ Assert(ioh->target == PGAIO_TID_INVALID);
+
+ ioh->target = targetid;
+}
+
+PgAioTargetData *
+pgaio_io_get_target_data(PgAioHandle *ioh)
+{
+ return &ioh->target_data;
+}
+
+/*
+ * Return a stringified description of the IO's target.
+ *
+ * The string is localized and allocated in the current memory context.
+ */
+char *
+pgaio_io_get_target_description(PgAioHandle *ioh)
+{
+ return pgaio_target_info[ioh->target]->describe_identity(&ioh->target_data);
+}
+
+/*
+ * Internal: Check if pgaio_io_reopen() is available for the IO.
+ */
+bool
+pgaio_io_can_reopen(PgAioHandle *ioh)
+{
+ return pgaio_target_info[ioh->target]->reopen != NULL;
+}
+
+/*
+ * Internal: Before executing an IO outside of the context of the process the
+ * IO has been prepared in, the file descriptor has to be reopened - any FD
+ * referenced in the IO itself, won't be valid in the separate process.
+ */
+void
+pgaio_io_reopen(PgAioHandle *ioh)
+{
+ Assert(ioh->target >= 0 && ioh->target < PGAIO_TID_COUNT);
+ Assert(ioh->op >= 0 && ioh->op < PGAIO_OP_COUNT);
+
+ pgaio_target_info[ioh->target]->reopen(ioh);
+}
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index c822fd4ddf7..2c26089d52e 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -2,6 +2,10 @@
backend_sources += files(
'aio.c',
+ 'aio_callback.c',
'aio_init.c',
+ 'aio_io.c',
+ 'aio_target.c',
+ 'method_sync.c',
'read_stream.c',
)
diff --git a/src/backend/storage/aio/method_sync.c b/src/backend/storage/aio/method_sync.c
new file mode 100644
index 00000000000..43f9c8bd0b3
--- /dev/null
+++ b/src/backend/storage/aio/method_sync.c
@@ -0,0 +1,47 @@
+/*-------------------------------------------------------------------------
+ *
+ * method_sync.c
+ * AIO - perform "AIO" by executing it synchronously
+ *
+ * This method is mainly to check if AIO use causes regressions. Other IO
+ * methods might also fall back to the synchronous method for functionality
+ * they cannot provide.
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/method_sync.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+
+static bool pgaio_sync_needs_synchronous_execution(PgAioHandle *ioh);
+static int pgaio_sync_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+
+
+const IoMethodOps pgaio_sync_ops = {
+ .needs_synchronous_execution = pgaio_sync_needs_synchronous_execution,
+ .submit = pgaio_sync_submit,
+};
+
+
+
+static bool
+pgaio_sync_needs_synchronous_execution(PgAioHandle *ioh)
+{
+ return true;
+}
+
+static int
+pgaio_sync_submit(uint16 num_staged_ios, PgAioHandle **staged_ios)
+{
+ elog(ERROR, "should be unreachable");
+
+ return 0;
+}
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index e199f071628..6f3ca878bd1 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -191,6 +191,7 @@ ABI_compatibility:
Section: ClassName - WaitEventIO
+AIO_IO_COMPLETION "Waiting for IO completion."
BASEBACKUP_READ "Waiting for base backup to read from a file."
BASEBACKUP_SYNC "Waiting for data written by a base backup to reach durable storage."
BASEBACKUP_WRITE "Waiting for base backup to write to a file."
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 406c0893440..14a338e4308 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1268,6 +1268,7 @@ InvalMessageArray
InvalidationInfo
InvalidationMsgsGroup
IoMethod
+IoMethodOps
IpcMemoryId
IpcMemoryKey
IpcMemoryState
@@ -2108,6 +2109,26 @@ Permutation
PermutationStep
PermutationStepBlocker
PermutationStepBlockerType
+PgAioBackend
+PgAioCtl
+PgAioHandle
+PgAioHandleCallbackID
+PgAioHandleCallbackStage
+PgAioHandleCallbackComplete
+PgAioHandleCallbackReport
+PgAioHandleCallbacks
+PgAioHandleCallbacksEntry
+PgAioHandleFlags
+PgAioHandleState
+PgAioOp
+PgAioOpData
+PgAioResult
+PgAioResultStatus
+PgAioReturn
+PgAioTargetData
+PgAioTargetID
+PgAioTargetInfo
+PgAioWaitRef
PgArchData
PgBackendGSSStatus
PgBackendSSLStatus
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0011-aio-Skeleton-IO-worker-infrastructure.patch (20.5K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/12-v2.4-0011-aio-Skeleton-IO-worker-infrastructure.patch)
download | inline diff:
From 27bdac00ef8da7a01c401f2f10b4017c9ee0b1ba Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 17 Feb 2025 17:57:00 -0500
Subject: [PATCH v2.4 11/29] aio: Skeleton IO worker infrastructure
This doesn't do anything useful on its own, but the code that needs to be
touched is independent of other changes.
Remarks:
- dynamic increase / decrease of workers based on IO load
Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/miscadmin.h | 2 +
src/include/postmaster/postmaster.h | 1 +
src/include/storage/aio_subsys.h | 4 +
src/include/storage/io_worker.h | 22 +++
src/include/storage/proc.h | 4 +-
src/backend/postmaster/launch_backend.c | 2 +
src/backend/postmaster/pmchild.c | 1 +
src/backend/postmaster/postmaster.c | 165 ++++++++++++++++--
src/backend/storage/aio/Makefile | 1 +
src/backend/storage/aio/meson.build | 1 +
src/backend/storage/aio/method_worker.c | 88 ++++++++++
src/backend/tcop/postgres.c | 2 +
src/backend/utils/activity/pgstat_backend.c | 1 +
src/backend/utils/activity/pgstat_io.c | 1 +
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/init/miscinit.c | 3 +
16 files changed, 286 insertions(+), 13 deletions(-)
create mode 100644 src/include/storage/io_worker.h
create mode 100644 src/backend/storage/aio/method_worker.c
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index a2b63495eec..54429e046a9 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -360,6 +360,7 @@ typedef enum BackendType
B_ARCHIVER,
B_BG_WRITER,
B_CHECKPOINTER,
+ B_IO_WORKER,
B_STARTUP,
B_WAL_RECEIVER,
B_WAL_SUMMARIZER,
@@ -389,6 +390,7 @@ extern PGDLLIMPORT BackendType MyBackendType;
#define AmWalReceiverProcess() (MyBackendType == B_WAL_RECEIVER)
#define AmWalSummarizerProcess() (MyBackendType == B_WAL_SUMMARIZER)
#define AmWalWriterProcess() (MyBackendType == B_WAL_WRITER)
+#define AmIoWorkerProcess() (MyBackendType == B_IO_WORKER)
#define AmSpecialWorkerProcess() \
(AmAutoVacuumLauncherProcess() || \
diff --git a/src/include/postmaster/postmaster.h b/src/include/postmaster/postmaster.h
index 188a06e2379..253dc98c50e 100644
--- a/src/include/postmaster/postmaster.h
+++ b/src/include/postmaster/postmaster.h
@@ -98,6 +98,7 @@ extern void InitProcessGlobals(void);
extern int MaxLivePostmasterChildren(void);
extern bool PostmasterMarkPIDForWorkerNotify(int);
+extern void assign_io_workers(int newval, void *extra);
#ifdef WIN32
extern void pgwin32_register_deadchild_callback(HANDLE procHandle, DWORD procId);
diff --git a/src/include/storage/aio_subsys.h b/src/include/storage/aio_subsys.h
index e4faf692a38..ed00d5c47cd 100644
--- a/src/include/storage/aio_subsys.h
+++ b/src/include/storage/aio_subsys.h
@@ -30,4 +30,8 @@ extern void pgaio_init_backend(void);
extern void pgaio_error_cleanup(void);
extern void AtEOXact_Aio(bool is_commit);
+
+/* aio_worker.c */
+extern bool pgaio_workers_enabled(void);
+
#endif /* AIO_SUBSYS_H */
diff --git a/src/include/storage/io_worker.h b/src/include/storage/io_worker.h
new file mode 100644
index 00000000000..223d614dc4a
--- /dev/null
+++ b/src/include/storage/io_worker.h
@@ -0,0 +1,22 @@
+/*-------------------------------------------------------------------------
+ *
+ * io_worker.h
+ * IO worker for implementing AIO "ourselves"
+ *
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/io.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef IO_WORKER_H
+#define IO_WORKER_H
+
+
+extern void IoWorkerMain(char *startup_data, size_t startup_data_len) pg_attribute_noreturn();
+
+extern int io_workers;
+
+#endif /* IO_WORKER_H */
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 20777f7d5ae..64e9b8ff8c5 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -448,7 +448,9 @@ extern PGDLLIMPORT PGPROC *PreparedXactProcs;
* 2 slots, but WAL writer is launched only after startup has exited, so we
* only need 6 slots.
*/
-#define NUM_AUXILIARY_PROCS 6
+#define MAX_IO_WORKERS 32
+#define NUM_AUXILIARY_PROCS (6 + MAX_IO_WORKERS)
+
/* configurable options */
extern PGDLLIMPORT int DeadlockTimeout;
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index a97a1eda6da..54b4c22bd63 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -48,6 +48,7 @@
#include "replication/slotsync.h"
#include "replication/walreceiver.h"
#include "storage/dsm.h"
+#include "storage/io_worker.h"
#include "storage/pg_shmem.h"
#include "tcop/backend_startup.h"
#include "utils/memutils.h"
@@ -197,6 +198,7 @@ static child_process_kind child_process_kinds[] = {
[B_ARCHIVER] = {"archiver", PgArchiverMain, true},
[B_BG_WRITER] = {"bgwriter", BackgroundWriterMain, true},
[B_CHECKPOINTER] = {"checkpointer", CheckpointerMain, true},
+ [B_IO_WORKER] = {"io_worker", IoWorkerMain, true},
[B_STARTUP] = {"startup", StartupProcessMain, true},
[B_WAL_RECEIVER] = {"wal_receiver", WalReceiverMain, true},
[B_WAL_SUMMARIZER] = {"wal_summarizer", WalSummarizerMain, true},
diff --git a/src/backend/postmaster/pmchild.c b/src/backend/postmaster/pmchild.c
index 0d473226c3a..cde1d23a4ca 100644
--- a/src/backend/postmaster/pmchild.c
+++ b/src/backend/postmaster/pmchild.c
@@ -101,6 +101,7 @@ InitPostmasterChildSlots(void)
pmchild_pools[B_AUTOVAC_WORKER].size = autovacuum_worker_slots;
pmchild_pools[B_BG_WORKER].size = max_worker_processes;
+ pmchild_pools[B_IO_WORKER].size = MAX_IO_WORKERS;
/*
* There can be only one of each of these running at a time. They each
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index bb22b13adef..abef9d941c0 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -108,9 +108,12 @@
#include "replication/logicallauncher.h"
#include "replication/slotsync.h"
#include "replication/walsender.h"
+#include "storage/aio_subsys.h"
#include "storage/fd.h"
+#include "storage/io_worker.h"
#include "storage/ipc.h"
#include "storage/pmsignal.h"
+#include "storage/proc.h"
#include "tcop/backend_startup.h"
#include "tcop/tcopprot.h"
#include "utils/datetime.h"
@@ -341,6 +344,7 @@ typedef enum
* ckpt */
PM_WAIT_XLOG_ARCHIVAL, /* waiting for archiver and walsenders to
* finish */
+ PM_WAIT_IO_WORKERS, /* waiting for io workers to exit */
PM_WAIT_CHECKPOINTER, /* waiting for checkpointer to shut down */
PM_WAIT_DEAD_END, /* waiting for dead-end children to exit */
PM_NO_CHILDREN, /* all important children have exited */
@@ -403,6 +407,10 @@ bool LoadedSSL = false;
static DNSServiceRef bonjour_sdref = NULL;
#endif
+/* State for IO worker management. */
+static int io_worker_count = 0;
+static PMChild *io_worker_children[MAX_IO_WORKERS];
+
/*
* postmaster.c - function prototypes
*/
@@ -437,6 +445,8 @@ static void TerminateChildren(int signal);
static int CountChildren(BackendTypeMask targetMask);
static void LaunchMissingBackgroundProcesses(void);
static void maybe_start_bgworkers(void);
+static bool maybe_reap_io_worker(int pid);
+static void maybe_adjust_io_workers(void);
static bool CreateOptsFile(int argc, char *argv[], char *fullprogname);
static PMChild *StartChildProcess(BackendType type);
static void StartSysLogger(void);
@@ -1366,6 +1376,11 @@ PostmasterMain(int argc, char *argv[])
*/
AddToDataDirLockFile(LOCK_FILE_LINE_PM_STATUS, PM_STATUS_STARTING);
+ UpdatePMState(PM_STARTUP);
+
+ /* Make sure we can perform I/O while starting up. */
+ maybe_adjust_io_workers();
+
/* Start bgwriter and checkpointer so they can help with recovery */
if (CheckpointerPMChild == NULL)
CheckpointerPMChild = StartChildProcess(B_CHECKPOINTER);
@@ -1378,7 +1393,6 @@ PostmasterMain(int argc, char *argv[])
StartupPMChild = StartChildProcess(B_STARTUP);
Assert(StartupPMChild != NULL);
StartupStatus = STARTUP_RUNNING;
- UpdatePMState(PM_STARTUP);
/* Some workers may be scheduled to start now */
maybe_start_bgworkers();
@@ -2503,6 +2517,16 @@ process_pm_child_exit(void)
continue;
}
+ /* Was it an IO worker? */
+ if (maybe_reap_io_worker(pid))
+ {
+ if (!EXIT_STATUS_0(exitstatus) && !EXIT_STATUS_1(exitstatus))
+ HandleChildCrash(pid, exitstatus, _("io worker"));
+
+ maybe_adjust_io_workers();
+ continue;
+ }
+
/*
* Was it a backend or a background worker?
*/
@@ -2724,6 +2748,7 @@ HandleFatalError(QuitSignalReason reason, bool consider_sigabrt)
case PM_WAIT_XLOG_SHUTDOWN:
case PM_WAIT_XLOG_ARCHIVAL:
case PM_WAIT_CHECKPOINTER:
+ case PM_WAIT_IO_WORKERS:
/*
* NB: Similar code exists in PostmasterStateMachine()'s handling
@@ -2906,20 +2931,21 @@ PostmasterStateMachine(void)
/*
* If we are doing crash recovery or an immediate shutdown then we
- * expect archiver, checkpointer and walsender to exit as well,
- * otherwise not.
+ * expect archiver, checkpointer, io workers and walsender to exit as
+ * well, otherwise not.
*/
if (FatalError || Shutdown >= ImmediateShutdown)
targetMask = btmask_add(targetMask,
B_CHECKPOINTER,
B_ARCHIVER,
+ B_IO_WORKER,
B_WAL_SENDER);
/*
- * Normally walsenders and archiver will continue running; they will
- * be terminated later after writing the checkpoint record. We also
- * let dead-end children to keep running for now. The syslogger
- * process exits last.
+ * Normally archiver, checkpointer, IO workers and walsenders will
+ * continue running; they will be terminated later after writing the
+ * checkpoint record. We also let dead-end children to keep running
+ * for now. The syslogger process exits last.
*
* This assertion checks that we have covered all backend types,
* either by including them in targetMask, or by noting here that they
@@ -2934,12 +2960,13 @@ PostmasterStateMachine(void)
B_LOGGER);
/*
- * Archiver, checkpointer and walsender may or may not be in
- * targetMask already.
+ * Archiver, checkpointer, IO workers, and walsender may or may
+ * not be in targetMask already.
*/
remainMask = btmask_add(remainMask,
B_ARCHIVER,
B_CHECKPOINTER,
+ B_IO_WORKER,
B_WAL_SENDER);
/* these are not real postmaster children */
@@ -3040,11 +3067,25 @@ PostmasterStateMachine(void)
{
/*
* PM_WAIT_XLOG_ARCHIVAL state ends when there are no children other
- * than checkpointer, dead-end children and logger left. There
+ * than checkpointer, io workers and dead-end children left. There
* shouldn't be any regular backends left by now anyway; what we're
* really waiting for is for walsenders and archiver to exit.
*/
- if (CountChildren(btmask_all_except(B_CHECKPOINTER, B_LOGGER, B_DEAD_END_BACKEND)) == 0)
+ if (CountChildren(btmask_all_except(B_CHECKPOINTER, B_IO_WORKER,
+ B_LOGGER, B_DEAD_END_BACKEND)) == 0)
+ {
+ UpdatePMState(PM_WAIT_IO_WORKERS);
+ SignalChildren(SIGUSR2, btmask(B_IO_WORKER));
+ }
+ }
+
+ if (pmState == PM_WAIT_IO_WORKERS)
+ {
+ /*
+ * PM_WAIT_IO_WORKERS state ends when there's only checkpointer and
+ * dead_end children left.
+ */
+ if (io_worker_count == 0)
{
UpdatePMState(PM_WAIT_CHECKPOINTER);
@@ -3172,10 +3213,14 @@ PostmasterStateMachine(void)
/* re-create shared memory and semaphores */
CreateSharedMemoryAndSemaphores();
+ UpdatePMState(PM_STARTUP);
+
+ /* Make sure we can perform I/O while starting up. */
+ maybe_adjust_io_workers();
+
StartupPMChild = StartChildProcess(B_STARTUP);
Assert(StartupPMChild != NULL);
StartupStatus = STARTUP_RUNNING;
- UpdatePMState(PM_STARTUP);
/* crash recovery started, reset SIGKILL flag */
AbortStartTime = 0;
@@ -3199,6 +3244,7 @@ pmstate_name(PMState state)
PM_TOSTR_CASE(PM_WAIT_BACKENDS);
PM_TOSTR_CASE(PM_WAIT_XLOG_SHUTDOWN);
PM_TOSTR_CASE(PM_WAIT_XLOG_ARCHIVAL);
+ PM_TOSTR_CASE(PM_WAIT_IO_WORKERS);
PM_TOSTR_CASE(PM_WAIT_DEAD_END);
PM_TOSTR_CASE(PM_WAIT_CHECKPOINTER);
PM_TOSTR_CASE(PM_NO_CHILDREN);
@@ -4115,6 +4161,7 @@ bgworker_should_start_now(BgWorkerStartTime start_time)
case PM_WAIT_DEAD_END:
case PM_WAIT_XLOG_ARCHIVAL:
case PM_WAIT_XLOG_SHUTDOWN:
+ case PM_WAIT_IO_WORKERS:
case PM_WAIT_BACKENDS:
case PM_STOP_BACKENDS:
break;
@@ -4265,6 +4312,100 @@ maybe_start_bgworkers(void)
}
}
+static bool
+maybe_reap_io_worker(int pid)
+{
+ for (int id = 0; id < MAX_IO_WORKERS; ++id)
+ {
+ if (io_worker_children[id] &&
+ io_worker_children[id]->pid == pid)
+ {
+ ReleasePostmasterChildSlot(io_worker_children[id]);
+
+ --io_worker_count;
+ io_worker_children[id] = NULL;
+ return true;
+ }
+ }
+ return false;
+}
+
+static void
+maybe_adjust_io_workers(void)
+{
+ if (!pgaio_workers_enabled())
+ return;
+
+ /*
+ * If we're in final shutting down state, then we're just waiting for all
+ * processes to exit.
+ */
+ if (pmState >= PM_WAIT_IO_WORKERS)
+ return;
+
+ /* Don't start new workers during an immediate shutdown either. */
+ if (Shutdown >= ImmediateShutdown)
+ return;
+
+ /*
+ * Don't start new workers if we're in the shutdown phase of a crash
+ * restart. But we *do* need to start if we're already starting up again.
+ */
+ if (FatalError && pmState >= PM_STOP_BACKENDS)
+ return;
+
+ Assert(pmState < PM_WAIT_IO_WORKERS);
+
+ /* Not enough running? */
+ while (io_worker_count < io_workers)
+ {
+ PMChild *child;
+ int id;
+
+ /* find unused entry in io_worker_children array */
+ for (id = 0; id < MAX_IO_WORKERS; ++id)
+ {
+ if (io_worker_children[id] == NULL)
+ break;
+ }
+ if (id == MAX_IO_WORKERS)
+ elog(ERROR, "could not find a free IO worker ID");
+
+ /* Try to launch one. */
+ child = StartChildProcess(B_IO_WORKER);
+ if (child != NULL)
+ {
+ io_worker_children[id] = child;
+ ++io_worker_count;
+ }
+ else
+ break; /* XXX try again soon? */
+ }
+
+ /* Too many running? */
+ if (io_worker_count > io_workers)
+ {
+ /* ask the IO worker in the highest slot to exit */
+ for (int id = MAX_IO_WORKERS - 1; id >= 0; --id)
+ {
+ if (io_worker_children[id] != NULL)
+ {
+ kill(io_worker_children[id]->pid, SIGUSR2);
+ break;
+ }
+ }
+ }
+}
+
+void
+assign_io_workers(int newval, void *extra)
+{
+ io_workers = newval;
+ if (!IsUnderPostmaster && pmState > PM_INIT)
+ maybe_adjust_io_workers();
+}
+
+
/*
* When a backend asks to be notified about worker state changes, we
* set a flag in its backend entry. The background worker machinery needs
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index 89f821ea7e1..f51c34a37f8 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -15,6 +15,7 @@ OBJS = \
aio_io.o \
aio_target.o \
method_sync.o \
+ method_worker.o \
read_stream.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 2c26089d52e..74f94c6e40b 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -7,5 +7,6 @@ backend_sources += files(
'aio_io.c',
'aio_target.c',
'method_sync.c',
+ 'method_worker.c',
'read_stream.c',
)
diff --git a/src/backend/storage/aio/method_worker.c b/src/backend/storage/aio/method_worker.c
new file mode 100644
index 00000000000..942f1609af2
--- /dev/null
+++ b/src/backend/storage/aio/method_worker.c
@@ -0,0 +1,88 @@
+/*-------------------------------------------------------------------------
+ *
+ * method_worker.c
+ * AIO - perform AIO using worker processes
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/method_worker.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "libpq/pqsignal.h"
+#include "miscadmin.h"
+#include "postmaster/auxprocess.h"
+#include "postmaster/interrupt.h"
+#include "storage/aio_subsys.h"
+#include "storage/io_worker.h"
+#include "storage/ipc.h"
+#include "storage/latch.h"
+#include "storage/proc.h"
+#include "tcop/tcopprot.h"
+#include "utils/wait_event.h"
+
+
+/* GUCs */
+int io_workers = 3;
+
+
+void
+IoWorkerMain(char *startup_data, size_t startup_data_len)
+{
+ sigjmp_buf local_sigjmp_buf;
+
+ MyBackendType = B_IO_WORKER;
+ AuxiliaryProcessMainCommon();
+
+ /* TODO review all signals */
+ pqsignal(SIGHUP, SignalHandlerForConfigReload);
+ pqsignal(SIGINT, die); /* to allow manually triggering worker restart */
+
+ /*
+ * Ignore SIGTERM, will get explicit shutdown via SIGUSR2 later in the
+ * shutdown sequence, similar to checkpointer.
+ */
+ pqsignal(SIGTERM, SIG_IGN);
+ /* SIGQUIT handler was already set up by InitPostmasterChild */
+ pqsignal(SIGALRM, SIG_IGN);
+ pqsignal(SIGPIPE, SIG_IGN);
+ pqsignal(SIGUSR1, procsignal_sigusr1_handler);
+ pqsignal(SIGUSR2, SignalHandlerForShutdownRequest);
+ sigprocmask(SIG_SETMASK, &UnBlockSig, NULL);
+
+ /* see PostgresMain() */
+ if (sigsetjmp(local_sigjmp_buf, 1) != 0)
+ {
+ error_context_stack = NULL;
+ HOLD_INTERRUPTS();
+
+ EmitErrorReport();
+
+ proc_exit(1);
+ }
+
+ /* We can now handle ereport(ERROR) */
+ PG_exception_stack = &local_sigjmp_buf;
+
+ while (!ShutdownRequestPending)
+ {
+ WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, -1,
+ WAIT_EVENT_IO_WORKER_MAIN);
+ ResetLatch(MyLatch);
+ CHECK_FOR_INTERRUPTS();
+ }
+
+ proc_exit(0);
+}
+
+bool
+pgaio_workers_enabled(void)
+{
+ /* placeholder for future commit */
+ return false;
+}
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 1149d89d7a1..bf6f204e71c 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3313,6 +3313,8 @@ ProcessInterrupts(void)
(errcode(ERRCODE_ADMIN_SHUTDOWN),
errmsg("terminating background worker \"%s\" due to administrator command",
MyBgworkerEntry->bgw_type)));
+ else if (AmIoWorkerProcess())
+ proc_exit(0);
else
ereport(FATAL,
(errcode(ERRCODE_ADMIN_SHUTDOWN),
diff --git a/src/backend/utils/activity/pgstat_backend.c b/src/backend/utils/activity/pgstat_backend.c
index 4a667e7019c..d7f7bc8d2e3 100644
--- a/src/backend/utils/activity/pgstat_backend.c
+++ b/src/backend/utils/activity/pgstat_backend.c
@@ -238,6 +238,7 @@ pgstat_tracks_backend_bktype(BackendType bktype)
case B_LOGGER:
case B_BG_WRITER:
case B_CHECKPOINTER:
+ case B_IO_WORKER:
case B_STARTUP:
return false;
diff --git a/src/backend/utils/activity/pgstat_io.c b/src/backend/utils/activity/pgstat_io.c
index 28a431084b8..713b1dc6cd0 100644
--- a/src/backend/utils/activity/pgstat_io.c
+++ b/src/backend/utils/activity/pgstat_io.c
@@ -375,6 +375,7 @@ pgstat_tracks_io_bktype(BackendType bktype)
case B_BG_WORKER:
case B_BG_WRITER:
case B_CHECKPOINTER:
+ case B_IO_WORKER:
case B_SLOTSYNC_WORKER:
case B_STANDALONE_BACKEND:
case B_STARTUP:
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 6f3ca878bd1..d18b788b8a1 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -57,6 +57,7 @@ BGWRITER_HIBERNATE "Waiting in background writer process, hibernating."
BGWRITER_MAIN "Waiting in main loop of background writer process."
CHECKPOINTER_MAIN "Waiting in main loop of checkpointer process."
CHECKPOINTER_SHUTDOWN "Waiting for checkpointer process to be terminated."
+IO_WORKER_MAIN "Waiting in main loop of IO Worker process."
LOGICAL_APPLY_MAIN "Waiting in main loop of logical replication apply process."
LOGICAL_LAUNCHER_MAIN "Waiting in main loop of logical replication launcher process."
LOGICAL_PARALLEL_APPLY_MAIN "Waiting in main loop of logical replication parallel apply process."
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index 0347fc11092..cbca090d2b0 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -293,6 +293,9 @@ GetBackendTypeDesc(BackendType backendType)
case B_CHECKPOINTER:
backendDesc = gettext_noop("checkpointer");
break;
+ case B_IO_WORKER:
+ backendDesc = "io worker";
+ break;
case B_LOGGER:
backendDesc = gettext_noop("logger");
break;
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0012-aio-Add-worker-method.patch (21.1K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/13-v2.4-0012-aio-Add-worker-method.patch)
download | inline diff:
From 0d79ede7a92225d90715a5370581c83b521cb63f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 31 Jan 2025 13:46:35 -0500
Subject: [PATCH v2.4 12/29] aio: Add worker method
---
src/include/storage/aio.h | 5 +-
src/include/storage/aio_internal.h | 1 +
src/include/storage/lwlocklist.h | 1 +
src/backend/storage/aio/aio.c | 2 +
src/backend/storage/aio/aio_init.c | 9 +
src/backend/storage/aio/method_worker.c | 435 +++++++++++++++++-
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/misc/guc_tables.c | 13 +
src/backend/utils/misc/postgresql.conf.sample | 3 +-
doc/src/sgml/config.sgml | 21 +
src/tools/pgindent/typedefs.list | 3 +
11 files changed, 483 insertions(+), 11 deletions(-)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index d87cfe96b20..ca0f9b64d97 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -23,10 +23,11 @@
typedef enum IoMethod
{
IOMETHOD_SYNC = 0,
+ IOMETHOD_WORKER,
} IoMethod;
-/* We'll default to synchronous execution. */
-#define DEFAULT_IO_METHOD IOMETHOD_SYNC
+/* We'll default to worker based execution. */
+#define DEFAULT_IO_METHOD IOMETHOD_WORKER
/*
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index e980b06c1f3..c6e7306ed61 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -338,6 +338,7 @@ extern PgAioHandle *pgaio_inj_io_get(void);
/* Declarations for the tables of function pointers exposed by each IO method. */
extern PGDLLIMPORT const IoMethodOps pgaio_sync_ops;
+extern PGDLLIMPORT const IoMethodOps pgaio_worker_ops;
extern PGDLLIMPORT const IoMethodOps *pgaio_method_ops;
extern PGDLLIMPORT PgAioCtl *pgaio_ctl;
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index cf565452382..932024b1b0b 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -83,3 +83,4 @@ PG_LWLOCK(49, WALSummarizer)
PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
+PG_LWLOCK(53, AioWorkerSubmissionQueue)
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 0483580d644..3eace131de2 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -63,6 +63,7 @@ static void pgaio_io_wait(PgAioHandle *ioh, uint64 ref_generation);
/* Options for io_method. */
const struct config_enum_entry io_method_options[] = {
{"sync", IOMETHOD_SYNC, false},
+ {"worker", IOMETHOD_WORKER, false},
{NULL, 0, false}
};
@@ -79,6 +80,7 @@ PgAioBackend *pgaio_my_backend;
static const IoMethodOps *const pgaio_method_ops_table[] = {
[IOMETHOD_SYNC] = &pgaio_sync_ops,
+ [IOMETHOD_WORKER] = &pgaio_worker_ops,
};
/* callbacks for the configured io_method, set by assign_io_method */
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index 4223cd1bfd6..87eac5e961c 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -18,6 +18,7 @@
#include "storage/aio.h"
#include "storage/aio_internal.h"
#include "storage/aio_subsys.h"
+#include "storage/io_worker.h"
#include "storage/ipc.h"
#include "storage/proc.h"
#include "storage/shmem.h"
@@ -39,6 +40,11 @@ AioCtlShmemSize(void)
static uint32
AioProcs(void)
{
+ /*
+ * While AIO workers don't need their own AIO context, we can't currently
+ * guarantee nothing gets assigned to the a ProcNumber for an IO worker if
+ * we just subtracted MAX_IO_WORKERS.
+ */
return MaxBackends + NUM_AUXILIARY_PROCS;
}
@@ -211,6 +217,9 @@ pgaio_init_backend(void)
/* shouldn't be initialized twice */
Assert(!pgaio_my_backend);
+ if (MyBackendType == B_IO_WORKER)
+ return;
+
if (MyProc == NULL || MyProcNumber >= AioProcs())
elog(ERROR, "aio requires a normal PGPROC");
diff --git a/src/backend/storage/aio/method_worker.c b/src/backend/storage/aio/method_worker.c
index 942f1609af2..d1b52cba2cd 100644
--- a/src/backend/storage/aio/method_worker.c
+++ b/src/backend/storage/aio/method_worker.c
@@ -3,6 +3,21 @@
* method_worker.c
* AIO - perform AIO using worker processes
*
+ * Worker processes consume IOs from a shared memory submission queue, run
+ * traditional synchronous system calls, and perform the shared completion
+ * handling immediately. Client code submits most requests by pushing IOs
+ * into the submission queue, and waits (if necessary) using condition
+ * variables. Some IOs cannot be performed in another process due to lack of
+ * infrastructure for reopening the file, and must processed synchronously by
+ * the client code when submitted.
+ *
+ * So that the submitter can make just one system call when submitting a batch
+ * of IOs, wakeups "fan out"; each woken backend can wake two more. XXX This
+ * could be improved by using futexes instead of latches to wake N waiters.
+ *
+ * This method of AIO is available in all builds on all operating systems, and
+ * is the default.
+ *
* Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
* Portions Copyright (c) 1994, Regents of the University of California
*
@@ -16,30 +31,329 @@
#include "libpq/pqsignal.h"
#include "miscadmin.h"
+#include "port/pg_bitutils.h"
#include "postmaster/auxprocess.h"
#include "postmaster/interrupt.h"
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
#include "storage/aio_subsys.h"
#include "storage/io_worker.h"
#include "storage/ipc.h"
#include "storage/latch.h"
#include "storage/proc.h"
#include "tcop/tcopprot.h"
+#include "utils/ps_status.h"
#include "utils/wait_event.h"
+/* How many workers should each worker wake up if needed? */
+#define IO_WORKER_WAKEUP_FANOUT 2
+
+
+typedef struct AioWorkerSubmissionQueue
+{
+ uint32 size;
+ uint32 mask;
+ uint32 head;
+ uint32 tail;
+ uint32 ios[FLEXIBLE_ARRAY_MEMBER];
+} AioWorkerSubmissionQueue;
+
+typedef struct AioWorkerSlot
+{
+ Latch *latch;
+ bool in_use;
+} AioWorkerSlot;
+
+typedef struct AioWorkerControl
+{
+ uint64 idle_worker_mask;
+ AioWorkerSlot workers[FLEXIBLE_ARRAY_MEMBER];
+} AioWorkerControl;
+
+
+static size_t pgaio_worker_shmem_size(void);
+static void pgaio_worker_shmem_init(bool first_time);
+
+static bool pgaio_worker_needs_synchronous_execution(PgAioHandle *ioh);
+static int pgaio_worker_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+
+
+const IoMethodOps pgaio_worker_ops = {
+ .shmem_size = pgaio_worker_shmem_size,
+ .shmem_init = pgaio_worker_shmem_init,
+
+ .needs_synchronous_execution = pgaio_worker_needs_synchronous_execution,
+ .submit = pgaio_worker_submit,
+};
+
+
/* GUCs */
int io_workers = 3;
+static int io_worker_queue_size = 64;
+static int MyIoWorkerId;
+static AioWorkerSubmissionQueue *io_worker_submission_queue;
+static AioWorkerControl *io_worker_control;
+
+
+static size_t
+pgaio_worker_shmem_size(void)
+{
+ return
+ offsetof(AioWorkerSubmissionQueue, ios) +
+ sizeof(uint32) * MAX_IO_WORKERS * io_max_concurrency +
+ offsetof(AioWorkerControl, workers) +
+ sizeof(AioWorkerSlot) * MAX_IO_WORKERS;
+}
+
+static void
+pgaio_worker_shmem_init(bool first_time)
+{
+ bool found;
+ int size;
+
+ /* Round size up to next power of two so we can make a mask. */
+ size = pg_nextpower2_32(io_worker_queue_size);
+
+ io_worker_submission_queue =
+ ShmemInitStruct("AioWorkerSubmissionQueue",
+ offsetof(AioWorkerSubmissionQueue, ios) +
+ sizeof(uint32) * size,
+ &found);
+ if (!found)
+ {
+ io_worker_submission_queue->size = size;
+ io_worker_submission_queue->head = 0;
+ io_worker_submission_queue->tail = 0;
+ }
+
+ io_worker_control =
+ ShmemInitStruct("AioWorkerControl",
+ offsetof(AioWorkerControl, workers) +
+ sizeof(AioWorkerSlot) * io_workers,
+ &found);
+ if (!found)
+ {
+ io_worker_control->idle_worker_mask = 0;
+ for (int i = 0; i < io_workers; ++i)
+ {
+ io_worker_control->workers[i].latch = NULL;
+ io_worker_control->workers[i].in_use = false;
+ }
+ }
+}
+
+static int
+pgaio_choose_idle_worker(void)
+{
+ int worker;
+
+ if (io_worker_control->idle_worker_mask == 0)
+ return -1;
+
+ /* Find the lowest bit position, and clear it. */
+ worker = pg_rightmost_one_pos64(io_worker_control->idle_worker_mask);
+ io_worker_control->idle_worker_mask &= ~(UINT64_C(1) << worker);
+
+ return worker;
+}
+
+static bool
+pgaio_worker_submission_queue_insert(PgAioHandle *ioh)
+{
+ AioWorkerSubmissionQueue *queue;
+ uint32 new_head;
+
+ queue = io_worker_submission_queue;
+ new_head = (queue->head + 1) & (queue->size - 1);
+ if (new_head == queue->tail)
+ {
+ pgaio_debug(DEBUG1, "io queue is full, at %u elements",
+ io_worker_submission_queue->size);
+ return false; /* full */
+ }
+
+ queue->ios[queue->head] = pgaio_io_get_id(ioh);
+ queue->head = new_head;
+
+ return true;
+}
+
+static uint32
+pgaio_worker_submission_queue_consume(void)
+{
+ AioWorkerSubmissionQueue *queue;
+ uint32 result;
+
+ queue = io_worker_submission_queue;
+ if (queue->tail == queue->head)
+ return UINT32_MAX; /* empty */
+
+ result = queue->ios[queue->tail];
+ queue->tail = (queue->tail + 1) & (queue->size - 1);
+
+ return result;
+}
+
+static uint32
+pgaio_worker_submission_queue_depth(void)
+{
+ uint32 head;
+ uint32 tail;
+
+ head = io_worker_submission_queue->head;
+ tail = io_worker_submission_queue->tail;
+
+ if (tail > head)
+ head += io_worker_submission_queue->size;
+
+ Assert(head >= tail);
+
+ return head - tail;
+}
+
+static bool
+pgaio_worker_needs_synchronous_execution(PgAioHandle *ioh)
+{
+ return
+ !IsUnderPostmaster
+ || ioh->flags & PGAIO_HF_REFERENCES_LOCAL
+ || !pgaio_io_can_reopen(ioh);
+}
+
+static void
+pgaio_worker_submit_internal(int nios, PgAioHandle *ios[])
+{
+ PgAioHandle *synchronous_ios[PGAIO_SUBMIT_BATCH_SIZE];
+ int nsync = 0;
+ Latch *wakeup = NULL;
+ int worker;
+
+ Assert(nios <= PGAIO_SUBMIT_BATCH_SIZE);
+
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+ for (int i = 0; i < nios; ++i)
+ {
+ Assert(!pgaio_worker_needs_synchronous_execution(ios[i]));
+ if (!pgaio_worker_submission_queue_insert(ios[i]))
+ {
+ /*
+ * We'll do it synchronously, but only after we've sent as many as
+ * we can to workers, to maximize concurrency.
+ */
+ synchronous_ios[nsync++] = ios[i];
+ continue;
+ }
+
+ if (wakeup == NULL)
+ {
+ /* Choose an idle worker to wake up if we haven't already. */
+ worker = pgaio_choose_idle_worker();
+ if (worker >= 0)
+ wakeup = io_worker_control->workers[worker].latch;
+
+ pgaio_debug_io(DEBUG4, ios[i],
+ "choosing worker %d",
+ worker);
+ }
+ }
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+
+ if (wakeup)
+ SetLatch(wakeup);
+
+ /* Run whatever is left synchronously. */
+ if (nsync > 0)
+ {
+ for (int i = 0; i < nsync; ++i)
+ {
+ pgaio_io_perform_synchronously(synchronous_ios[i]);
+ }
+ }
+}
+
+static int
+pgaio_worker_submit(uint16 num_staged_ios, PgAioHandle **staged_ios)
+{
+ for (int i = 0; i < num_staged_ios; i++)
+ {
+ PgAioHandle *ioh = staged_ios[i];
+
+ pgaio_io_prepare_submit(ioh);
+ }
+
+ pgaio_worker_submit_internal(num_staged_ios, staged_ios);
+
+ return num_staged_ios;
+}
+
+/*
+ * on_shmem_exit() callback that releases the worker's slot in
+ * io_worker_control.
+ */
+static void
+pgaio_worker_die(int code, Datum arg)
+{
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+ Assert(io_worker_control->workers[MyIoWorkerId].in_use);
+ Assert(io_worker_control->workers[MyIoWorkerId].latch == MyLatch);
+
+ io_worker_control->workers[MyIoWorkerId].in_use = false;
+ io_worker_control->workers[MyIoWorkerId].latch = NULL;
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+}
+
+/*
+ * Register the worker in shared memory, assign MyWorkerId and register a
+ * shutdown callback to release registration.
+ */
+static void
+pgaio_worker_register(void)
+{
+ MyIoWorkerId = -1;
+
+ /*
+ * XXX: This could do with more fine-grained locking. But it's also not
+ * very common for the number of workers to change at the moment...
+ */
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+
+ for (int i = 0; i < io_workers; ++i)
+ {
+ if (!io_worker_control->workers[i].in_use)
+ {
+ Assert(io_worker_control->workers[i].latch == NULL);
+ io_worker_control->workers[i].in_use = true;
+ MyIoWorkerId = i;
+ break;
+ }
+ else
+ Assert(io_worker_control->workers[i].latch != NULL);
+ }
+
+ if (MyIoWorkerId == -1)
+ elog(ERROR, "couldn't find a free worker slot");
+
+ io_worker_control->idle_worker_mask |= (UINT64_C(1) << MyIoWorkerId);
+ io_worker_control->workers[MyIoWorkerId].latch = MyLatch;
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+
+ on_shmem_exit(pgaio_worker_die, 0);
+}
+
void
IoWorkerMain(char *startup_data, size_t startup_data_len)
{
sigjmp_buf local_sigjmp_buf;
+ PgAioHandle *volatile error_ioh = NULL;
+ volatile int error_errno = 0;
+ char cmd[128];
MyBackendType = B_IO_WORKER;
AuxiliaryProcessMainCommon();
- /* TODO review all signals */
pqsignal(SIGHUP, SignalHandlerForConfigReload);
pqsignal(SIGINT, die); /* to allow manually triggering worker restart */
@@ -53,7 +367,12 @@ IoWorkerMain(char *startup_data, size_t startup_data_len)
pqsignal(SIGPIPE, SIG_IGN);
pqsignal(SIGUSR1, procsignal_sigusr1_handler);
pqsignal(SIGUSR2, SignalHandlerForShutdownRequest);
- sigprocmask(SIG_SETMASK, &UnBlockSig, NULL);
+
+ /* also registers a shutdown callback to unregister */
+ pgaio_worker_register();
+
+ sprintf(cmd, "io worker: %d", MyIoWorkerId);
+ set_ps_display(cmd);
/* see PostgresMain() */
if (sigsetjmp(local_sigjmp_buf, 1) != 0)
@@ -61,6 +380,27 @@ IoWorkerMain(char *startup_data, size_t startup_data_len)
error_context_stack = NULL;
HOLD_INTERRUPTS();
+ /*
+ * In the - very unlikely - case that the IO failed in a way that
+ * raises an error we need to mark the IO as failed.
+ *
+ * Need to do just enough error recovery so that we can mark the IO as
+ * failed and then exit (postmaster will start a new worker).
+ */
+ LWLockReleaseAll();
+
+ if (error_ioh != NULL)
+ {
+ /* should never fail without setting error_errno */
+ Assert(error_errno != 0);
+
+ errno = error_errno;
+
+ START_CRIT_SECTION();
+ pgaio_io_process_completion(error_ioh, -error_errno);
+ END_CRIT_SECTION();
+ }
+
EmitErrorReport();
proc_exit(1);
@@ -69,12 +409,92 @@ IoWorkerMain(char *startup_data, size_t startup_data_len)
/* We can now handle ereport(ERROR) */
PG_exception_stack = &local_sigjmp_buf;
+ sigprocmask(SIG_SETMASK, &UnBlockSig, NULL);
+
while (!ShutdownRequestPending)
{
- WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, -1,
- WAIT_EVENT_IO_WORKER_MAIN);
- ResetLatch(MyLatch);
- CHECK_FOR_INTERRUPTS();
+ uint32 io_index;
+ Latch *latches[IO_WORKER_WAKEUP_FANOUT];
+ int nlatches = 0;
+ int nwakeups = 0;
+ int worker;
+
+ /* Try to get a job to do. */
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+ if ((io_index = pgaio_worker_submission_queue_consume()) == UINT32_MAX)
+ {
+ /*
+ * Nothing to do. Mark self idle.
+ *
+ * XXX: Invent some kind of back pressure to reduce useless
+ * wakeups?
+ */
+ io_worker_control->idle_worker_mask |= (UINT64_C(1) << MyIoWorkerId);
+ }
+ else
+ {
+ /* Got one. Clear idle flag. */
+ io_worker_control->idle_worker_mask &= ~(UINT64_C(1) << MyIoWorkerId);
+
+ /* See if we can wake up some peers. */
+ nwakeups = Min(pgaio_worker_submission_queue_depth(),
+ IO_WORKER_WAKEUP_FANOUT);
+ for (int i = 0; i < nwakeups; ++i)
+ {
+ if ((worker = pgaio_choose_idle_worker()) < 0)
+ break;
+ latches[nlatches++] = io_worker_control->workers[worker].latch;
+ }
+ }
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+
+ for (int i = 0; i < nlatches; ++i)
+ SetLatch(latches[i]);
+
+ if (io_index != UINT32_MAX)
+ {
+ PgAioHandle *ioh = NULL;
+
+ ioh = &pgaio_ctl->io_handles[io_index];
+ error_ioh = ioh;
+
+ pgaio_debug_io(DEBUG4, ioh,
+ "worker %d processing IO",
+ MyIoWorkerId);
+
+ /*
+ * It's very unlikely, but possible, that reopen fails. E.g. due
+ * to memory allocations failing or file permissions changing or
+ * such. In that case we need to fail the IO.
+ *
+ * There's not really a good errno we can report here.
+ */
+ error_errno = ENOENT;
+ pgaio_io_reopen(ioh);
+
+ /*
+ * To be able to exercise the reopen-fails path, allow injection
+ * points to trigger a failure at this point.
+ */
+ pgaio_io_call_inj(ioh, "AIO_WORKER_AFTER_REOPEN");
+
+ error_errno = 0;
+ error_ioh = NULL;
+
+ /*
+ * We don't expect this to ever fail, no need to keep error_ioh
+ * around. pgaio_io_perform_synchronously() contains a critical
+ * section.
+ */
+ pgaio_io_perform_synchronously(ioh);
+ }
+ else
+ {
+ WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, -1,
+ WAIT_EVENT_IO_WORKER_MAIN);
+ ResetLatch(MyLatch);
+ CHECK_FOR_INTERRUPTS();
+ }
}
proc_exit(0);
@@ -83,6 +503,5 @@ IoWorkerMain(char *startup_data, size_t startup_data_len)
bool
pgaio_workers_enabled(void)
{
- /* placeholder for future commit */
- return false;
+ return io_method == IOMETHOD_WORKER;
}
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index d18b788b8a1..cb977b049d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -348,6 +348,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index b7c84d061e2..15954f42d4e 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -74,6 +74,7 @@
#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/bufpage.h"
+#include "storage/io_worker.h"
#include "storage/large_object.h"
#include "storage/pg_shmem.h"
#include "storage/predicate.h"
@@ -3265,6 +3266,18 @@ struct config_int ConfigureNamesInt[] =
check_io_max_concurrency, NULL, NULL
},
+ {
+ {"io_workers",
+ PGC_SIGHUP,
+ RESOURCES_IO,
+ gettext_noop("Number of IO worker processes, for io_method=worker."),
+ NULL,
+ },
+ &io_workers,
+ 3, 1, MAX_IO_WORKERS,
+ NULL, assign_io_workers, NULL
+ },
+
{
{"backend_flush_after", PGC_USERSET, RESOURCES_IO,
gettext_noop("Number of pages after which previously performed writes are flushed to disk."),
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index 186bc47b700..3bb4e0d4d7d 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -199,11 +199,12 @@
#maintenance_io_concurrency = 10 # 1-1000; 0 disables prefetching
#io_combine_limit = 128kB # usually 1-32 blocks (depends on OS)
-#io_method = sync # sync (change requires restart)
+#io_method = worker # worker, sync (change requires restart)
#io_max_concurrency = -1 # Max number of IOs that one process
# can execute simultaneously
# -1 sets based on shared_buffers
# (change requires restart)
+#io_workers = 3 # 1-32;
# - Worker Processes -
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 0306827afbd..2cf0120bb56 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -2656,6 +2656,11 @@ include_dir 'conf.d'
Selects the method for executing asynchronous I/O.
Possible values are:
<itemizedlist>
+ <listitem>
+ <para>
+ <literal>worker</literal> (execute asynchronous I/O using worker processes)
+ </para>
+ </listitem>
<listitem>
<para>
<literal>sync</literal> (execute asynchronous I/O synchronously)
@@ -2669,6 +2674,22 @@ include_dir 'conf.d'
</listitem>
</varlistentry>
+ <varlistentry id="guc-io-workers" xreflabel="io_workers">
+ <term><varname>io_workers</varname> (<type>int</type>)
+ <indexterm>
+ <primary><varname>io_workers</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ Selects the number of I/O worker processes to use. The default is 3.
+ </para>
+ <para>
+ Only has an effect if <xref linkend="guc-max-wal-senders"/> is set to
+ <literal>worker</literal>.
+ </para>
+ </listitem>
+ </varlistentry>
</variablelist>
</sect2>
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 14a338e4308..e34727e269a 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -55,6 +55,9 @@ AggStrategy
AggTransInfo
Aggref
AggregateInstrumentation
+AioWorkerControl
+AioWorkerSlot
+AioWorkerSubmissionQueue
AlenState
Alias
AllocBlock
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0013-aio-Add-liburing-dependency.patch (13.2K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/14-v2.4-0013-aio-Add-liburing-dependency.patch)
download | inline diff:
From be83ff6fb350b9900428e4c39c4ce339b354018c Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 5 Jun 2024 19:37:25 -0700
Subject: [PATCH v2.4 13/29] aio: Add liburing dependency
Not yet used.
Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
meson.build | 14 ++++
meson_options.txt | 3 +
configure.ac | 11 +++
src/makefiles/meson.build | 3 +
src/include/pg_config.h.in | 3 +
src/backend/Makefile | 7 +-
doc/src/sgml/installation.sgml | 34 ++++++++
configure | 138 +++++++++++++++++++++++++++++++++
.cirrus.tasks.yml | 1 +
src/Makefile.global.in | 4 +
10 files changed, 215 insertions(+), 3 deletions(-)
diff --git a/meson.build b/meson.build
index 7dd7110318d..53a1365f5de 100644
--- a/meson.build
+++ b/meson.build
@@ -855,6 +855,18 @@ endif
+###############################################################
+# Library: liburing
+###############################################################
+
+liburingopt = get_option('liburing')
+liburing = dependency('liburing', required: liburingopt)
+if liburing.found()
+ cdata.set('USE_LIBURING', 1)
+endif
+
+
+
###############################################################
# Library: libxml
###############################################################
@@ -3069,6 +3081,7 @@ backend_both_deps += [
icu_i18n,
ldap,
libintl,
+ liburing,
libxml,
lz4,
pam,
@@ -3721,6 +3734,7 @@ if meson.version().version_compare('>=0.57')
'gss': gssapi,
'icu': icu,
'ldap': ldap,
+ 'liburing': liburing,
'libxml': libxml,
'libxslt': libxslt,
'llvm': llvm,
diff --git a/meson_options.txt b/meson_options.txt
index d9c7ddccbc4..8e0c63cc782 100644
--- a/meson_options.txt
+++ b/meson_options.txt
@@ -103,6 +103,9 @@ option('ldap', type: 'feature', value: 'auto',
option('libedit_preferred', type: 'boolean', value: false,
description: 'Prefer BSD Libedit over GNU Readline')
+option('liburing', type : 'feature', value: 'auto',
+ description: 'io_uring support, for asynchronous I/O')
+
option('libxml', type: 'feature', value: 'auto',
description: 'XML support')
diff --git a/configure.ac b/configure.ac
index f56681e0d91..065bd89fbbb 100644
--- a/configure.ac
+++ b/configure.ac
@@ -975,6 +975,14 @@ AC_SUBST(with_readline)
PGAC_ARG_BOOL(with, libedit-preferred, no,
[prefer BSD Libedit over GNU Readline])
+#
+# liburing
+#
+AC_MSG_CHECKING([whether to build with liburing support])
+PGAC_ARG_BOOL(with, liburing, no, [io_uring support, for asynchronous I/O],
+ [AC_DEFINE([USE_LIBURING], 1, [Define to build with io_uring support. (--with-liburing)])])
+AC_MSG_RESULT([$with_liburing])
+AC_SUBST(with_liburing)
#
# UUID library
@@ -1422,6 +1430,9 @@ elif test "$with_uuid" = ossp ; then
fi
AC_SUBST(UUID_LIBS)
+if test "$with_liburing" = yes; then
+ PKG_CHECK_MODULES(LIBURING, liburing)
+fi
##
## Header files
diff --git a/src/makefiles/meson.build b/src/makefiles/meson.build
index d49b2079a44..714b7ccaa4e 100644
--- a/src/makefiles/meson.build
+++ b/src/makefiles/meson.build
@@ -199,6 +199,8 @@ pgxs_empty = [
'PTHREAD_CFLAGS', 'PTHREAD_LIBS',
'ICU_LIBS',
+
+ 'LIBURING_CFLAGS', 'LIBURING_LIBS',
]
if host_system == 'windows' and cc.get_argument_syntax() != 'msvc'
@@ -229,6 +231,7 @@ pgxs_deps = {
'gssapi': gssapi,
'icu': icu,
'ldap': ldap,
+ 'liburing': liburing,
'libxml': libxml,
'libxslt': libxslt,
'llvm': llvm,
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 07b2f798abd..df7584a6187 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -663,6 +663,9 @@
/* Define to 1 to build with LDAP support. (--with-ldap) */
#undef USE_LDAP
+/* Define to build with io_uring support. (--with-liburing) */
+#undef USE_LIBURING
+
/* Define to 1 to build with XML support. (--with-libxml) */
#undef USE_LIBXML
diff --git a/src/backend/Makefile b/src/backend/Makefile
index 42d4a28e5aa..7344c8c7f5c 100644
--- a/src/backend/Makefile
+++ b/src/backend/Makefile
@@ -43,9 +43,10 @@ OBJS = \
$(top_builddir)/src/common/libpgcommon_srv.a \
$(top_builddir)/src/port/libpgport_srv.a
-# We put libpgport and libpgcommon into OBJS, so remove it from LIBS; also add
-# libldap and ICU
-LIBS := $(filter-out -lpgport -lpgcommon, $(LIBS)) $(LDAP_LIBS_BE) $(ICU_LIBS)
+# We put libpgport and libpgcommon into OBJS, so remove it from LIBS.
+LIBS := $(filter-out -lpgport -lpgcommon, $(LIBS))
+# The backend conditionally needs libraries that most executables don't need.
+LIBS += $(LDAP_LIBS_BE) $(ICU_LIBS) $(LIBURING_LIBS)
# The backend doesn't need everything that's in LIBS, however
LIBS := $(filter-out -lreadline -ledit -ltermcap -lncurses -lcurses, $(LIBS))
diff --git a/doc/src/sgml/installation.sgml b/doc/src/sgml/installation.sgml
index 3f0a7e9c069..9989fbc5936 100644
--- a/doc/src/sgml/installation.sgml
+++ b/doc/src/sgml/installation.sgml
@@ -1143,6 +1143,24 @@ build-postgresql:
</listitem>
</varlistentry>
+ <varlistentry id="configure-option-with-liburing">
+ <term><option>--with-liburing</option></term>
+ <listitem>
+ <para>
+ Build with liburing, enabling io_uring support for asynchronous I/O.
+ </para>
+ <para>
+ To detect the required compiler and linker options, PostgreSQL will
+ query <command>pkg-config</command>.
+ </para>
+ <para>
+ To use a liburing installation that is in an unusual location, you
+ can set <command>pkg-config</command>-related environment
+ variables (see its documentation).
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry id="configure-option-with-libxml">
<term><option>--with-libxml</option></term>
<listitem>
@@ -2584,6 +2602,22 @@ ninja install
</listitem>
</varlistentry>
+ <varlistentry id="configure-with-liburing-meson">
+ <term><option>-Dliburing={ auto | enabled | disabled }</option></term>
+ <listitem>
+ <para>
+ Build with liburing, enabling io_uring support for asynchronous I/O.
+ Defaults to auto.
+ </para>
+
+ <para>
+ To use a liburing installation that is in an unusual location, you
+ can set <command>pkg-config</command>-related environment
+ variables (see its documentation).
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry id="configure-with-libxml-meson">
<term><option>-Dlibxml={ auto | enabled | disabled }</option></term>
<listitem>
diff --git a/configure b/configure
index 0ffcaeb4367..400a03ce0e1 100755
--- a/configure
+++ b/configure
@@ -651,6 +651,8 @@ LIBOBJS
OPENSSL
ZSTD
LZ4
+LIBURING_LIBS
+LIBURING_CFLAGS
UUID_LIBS
LDAP_LIBS_BE
LDAP_LIBS_FE
@@ -709,6 +711,7 @@ XML2_CFLAGS
XML2_CONFIG
with_libxml
with_uuid
+with_liburing
with_readline
with_systemd
with_selinux
@@ -862,6 +865,7 @@ with_selinux
with_systemd
with_readline
with_libedit_preferred
+with_liburing
with_uuid
with_ossp_uuid
with_libxml
@@ -905,6 +909,8 @@ LDFLAGS_EX
LDFLAGS_SL
PERL
PYTHON
+LIBURING_CFLAGS
+LIBURING_LIBS
MSGFMT
TCLSH'
@@ -1572,6 +1578,7 @@ Optional Packages:
--without-readline do not use GNU Readline nor BSD Libedit for editing
--with-libedit-preferred
prefer BSD Libedit over GNU Readline
+ --with-liburing io_uring support, for asynchronous I/O
--with-uuid=LIB build contrib/uuid-ossp using LIB (bsd,e2fs,ossp)
--with-ossp-uuid obsolete spelling of --with-uuid=ossp
--with-libxml build with XML support
@@ -1618,6 +1625,10 @@ Some influential environment variables:
LDFLAGS_SL extra linker flags for linking shared libraries only
PERL Perl program
PYTHON Python program
+ LIBURING_CFLAGS
+ C compiler flags for LIBURING, overriding pkg-config
+ LIBURING_LIBS
+ linker flags for LIBURING, overriding pkg-config
MSGFMT msgfmt program for NLS
TCLSH Tcl interpreter program (tclsh)
@@ -8681,6 +8692,40 @@ fi
+#
+# liburing
+#
+{ $as_echo "$as_me:${as_lineno-$LINENO}: checking whether to build with liburing support" >&5
+$as_echo_n "checking whether to build with liburing support... " >&6; }
+
+
+
+# Check whether --with-liburing was given.
+if test "${with_liburing+set}" = set; then :
+ withval=$with_liburing;
+ case $withval in
+ yes)
+
+$as_echo "#define USE_LIBURING 1" >>confdefs.h
+
+ ;;
+ no)
+ :
+ ;;
+ *)
+ as_fn_error $? "no argument expected for --with-liburing option" "$LINENO" 5
+ ;;
+ esac
+
+else
+ with_liburing=no
+
+fi
+
+
+{ $as_echo "$as_me:${as_lineno-$LINENO}: result: $with_liburing" >&5
+$as_echo "$with_liburing" >&6; }
+
#
# UUID library
@@ -13112,6 +13157,99 @@ fi
fi
+if test "$with_liburing" = yes; then
+
+pkg_failed=no
+{ $as_echo "$as_me:${as_lineno-$LINENO}: checking for liburing" >&5
+$as_echo_n "checking for liburing... " >&6; }
+
+if test -n "$LIBURING_CFLAGS"; then
+ pkg_cv_LIBURING_CFLAGS="$LIBURING_CFLAGS"
+ elif test -n "$PKG_CONFIG"; then
+ if test -n "$PKG_CONFIG" && \
+ { { $as_echo "$as_me:${as_lineno-$LINENO}: \$PKG_CONFIG --exists --print-errors \"liburing\""; } >&5
+ ($PKG_CONFIG --exists --print-errors "liburing") 2>&5
+ ac_status=$?
+ $as_echo "$as_me:${as_lineno-$LINENO}: \$? = $ac_status" >&5
+ test $ac_status = 0; }; then
+ pkg_cv_LIBURING_CFLAGS=`$PKG_CONFIG --cflags "liburing" 2>/dev/null`
+ test "x$?" != "x0" && pkg_failed=yes
+else
+ pkg_failed=yes
+fi
+ else
+ pkg_failed=untried
+fi
+if test -n "$LIBURING_LIBS"; then
+ pkg_cv_LIBURING_LIBS="$LIBURING_LIBS"
+ elif test -n "$PKG_CONFIG"; then
+ if test -n "$PKG_CONFIG" && \
+ { { $as_echo "$as_me:${as_lineno-$LINENO}: \$PKG_CONFIG --exists --print-errors \"liburing\""; } >&5
+ ($PKG_CONFIG --exists --print-errors "liburing") 2>&5
+ ac_status=$?
+ $as_echo "$as_me:${as_lineno-$LINENO}: \$? = $ac_status" >&5
+ test $ac_status = 0; }; then
+ pkg_cv_LIBURING_LIBS=`$PKG_CONFIG --libs "liburing" 2>/dev/null`
+ test "x$?" != "x0" && pkg_failed=yes
+else
+ pkg_failed=yes
+fi
+ else
+ pkg_failed=untried
+fi
+
+
+
+if test $pkg_failed = yes; then
+ { $as_echo "$as_me:${as_lineno-$LINENO}: result: no" >&5
+$as_echo "no" >&6; }
+
+if $PKG_CONFIG --atleast-pkgconfig-version 0.20; then
+ _pkg_short_errors_supported=yes
+else
+ _pkg_short_errors_supported=no
+fi
+ if test $_pkg_short_errors_supported = yes; then
+ LIBURING_PKG_ERRORS=`$PKG_CONFIG --short-errors --print-errors --cflags --libs "liburing" 2>&1`
+ else
+ LIBURING_PKG_ERRORS=`$PKG_CONFIG --print-errors --cflags --libs "liburing" 2>&1`
+ fi
+ # Put the nasty error message in config.log where it belongs
+ echo "$LIBURING_PKG_ERRORS" >&5
+
+ as_fn_error $? "Package requirements (liburing) were not met:
+
+$LIBURING_PKG_ERRORS
+
+Consider adjusting the PKG_CONFIG_PATH environment variable if you
+installed software in a non-standard prefix.
+
+Alternatively, you may set the environment variables LIBURING_CFLAGS
+and LIBURING_LIBS to avoid the need to call pkg-config.
+See the pkg-config man page for more details." "$LINENO" 5
+elif test $pkg_failed = untried; then
+ { $as_echo "$as_me:${as_lineno-$LINENO}: result: no" >&5
+$as_echo "no" >&6; }
+ { { $as_echo "$as_me:${as_lineno-$LINENO}: error: in \`$ac_pwd':" >&5
+$as_echo "$as_me: error: in \`$ac_pwd':" >&2;}
+as_fn_error $? "The pkg-config script could not be found or is too old. Make sure it
+is in your PATH or set the PKG_CONFIG environment variable to the full
+path to pkg-config.
+
+Alternatively, you may set the environment variables LIBURING_CFLAGS
+and LIBURING_LIBS to avoid the need to call pkg-config.
+See the pkg-config man page for more details.
+
+To get pkg-config, see <http://pkg-config.freedesktop.org/>.
+See \`config.log' for more details" "$LINENO" 5; }
+else
+ LIBURING_CFLAGS=$pkg_cv_LIBURING_CFLAGS
+ LIBURING_LIBS=$pkg_cv_LIBURING_LIBS
+ { $as_echo "$as_me:${as_lineno-$LINENO}: result: yes" >&5
+$as_echo "yes" >&6; }
+
+fi
+fi
##
## Header files
diff --git a/.cirrus.tasks.yml b/.cirrus.tasks.yml
index fffa438cec1..f4a5a50c3a3 100644
--- a/.cirrus.tasks.yml
+++ b/.cirrus.tasks.yml
@@ -444,6 +444,7 @@ task:
--enable-cassert --enable-injection-points --enable-debug \
--enable-tap-tests --enable-nls \
--with-segsize-blocks=6 \
+ --with-liburing \
\
${LINUX_CONFIGURE_FEATURES} \
\
diff --git a/src/Makefile.global.in b/src/Makefile.global.in
index bbe11e75bf0..ca6106d37c4 100644
--- a/src/Makefile.global.in
+++ b/src/Makefile.global.in
@@ -190,6 +190,7 @@ with_systemd = @with_systemd@
with_gssapi = @with_gssapi@
with_krb_srvnam = @with_krb_srvnam@
with_ldap = @with_ldap@
+with_liburing = @with_liburing@
with_libxml = @with_libxml@
with_libxslt = @with_libxslt@
with_llvm = @with_llvm@
@@ -216,6 +217,9 @@ krb_srvtab = @krb_srvtab@
ICU_CFLAGS = @ICU_CFLAGS@
ICU_LIBS = @ICU_LIBS@
+LIBURING_CFLAGS = @LIBURING_CFLAGS@
+LIBURING_LIBS = @LIBURING_LIBS@
+
TCLSH = @TCLSH@
TCL_LIBS = @TCL_LIBS@
TCL_LIB_SPEC = @TCL_LIB_SPEC@
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0014-aio-Add-io_uring-method.patch (16.4K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/15-v2.4-0014-aio-Add-io_uring-method.patch)
download | inline diff:
From f6e3c30c331f385a7bed2ee78704f6393b8d7f82 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 11 Feb 2025 14:29:37 -0500
Subject: [PATCH v2.4 14/29] aio: Add io_uring method
---
src/include/storage/aio.h | 3 +
src/include/storage/aio_internal.h | 3 +
src/include/storage/lwlock.h | 1 +
src/backend/storage/aio/Makefile | 1 +
src/backend/storage/aio/aio.c | 6 +
src/backend/storage/aio/meson.build | 1 +
src/backend/storage/aio/method_io_uring.c | 382 ++++++++++++++++++
src/backend/storage/lmgr/lwlock.c | 1 +
.../utils/activity/wait_event_names.txt | 2 +
src/backend/utils/misc/postgresql.conf.sample | 3 +-
doc/src/sgml/config.sgml | 6 +
src/tools/pgindent/typedefs.list | 1 +
12 files changed, 409 insertions(+), 1 deletion(-)
create mode 100644 src/backend/storage/aio/method_io_uring.c
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index ca0f9b64d97..3c058c84003 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -24,6 +24,9 @@ typedef enum IoMethod
{
IOMETHOD_SYNC = 0,
IOMETHOD_WORKER,
+#ifdef USE_LIBURING
+ IOMETHOD_IO_URING,
+#endif
} IoMethod;
/* We'll default to worker based execution. */
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index c6e7306ed61..ac26aff80b6 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -339,6 +339,9 @@ extern PgAioHandle *pgaio_inj_io_get(void);
/* Declarations for the tables of function pointers exposed by each IO method. */
extern PGDLLIMPORT const IoMethodOps pgaio_sync_ops;
extern PGDLLIMPORT const IoMethodOps pgaio_worker_ops;
+#ifdef USE_LIBURING
+extern PGDLLIMPORT const IoMethodOps pgaio_uring_ops;
+#endif
extern PGDLLIMPORT const IoMethodOps *pgaio_method_ops;
extern PGDLLIMPORT PgAioCtl *pgaio_ctl;
diff --git a/src/include/storage/lwlock.h b/src/include/storage/lwlock.h
index 13a7dc89980..043e8bae7a9 100644
--- a/src/include/storage/lwlock.h
+++ b/src/include/storage/lwlock.h
@@ -217,6 +217,7 @@ typedef enum BuiltinTrancheIds
LWTRANCHE_SUBTRANS_SLRU,
LWTRANCHE_XACT_SLRU,
LWTRANCHE_PARALLEL_VACUUM_DSA,
+ LWTRANCHE_AIO_URING_COMPLETION,
LWTRANCHE_FIRST_USER_DEFINED,
} BuiltinTrancheIds;
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index f51c34a37f8..c06c50771e0 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -14,6 +14,7 @@ OBJS = \
aio_init.o \
aio_io.o \
aio_target.o \
+ method_io_uring.o \
method_sync.o \
method_worker.o \
read_stream.o
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 3eace131de2..a1282351436 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -64,6 +64,9 @@ static void pgaio_io_wait(PgAioHandle *ioh, uint64 ref_generation);
const struct config_enum_entry io_method_options[] = {
{"sync", IOMETHOD_SYNC, false},
{"worker", IOMETHOD_WORKER, false},
+#ifdef USE_LIBURING
+ {"io_uring", IOMETHOD_IO_URING, false},
+#endif
{NULL, 0, false}
};
@@ -81,6 +84,9 @@ PgAioBackend *pgaio_my_backend;
static const IoMethodOps *const pgaio_method_ops_table[] = {
[IOMETHOD_SYNC] = &pgaio_sync_ops,
[IOMETHOD_WORKER] = &pgaio_worker_ops,
+#ifdef USE_LIBURING
+ [IOMETHOD_IO_URING] = &pgaio_uring_ops,
+#endif
};
/* callbacks for the configured io_method, set by assign_io_method */
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 74f94c6e40b..2f0f03d8071 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -6,6 +6,7 @@ backend_sources += files(
'aio_init.c',
'aio_io.c',
'aio_target.c',
+ 'method_io_uring.c',
'method_sync.c',
'method_worker.c',
'read_stream.c',
diff --git a/src/backend/storage/aio/method_io_uring.c b/src/backend/storage/aio/method_io_uring.c
new file mode 100644
index 00000000000..43f7576498c
--- /dev/null
+++ b/src/backend/storage/aio/method_io_uring.c
@@ -0,0 +1,382 @@
+/*-------------------------------------------------------------------------
+ *
+ * method_io_uring.c
+ * AIO - perform AIO using Linux' io_uring
+ *
+ * XXX Write me
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/method_io_uring.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#ifdef USE_LIBURING
+
+#include <liburing.h>
+
+#include "pgstat.h"
+#include "port/pg_iovec.h"
+#include "storage/aio_internal.h"
+#include "storage/fd.h"
+#include "storage/proc.h"
+#include "storage/shmem.h"
+
+
+/* Entry points for IoMethodOps. */
+static size_t pgaio_uring_shmem_size(void);
+static void pgaio_uring_shmem_init(bool first_time);
+static void pgaio_uring_init_backend(void);
+
+static int pgaio_uring_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+static void pgaio_uring_wait_one(PgAioHandle *ioh, uint64 ref_generation);
+
+static void pgaio_uring_sq_from_io(PgAioHandle *ioh, struct io_uring_sqe *sqe);
+
+
+const IoMethodOps pgaio_uring_ops = {
+ .shmem_size = pgaio_uring_shmem_size,
+ .shmem_init = pgaio_uring_shmem_init,
+ .init_backend = pgaio_uring_init_backend,
+
+ .submit = pgaio_uring_submit,
+ .wait_one = pgaio_uring_wait_one,
+};
+
+typedef struct PgAioUringContext
+{
+ LWLock completion_lock;
+
+ struct io_uring io_uring_ring;
+ /* XXX: probably worth padding to a cacheline boundary here */
+} PgAioUringContext;
+
+
+static PgAioUringContext *pgaio_uring_contexts;
+static PgAioUringContext *pgaio_my_uring_context;
+
+/* io_uring local state */
+static struct io_uring local_ring;
+
+
+
+static Size
+pgaio_uring_context_shmem_size(void)
+{
+ uint32 TotalProcs = MaxBackends + NUM_AUXILIARY_PROCS - MAX_IO_WORKERS;
+
+ return mul_size(TotalProcs, sizeof(PgAioUringContext));
+}
+
+static size_t
+pgaio_uring_shmem_size(void)
+{
+ return pgaio_uring_context_shmem_size();
+}
+
+static void
+pgaio_uring_shmem_init(bool first_time)
+{
+ uint32 TotalProcs = MaxBackends + NUM_AUXILIARY_PROCS - MAX_IO_WORKERS;
+ bool found;
+
+ pgaio_uring_contexts = (PgAioUringContext *)
+ ShmemInitStruct("AioUring", pgaio_uring_shmem_size(), &found);
+
+ if (found)
+ return;
+
+ for (int contextno = 0; contextno < TotalProcs; contextno++)
+ {
+ PgAioUringContext *context = &pgaio_uring_contexts[contextno];
+ int ret;
+
+ /*
+ * XXX: Probably worth sharing the WQ between the different rings,
+ * when supported by the kernel. Could also cause additional
+ * contention, I guess?
+ */
+#if 0
+ if (!AcquireExternalFD())
+ elog(ERROR, "No external FD available");
+#endif
+ ret = io_uring_queue_init(io_max_concurrency, &context->io_uring_ring, 0);
+ if (ret < 0)
+ elog(ERROR, "io_uring_queue_init failed: %s", strerror(-ret));
+
+ LWLockInitialize(&context->completion_lock, LWTRANCHE_AIO_URING_COMPLETION);
+ }
+}
+
+static void
+pgaio_uring_init_backend(void)
+{
+ int ret;
+
+ pgaio_my_uring_context = &pgaio_uring_contexts[MyProcNumber];
+
+ ret = io_uring_queue_init(32, &local_ring, 0);
+ if (ret < 0)
+ elog(ERROR, "io_uring_queue_init failed: %s", strerror(-ret));
+}
+
+static int
+pgaio_uring_submit(uint16 num_staged_ios, PgAioHandle **staged_ios)
+{
+ struct io_uring *uring_instance = &pgaio_my_uring_context->io_uring_ring;
+ int in_flight_before = dclist_count(&pgaio_my_backend->in_flight_ios);
+
+ Assert(num_staged_ios <= PGAIO_SUBMIT_BATCH_SIZE);
+
+ for (int i = 0; i < num_staged_ios; i++)
+ {
+ PgAioHandle *ioh = staged_ios[i];
+ struct io_uring_sqe *sqe;
+
+ sqe = io_uring_get_sqe(uring_instance);
+
+ if (!sqe)
+ elog(ERROR, "io_uring submission queue is unexpectedly full");
+
+ pgaio_io_prepare_submit(ioh);
+ pgaio_uring_sq_from_io(ioh, sqe);
+
+ /*
+ * io_uring executes IO in process context if possible. That's
+ * generally good, as it reduces context switching. When performing a
+ * lot of buffered IO that means that copying between page cache and
+ * userspace memory happens in the foreground, as it can't be
+ * offloaded to DMA hardware as is possible when using direct IO. When
+ * executing a lot of buffered IO this causes io_uring to be slower
+ * than worker mode, as worker mode parallelizes the copying. io_uring
+ * can be told to offload work to worker threads instead.
+ *
+ * If an IO is buffered IO and we already have IOs in flight or
+ * multiple IOs are being submitted, we thus tell io_uring to execute
+ * the IO in the background. We don't do so for the first few IOs
+ * being submitted as executing in this process' context has lower
+ * latency.
+ */
+ if (in_flight_before > 4 && (ioh->flags & PGAIO_HF_BUFFERED))
+ io_uring_sqe_set_flags(sqe, IOSQE_ASYNC);
+
+ in_flight_before++;
+ }
+
+ while (true)
+ {
+ int ret;
+
+ pgstat_report_wait_start(WAIT_EVENT_AIO_IO_URING_SUBMIT);
+ ret = io_uring_submit(uring_instance);
+ pgstat_report_wait_end();
+
+ if (ret == -EINTR)
+ {
+ pgaio_debug(DEBUG3,
+ "aio method uring: submit EINTR, nios: %d",
+ num_staged_ios);
+ continue;
+ }
+ if (ret < 0)
+ elog(PANIC, "failed: %d/%s",
+ ret, strerror(-ret));
+ else if (ret != num_staged_ios)
+ {
+ /* likely unreachable, but if it is, we would need to re-submit */
+ elog(PANIC, "submitted only %d of %d",
+ ret, num_staged_ios);
+ }
+ else
+ {
+ pgaio_debug(DEBUG4,
+ "aio method uring: submitted %d IOs",
+ num_staged_ios);
+ }
+ break;
+ }
+
+ return num_staged_ios;
+}
+
+
+#define PGAIO_MAX_LOCAL_COMPLETED_IO 32
+
+static void
+pgaio_uring_drain_locked(PgAioUringContext *context)
+{
+ int ready;
+ int orig_ready;
+
+ /*
+ * Don't drain more events than available right now. Otherwise it's
+ * plausible that one backend could get stuck, for a while, receiving CQEs
+ * without actually processing them.
+ */
+ orig_ready = ready = io_uring_cq_ready(&context->io_uring_ring);
+
+ while (ready > 0)
+ {
+ struct io_uring_cqe *cqes[PGAIO_MAX_LOCAL_COMPLETED_IO];
+ uint32 ncqes;
+
+ START_CRIT_SECTION();
+ ncqes =
+ io_uring_peek_batch_cqe(&context->io_uring_ring,
+ cqes,
+ Min(PGAIO_MAX_LOCAL_COMPLETED_IO, ready));
+ Assert(ncqes <= ready);
+
+ ready -= ncqes;
+
+ for (int i = 0; i < ncqes; i++)
+ {
+ struct io_uring_cqe *cqe = cqes[i];
+ PgAioHandle *ioh;
+
+ ioh = io_uring_cqe_get_data(cqe);
+ io_uring_cqe_seen(&context->io_uring_ring, cqe);
+
+ pgaio_io_process_completion(ioh, cqe->res);
+ }
+
+ END_CRIT_SECTION();
+
+ pgaio_debug(DEBUG3,
+ "drained %d/%d, now expecting %d",
+ ncqes, orig_ready, io_uring_cq_ready(&context->io_uring_ring));
+ }
+}
+
+static void
+pgaio_uring_wait_one(PgAioHandle *ioh, uint64 ref_generation)
+{
+ PgAioHandleState state;
+ ProcNumber owner_procno = ioh->owner_procno;
+ PgAioUringContext *owner_context = &pgaio_uring_contexts[owner_procno];
+ bool expect_cqe;
+ int waited = 0;
+
+ /*
+ * We ought to have a smarter locking scheme, nearly all the time the
+ * backend owning the ring will consume the completions, making the
+ * locking unnecessarily expensive.
+ */
+ LWLockAcquire(&owner_context->completion_lock, LW_EXCLUSIVE);
+
+ while (true)
+ {
+ pgaio_debug_io(DEBUG3, ioh,
+ "wait_one io_gen: %llu, ref_gen: %llu, cycle %d",
+ (long long unsigned) ref_generation,
+ (long long unsigned) ioh->generation,
+ waited);
+
+ if (pgaio_io_was_recycled(ioh, ref_generation, &state) ||
+ state != PGAIO_HS_SUBMITTED)
+ {
+ break;
+ }
+ else if (io_uring_cq_ready(&owner_context->io_uring_ring))
+ {
+ expect_cqe = true;
+ }
+ else
+ {
+ int ret;
+ struct io_uring_cqe *cqes;
+
+ pgstat_report_wait_start(WAIT_EVENT_AIO_IO_URING_COMPLETION);
+ ret = io_uring_wait_cqes(&owner_context->io_uring_ring, &cqes, 1, NULL, NULL);
+ pgstat_report_wait_end();
+
+ if (ret == -EINTR)
+ {
+ continue;
+ }
+ else if (ret != 0)
+ {
+ elog(PANIC, "unexpected: %d/%s: %m", ret, strerror(-ret));
+ }
+ else
+ {
+ Assert(cqes != NULL);
+ expect_cqe = true;
+ waited++;
+ }
+ }
+
+ if (expect_cqe)
+ {
+ pgaio_uring_drain_locked(owner_context);
+ }
+ }
+
+ LWLockRelease(&owner_context->completion_lock);
+
+ pgaio_debug(DEBUG3,
+ "wait_one with %d sleeps",
+ waited);
+}
+
+static void
+pgaio_uring_sq_from_io(PgAioHandle *ioh, struct io_uring_sqe *sqe)
+{
+ struct iovec *iov;
+
+ switch (ioh->op)
+ {
+ case PGAIO_OP_READV:
+ iov = &pgaio_ctl->iovecs[ioh->iovec_off];
+ if (ioh->op_data.read.iov_length == 1)
+ {
+ io_uring_prep_read(sqe,
+ ioh->op_data.read.fd,
+ iov->iov_base,
+ iov->iov_len,
+ ioh->op_data.read.offset);
+ }
+ else
+ {
+ io_uring_prep_readv(sqe,
+ ioh->op_data.read.fd,
+ iov,
+ ioh->op_data.read.iov_length,
+ ioh->op_data.read.offset);
+
+ }
+ break;
+
+ case PGAIO_OP_WRITEV:
+ iov = &pgaio_ctl->iovecs[ioh->iovec_off];
+ if (ioh->op_data.write.iov_length == 1)
+ {
+ io_uring_prep_write(sqe,
+ ioh->op_data.write.fd,
+ iov->iov_base,
+ iov->iov_len,
+ ioh->op_data.write.offset);
+ }
+ else
+ {
+ io_uring_prep_writev(sqe,
+ ioh->op_data.write.fd,
+ iov,
+ ioh->op_data.write.iov_length,
+ ioh->op_data.write.offset);
+ }
+ break;
+
+ case PGAIO_OP_INVALID:
+ elog(ERROR, "trying to prepare invalid IO operation for execution");
+ }
+
+ io_uring_sqe_set_data(sqe, ioh);
+}
+
+#endif /* USE_LIBURING */
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index b02625194be..9fec95dd4b7 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -166,6 +166,7 @@ static const char *const BuiltinTrancheNames[] = {
[LWTRANCHE_SUBTRANS_SLRU] = "SubtransSLRU",
[LWTRANCHE_XACT_SLRU] = "XactSLRU",
[LWTRANCHE_PARALLEL_VACUUM_DSA] = "ParallelVacuumDSA",
+ [LWTRANCHE_AIO_URING_COMPLETION] = "AioUringCompletion",
};
StaticAssertDecl(lengthof(BuiltinTrancheNames) ==
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index cb977b049d8..a2d1a9fa4ec 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -192,6 +192,8 @@ ABI_compatibility:
Section: ClassName - WaitEventIO
+AIO_IO_URING_SUBMIT "Waiting for IO submission via io_uring."
+AIO_IO_URING_COMPLETION "Waiting for IO completion via io_uring."
AIO_IO_COMPLETION "Waiting for IO completion."
BASEBACKUP_READ "Waiting for base backup to read from a file."
BASEBACKUP_SYNC "Waiting for data written by a base backup to reach durable storage."
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index 3bb4e0d4d7d..50fde0ba2c3 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -199,7 +199,8 @@
#maintenance_io_concurrency = 10 # 1-1000; 0 disables prefetching
#io_combine_limit = 128kB # usually 1-32 blocks (depends on OS)
-#io_method = worker # worker, sync (change requires restart)
+#io_method = worker # worker, io_uring, sync
+ # (change requires restart)
#io_max_concurrency = -1 # Max number of IOs that one process
# can execute simultaneously
# -1 sets based on shared_buffers
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 2cf0120bb56..c13c1e9e95b 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -2661,6 +2661,12 @@ include_dir 'conf.d'
<literal>worker</literal> (execute asynchronous I/O using worker processes)
</para>
</listitem>
+ <listitem>
+ <para>
+ <literal>io_uring</literal> (execute asynchronous I/O using
+ io_uring, if available)
+ </para>
+ </listitem>
<listitem>
<para>
<literal>sync</literal> (execute asynchronous I/O synchronously)
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index e34727e269a..d4734b85c0d 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2131,6 +2131,7 @@ PgAioReturn
PgAioTargetData
PgAioTargetID
PgAioTargetInfo
+PgAioUringContext
PgAioWaitRef
PgArchData
PgBackendGSSStatus
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0015-aio-Add-README.md-explaining-higher-level-desig.patch (18.6K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/16-v2.4-0015-aio-Add-README.md-explaining-higher-level-desig.patch)
download | inline diff:
From 21f1e178886b748049ef0f8fc2ff279ace66092f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 13:44:37 -0500
Subject: [PATCH v2.4 15/29] aio: Add README.md explaining higher level design
---
src/backend/storage/aio/README.md | 422 ++++++++++++++++++++++++++++++
src/backend/storage/aio/aio.c | 2 +
2 files changed, 424 insertions(+)
create mode 100644 src/backend/storage/aio/README.md
diff --git a/src/backend/storage/aio/README.md b/src/backend/storage/aio/README.md
new file mode 100644
index 00000000000..55e64194ded
--- /dev/null
+++ b/src/backend/storage/aio/README.md
@@ -0,0 +1,422 @@
+# Asynchronous & Direct IO
+
+## Motivation
+
+### Why Asynchronous IO
+
+Until the introduction of asynchronous IO Postgres relied on the operating
+system to hide the cost of synchronous IO from Postgres. While this worked
+surprisingly well in a lot of workloads, it does not do as good a job on
+prefetching and controlled writeback as we would like.
+
+There are important expensive operations like `fdatasync()` where the operating
+system cannot hide the storage latency. This is particularly important for WAL
+writes, where the ability to asynchronously issue `fdatasync()` or O_DSYNC
+writes can yield significantly higher throughput.
+
+
+### Why Direct / unbuffered IO
+
+The main reason to want to use Direct IO are:
+
+- Lower CPU usage / higher throughput. Particularly on modern storage buffered
+ writes are bottlenecked by the operating system having to copy data from the
+ kernel's page cache to postgres buffer pool using the CPU. Whereas direct IO
+ can often move the data directly between the storage devices and postgres'
+ buffer cache, using DMA. While that transfer is ongoing, the CPU is free to
+ perform other work.
+- Reduced latency - Direct IO can have substantially lower latency than
+ buffered IO, which can be impactful for OLTP workloads bottlenecked by WAL
+ write latency.
+- Avoiding double buffering between operating system cache and postgres'
+ shared_buffers.
+- Better control over the timing and pace of dirty data writeback.
+
+
+The main reason *not* to use Direct IO are:
+
+- Without AIO, Direct IO is unusably slow for most purposes.
+- Even with AIO, many parts of postgres need to be modified to perform
+ explicit prefetching.
+- In situations where shared_buffers cannot be set appropriately large,
+ e.g. because there are many different postgres instances hosted on shared
+ hardware, performance will often be worse than when using buffered IO.
+
+
+## AIO Usage Example
+
+In many cases code that can benefit from AIO does not directly have to
+interact with the AIO interface, but can use AIO via higher-level
+abstractions. See [Helpers](#helpers).
+
+In this example, a buffer will be read into shared buffers.
+
+```C
+/*
+ * Result of the operation, only to be accessed in this backend.
+ */
+PgAioReturn ioret;
+
+/*
+ * Acquire an AIO Handle, ioret will get result upon completion.
+ *
+ * Note that ioret needs to stay alive until the IO completes or
+ * CurrentResourceOwner is released (i.e. an error is thrown).
+ */
+PgAioHandle *ioh = pgaio_io_acquire(CurrentResourceOwner, &ioret);
+
+/*
+ * Reference that can be used to wait for the IO we initiate below. This
+ * reference can reside in local or shared memory and waited upon by any
+ * process. An arbitrary number of references can be made for each IO.
+ */
+PgAioWaitRef iow;
+
+pgaio_io_get_wref(ioh, &iow);
+
+/*
+ * Arrange for shared buffer completion callbacks to be called upon completion
+ * of the IO. This callback will update the buffer descriptors associated with
+ * the AioHandle, which e.g. allows other backends to access the buffer.
+ *
+ * Multiple completion callbacks can be registered for each handle.
+ */
+pgaio_io_register_callbacks(ioh, PGAIO_HCB_SHARED_BUFFER_READV);
+
+/*
+ * The completion callback needs to know which buffers to update when the IO
+ * completes. As the AIO subsystem does not know about buffers, we have to
+ * associate this information with the AioHandle, for use by the completion
+ * callback registered above.
+ *
+ * In this example we're reading only a single buffer, hence the 1.
+ */
+pgaio_io_set_handle_data_32(ioh, (uint32 *) buffer, 1);
+
+/*
+ * Pass the AIO handle to lower-level function. When operating on the level of
+ * buffers, we don't know how exactly the IO is performed, that is the
+ * responsibility of the storage manager implementation.
+ *
+ * E.g. md.c needs to translate block numbers into offsets in segments.
+ *
+ * Once the IO handle has been handed off to smgstartreadv(), it may not
+ * further be used, as the IO may immediately get executed below
+ * smgrstartreadv() and the handle reused for another IO.
+ *
+ * To issue multiple IOs in an efficient way, a caller can call
+ * pgaio_enter_batchmode() before starting multiple IOs, and end that batch
+ * with pgaio_exit_batchmode(). Note that one needs to be careful while there
+ * may be unsubmitted IOs, as another backend may need to wait for one of the
+ * unsubmitted IOs. If this backend then had to to wait for the other backend,
+ * it'd end in an undetected deadlock. See pgaio_enter_batchmode() for more
+ * details.
+ *
+ * Note that even while in batchmode an IO might get submitted immediately,
+ * e.g. due to reaching a limit on the number of unsubmitted IOs, and even
+ * complete before smgrstartreadv() returns.
+ */
+smgrstartreadv(ioh, operation->smgr, forknum, blkno,
+ BufferGetBlock(buffer), 1);
+
+/*
+ * To benefit from AIO, it is beneficial to perform other work, including
+ * submitting other IOs, before waiting for the IO to complete. Otherwise
+ * we could just have used synchronous, blocking IO.
+ */
+perform_other_work();
+
+/*
+ * We did some other work and now need the IO operation to have completed to
+ * continue.
+ */
+pgaio_wref_wait(&iow);
+
+/*
+ * At this point the IO has completed. We do not yet know whether it succeeded
+ * or failed, however. The buffer's state has been updated, which allows other
+ * backends to use the buffer (if the IO succeeded), or retry the IO (if it
+ * failed).
+ *
+ * Note that in case the IO has failed, a LOG message may have been emitted,
+ * but no ERROR has been raised. This is crucial, as another backend waiting
+ * for this IO should not see an ERROR.
+ *
+ * To check whether the operation succeeded, and to raise an ERROR, or if more
+ * appropriate LOG, the PgAioReturn we passed to pgaio_io_acquire() is used.
+ */
+if (ioret.result.status == ARS_ERROR)
+ pgaio_result_report(aio_ret.result, &aio_ret.target_data, ERROR);
+
+/*
+ * Besides having succeeded completely, the IO could also have partially
+ * completed. If we e.g. tried to read many blocks at once, the read might have
+ * only succeeded for the first few blocks.
+ *
+ * If the IO partially succeeded and this backend needs all blocks to have
+ * completed, this backend needs to reissue the IO for the remaining buffers.
+ * The AIO subsystem cannot handle this retry transparently.
+ *
+ * As this example is already long, and we only read a single block, we'll just
+ * error out if there's a partial read.
+ */
+if (ioret.result.status == ARS_PARTIAL)
+ pgaio_result_report(aio_ret.result, &aio_ret.target_data, ERROR);
+
+/*
+ * The IO succeeded, so we can use the buffer now.
+ */
+```
+
+
+## Design Criteria & Motivation
+
+### Deadlock and Starvation Dangers due to AIO
+
+Using AIO in a naive way can easily lead to deadlocks in an environment where
+the source/target of AIO are shared resources, like pages in postgres'
+shared_buffers.
+
+Consider one backend performing readahead on a table, initiating IO for a
+number of buffers ahead of the current "scan position". If that backend then
+performs some operation that blocks, or even just is slow, the IO completion
+for the asynchronously initiated read may not be processed.
+
+This AIO implementation solves this problem by requiring that AIO methods
+either allow AIO completions to be processed by any backend in the system
+(e.g. io_uring), or to guarantee that AIO processing will happen even when the
+issuing backend is blocked (e.g. worker mode, which offloads completion
+processing to the AIO workers).
+
+
+### IO can be started in critical sections
+
+Using AIO for WAL writes can reduce the overhead of WAL logging substantially:
+
+- AIO allows to start WAL writes eagerly, so they complete before needing to
+ wait
+- AIO allows to have multiple WAL flushes in progress at the same time
+- AIO makes it more realistic to use O\_DIRECT + O\_DSYNC, which can reduce
+ the number of roundtrips to storage on some OSs and storage HW (buffered IO
+ and direct IO without O_DSYNC needs to issue a write and after the writes
+ completion a cache cache flush, whereas O\_DIRECT + O\_DSYNC can use a
+ single FUA write).
+
+The need to be able to execute IO in critical sections has substantial design
+implication on the AIO subsystem. Mainly because completing IOs (see prior
+section) needs to be possible within a critical section, even if the
+to-be-completed IO itself was not issued in a critical section. Consider
+e.g. the case of a backend first starting a number of writes from shared
+buffers and then starting to flush the WAL. Because only a limited amount of
+IO can be in-progress at the same time, initiating IO for flushing the WAL may
+require to first complete IO that was started earlier.
+
+
+### State for AIO needs to live in shared memory
+
+Because postgres uses a process model and because AIOs need to be
+complete-able by any backend much of the state of the AIO subsystem needs to
+live in shared memory.
+
+In an `EXEC_BACKEND` build backends executable code and other process local
+state is not necessarily mapped to the same addresses in each process due to
+ASLR. This means that the shared memory cannot contain pointer to callbacks.
+
+
+## Design of the AIO Subsystem
+
+
+### AIO Methods
+
+To achieve portability and performance, multiple methods of performing AIO are
+implemented and others are likely worth adding in the future.
+
+
+#### Synchronous Mode
+
+`io_method=sync` does not actually perform AIO but allows to use the AIO API
+while performing synchronous IO. This can be useful for debugging. The code
+for the synchronous mode is also used as a fallback by e.g. the [worker
+mode](#worker) uses it to execute IO that cannot be executed by workers.
+
+
+#### Worker
+
+`io_method=worker` is available on every platform postgres runs on, and
+implements asynchronous IO - from the view of the issuing process - by
+dispatching the IO to one of several worker processes performing the IO in a
+synchronous manner.
+
+
+#### io_uring
+
+`io_method=io_uring` is available on Linux 5.1+. In contrast to worker mode it
+dispatches all IO from within the process, lowering context switch rate /
+latency.
+
+
+### AIO Handles
+
+The central API piece for postgres' AIO abstraction are AIO handles. To
+execute an IO one first has to acquire an IO handle (`pgaio_io_acquire()`) and
+then "defined", i.e. associate an IO operation with the handle.
+
+Often AIO handles are acquired on a higher level and then passed to a lower
+level to be fully defined. E.g., for IO to/from shared buffers, bufmgr.c
+routines acquire the handle, which is then passed through smgr.c, md.c to be
+finally fully defined in fd.c.
+
+The functions used at the lowest level to define the operation are
+`pgaio_io_prep_*()`.
+
+Because acquisition of an IO handle
+[must always succeed](#io-can-be-started-in-critical-sections)
+and the number of AIO Handles
+[has to be limited](#state-for-aio-needs-to-live-in-shared-memory)
+AIO handles can be reused as soon as they have completed. Obviously code needs
+to be able to react to IO completion. Shared state can be updated using
+[AIO Completion callbacks](#aio-callbacks)
+and the issuing backend can provide a backend local variable to receive the
+result of the IO, as described in
+[AIO Result](#aio-results)
+. An IO can be waited for, by both the issuing and any other backend, using
+[AIO References](#aio-wait-references).
+
+
+Because an AIO Handle is not executable just after calling `pgaio_io_acquire()`
+and because `pgaio_io_acquire()` needs to be able to succeed, only a single AIO
+Handle may be acquired (i.e. returned by `pgaio_io_acquire()`) without causing
+the IO to have been defined (by, potentially indirectly, causing
+`pgaio_io_prep_*()` to have been called). Otherwise a backend could trivially
+self-deadlock by using up all AIO Handles without the ability to wait for some
+of the IOs to complete.
+
+If it turns out that an AIO Handle is not needed, e.g., because the handle was
+acquired before holding a contended lock, it can be released without being
+defined using `pgaio_io_release()`.
+
+
+### AIO Callbacks
+
+Commonly several layers need to react to completion of an IO. E.g. for a read
+md.c needs to check if the IO outright failed or was shorter than needed,
+bufmgr.c needs to verify the page looks valid and bufmgr.c needs to update the
+BufferDesc to update the buffer's state.
+
+The fact that several layers / subsystems need to react to IO completion poses
+a few challenges:
+
+- Upper layers should not need to know details of lower layers. E.g. bufmgr.c
+ should not assume the IO will pass through md.c. Therefore upper levels
+ cannot know what lower layers would consider an error.
+
+- Lower layers should not need to know about upper layers. E.g. smgr APIs are
+ used going through shared buffers but are also used bypassing shared
+ buffers. This means that e.g. md.c is not in a position to validate
+ checksums.
+
+- Having code in the AIO subsystem for every possible combination of layers
+ would lead to a lot of duplication.
+
+The "solution" to this the ability to associate multiple completion callbacks
+with a handle. E.g. bufmgr.c can have a callback to update the BufferDesc
+state and to verify the page and md.c. another callback to check if the IO
+operation was successful.
+
+As [mentioned](#state-for-aio-needs-to-live-in-shared-memory), shared memory
+currently cannot contain function pointers. Because of that completion
+callbacks are not directly identified by function pointers but by IDs
+(`PgAioHandleCallbackID`). A substantial added benefit is that that
+allows callbacks to be identified by much smaller amount of memory (a single
+byte currently).
+
+In addition to completion, AIO callbacks also are called to "prepare" an
+IO. This is, e.g., used to increase buffer reference counts to account for the
+AIO subsystem referencing the buffer, which is required to handle the case
+where the issuing backend errors out and releases its own pins while the IO is
+still ongoing.
+
+As [explained earlier](#io-can-be-started-in-critical-sections) IO completions
+need to be safe to execute in critical sections. To allow the backend that
+issued the IO to error out in case of failure [AIO Result](#aio-results) can
+be used.
+
+
+### AIO Targets
+
+In addition to the completion callbacks describe above, each AIO Handle has
+exactly one "target". Each target has some space inside an AIO Handle with
+information specific to the target and can provide callbacks to allow to
+reopen the underlying file (required for worker mode) and to describe the IO
+operation (used for debug logging and error messages).
+
+I.e., if two different uses of AIO can describe the identity of the file being
+operated on the same way, it likely makes sense to use the same
+target. E.g. different smgr implementations can describe IO with
+RelFileLocator, ForkNumber and BlockNumber and can thus share a target. In
+contrast, IO for a WAL file would be described with TimeLineID and XLogRecPtr
+and it would not make sense to use the same target for smgr and WAL.
+
+
+### AIO Wait References
+
+As [described above](#aio-handles), AIO Handles can be reused immediately
+after completion and therefore cannot be used to wait for completion of the
+IO. Waiting is enabled using AIO wait references, which do not just identify
+an AIO Handle but also include the handles "generation".
+
+A reference to an AIO Handle can be acquired using `pgaio_io_get_wref()` and
+then waited upon using `pgaio_wref_wait()`.
+
+
+### AIO Results
+
+As AIO completion callbacks
+[are executed in critical sections](#io-can-be-started-in-critical-sections)
+and [may be executed by any backend](#deadlock-and-starvation-dangers-due-to-aio)
+completion callbacks cannot be used to, e.g., make the query that triggered an
+IO ERROR out.
+
+To allow to react to failing IOs the issuing backend can pass a pointer to a
+`PgAioReturn` in backend local memory. Before an AIO Handle is reused the
+`PgAioReturn` is filled with information about the IO. This includes
+information about whether the IO was successful (as a value of
+`PgAioResultStatus`) and enough information to raise an error in case of a
+failure (via `pgaio_result_report()`, with the error details encoded in
+`PgAioResult`).
+
+XXX: "return" vs "result" vs "result status" seems quite confusing. The naming
+should be improved.
+
+
+### AIO Errors
+
+It would be very convenient to have shared completion callbacks encode the
+details of errors as an `ErrorData` that could be raised at a later
+time. Unfortunately doing so would require allocating memory. While elog.c can
+guarantee (well, kinda) that logging a message will not run out of memory,
+that only works because a very limited number of messages are in the process
+of being logged. With AIO a large number of concurrently issued AIOs might
+fail.
+
+To avoid the need for preallocating a potentially large amount of memory (in
+shared memory no less!), completion callbacks instead have to encode errors in
+a more compact format that can be converted into an error message.
+
+
+## Helpers
+
+Using the low-level AIO API introduces too much complexity to do so all over
+the tree. Most uses of AIO should be done via reusable, higher-level,
+helpers.
+
+
+### Read Stream
+
+A common and very beneficial use of AIO are reads where a substantial number
+of to-be-read locations are known ahead of time. E.g., for a sequential scan
+the set of blocks that need to be read can be determined solely by knowing the
+current position and checking the buffer mapping table.
+
+The [Read Stream](../../../include/storage/read_stream.h) interface makes it
+comparatively easy to use AIO for such use cases.
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index a1282351436..f2d763180d1 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -24,6 +24,8 @@
*
* - read_stream.c - helper for reading buffered relation data
*
+ * - README.md - higher-level overview over AIO
+ *
*
* Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
* Portions Copyright (c) 1994, Regents of the University of California
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0016-aio-Implement-smgr-md-fd-aio-methods.patch (28.5K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/17-v2.4-0016-aio-Implement-smgr-md-fd-aio-methods.patch)
download | inline diff:
From 13f0bd1fc0bc86addb30fab960853d61d137e4bd Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 16:06:51 -0500
Subject: [PATCH v2.4 16/29] aio: Implement smgr/md/fd aio methods
---
src/include/storage/aio.h | 6 +-
src/include/storage/aio_types.h | 12 +-
src/include/storage/fd.h | 6 +
src/include/storage/md.h | 12 +
src/include/storage/smgr.h | 22 ++
src/backend/storage/aio/aio_callback.c | 4 +
src/backend/storage/aio/aio_target.c | 2 +
src/backend/storage/file/fd.c | 68 +++++
src/backend/storage/smgr/md.c | 362 +++++++++++++++++++++++++
src/backend/storage/smgr/smgr.c | 126 +++++++++
10 files changed, 616 insertions(+), 4 deletions(-)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index 3c058c84003..4c5fb7bcfce 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -108,9 +108,10 @@ typedef enum PgAioTargetID
{
/* intentionally the zero value, to help catch zeroed memory etc */
PGAIO_TID_INVALID = 0,
+ PGAIO_TID_SMGR,
} PgAioTargetID;
-#define PGAIO_TID_COUNT (PGAIO_TID_INVALID + 1)
+#define PGAIO_TID_COUNT (PGAIO_TID_SMGR + 1)
/*
@@ -174,6 +175,9 @@ typedef struct PgAioTargetInfo
typedef enum PgAioHandleCallbackID
{
PGAIO_HCB_INVALID,
+
+ PGAIO_HCB_MD_READV,
+ PGAIO_HCB_MD_WRITEV,
} PgAioHandleCallbackID;
diff --git a/src/include/storage/aio_types.h b/src/include/storage/aio_types.h
index d2617139a25..762fce3f075 100644
--- a/src/include/storage/aio_types.h
+++ b/src/include/storage/aio_types.h
@@ -58,11 +58,17 @@ typedef struct PgAioWaitRef
*/
typedef union PgAioTargetData
{
- /* just as an example placeholder for later */
struct
{
- uint32 queue_id;
- } wal;
+ RelFileLocator rlocator; /* physical relation identifier */
+ BlockNumber blockNum; /* blknum relative to begin of reln */
+ BlockNumber nblocks;
+ ForkNumber forkNum:8; /* don't waste 4 byte for four values */
+ bool is_temp:1; /* proc can be inferred by owning AIO */
+ bool release_lock:1;
+ bool skip_fsync:1;
+ uint8 mode;
+ } smgr;
} PgAioTargetData;
diff --git a/src/include/storage/fd.h b/src/include/storage/fd.h
index e3067ab6597..e2fd896646e 100644
--- a/src/include/storage/fd.h
+++ b/src/include/storage/fd.h
@@ -101,6 +101,8 @@ extern PGDLLIMPORT int max_safe_fds;
* prototypes for functions in fd.c
*/
+struct PgAioHandle;
+
/* Operations on virtual Files --- equivalent to Unix kernel file ops */
extern File PathNameOpenFile(const char *fileName, int fileFlags);
extern File PathNameOpenFilePerm(const char *fileName, int fileFlags, mode_t fileMode);
@@ -109,6 +111,10 @@ extern void FileClose(File file);
extern int FilePrefetch(File file, off_t offset, off_t amount, uint32 wait_event_info);
extern ssize_t FileReadV(File file, const struct iovec *iov, int iovcnt, off_t offset, uint32 wait_event_info);
extern ssize_t FileWriteV(File file, const struct iovec *iov, int iovcnt, off_t offset, uint32 wait_event_info);
+extern ssize_t FileReadV(File file, const struct iovec *iov, int iovcnt, off_t offset, uint32 wait_event_info);
+extern int FileStartReadV(struct PgAioHandle *ioh, File file, int iovcnt, off_t offset, uint32 wait_event_info);
+extern ssize_t FileWriteV(File file, const struct iovec *iov, int iovcnt, off_t offset, uint32 wait_event_info);
+extern int FileStartWriteV(struct PgAioHandle *ioh, File file, int iovcnt, off_t offset, uint32 wait_event_info);
extern int FileSync(File file, uint32 wait_event_info);
extern int FileZero(File file, off_t offset, off_t amount, uint32 wait_event_info);
extern int FileFallocate(File file, off_t offset, off_t amount, uint32 wait_event_info);
diff --git a/src/include/storage/md.h b/src/include/storage/md.h
index 05bf537066e..7b28c3d482c 100644
--- a/src/include/storage/md.h
+++ b/src/include/storage/md.h
@@ -19,6 +19,10 @@
#include "storage/smgr.h"
#include "storage/sync.h"
+struct PgAioHandleCallbacks;
+extern const struct PgAioHandleCallbacks aio_md_readv_cb;
+extern const struct PgAioHandleCallbacks aio_md_writev_cb;
+
/* md storage manager functionality */
extern void mdinit(void);
extern void mdopen(SMgrRelation reln);
@@ -36,9 +40,16 @@ extern uint32 mdmaxcombine(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum);
extern void mdreadv(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
void **buffers, BlockNumber nblocks);
+extern void mdstartreadv(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks);
extern void mdwritev(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
const void **buffers, BlockNumber nblocks, bool skipFsync);
+extern void mdstartwritev(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks, bool skipFsync);
extern void mdwriteback(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum, BlockNumber nblocks);
extern BlockNumber mdnblocks(SMgrRelation reln, ForkNumber forknum);
@@ -46,6 +57,7 @@ extern void mdtruncate(SMgrRelation reln, ForkNumber forknum,
BlockNumber old_blocks, BlockNumber nblocks);
extern void mdimmedsync(SMgrRelation reln, ForkNumber forknum);
extern void mdregistersync(SMgrRelation reln, ForkNumber forknum);
+extern int mdfd(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off);
extern void ForgetDatabaseSyncRequests(Oid dbid);
extern void DropRelationFiles(RelFileLocator *delrels, int ndelrels, bool isRedo);
diff --git a/src/include/storage/smgr.h b/src/include/storage/smgr.h
index 4016b206ad6..86fa07b110f 100644
--- a/src/include/storage/smgr.h
+++ b/src/include/storage/smgr.h
@@ -73,6 +73,11 @@ typedef SMgrRelationData *SMgrRelation;
#define SmgrIsTemp(smgr) \
RelFileLocatorBackendIsTemp((smgr)->smgr_rlocator)
+struct PgAioHandle;
+struct PgAioTargetInfo;
+
+extern const struct PgAioTargetInfo aio_smgr_target_info;
+
extern void smgrinit(void);
extern SMgrRelation smgropen(RelFileLocator rlocator, ProcNumber backend);
extern bool smgrexists(SMgrRelation reln, ForkNumber forknum);
@@ -97,10 +102,19 @@ extern uint32 smgrmaxcombine(SMgrRelation reln, ForkNumber forknum,
extern void smgrreadv(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
void **buffers, BlockNumber nblocks);
+extern void smgrstartreadv(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks);
extern void smgrwritev(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
const void **buffers, BlockNumber nblocks,
bool skipFsync);
+extern void smgrstartwritev(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks,
+ bool skipFsync);
extern void smgrwriteback(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum, BlockNumber nblocks);
extern BlockNumber smgrnblocks(SMgrRelation reln, ForkNumber forknum);
@@ -110,6 +124,7 @@ extern void smgrtruncate(SMgrRelation reln, ForkNumber *forknum, int nforks,
BlockNumber *nblocks);
extern void smgrimmedsync(SMgrRelation reln, ForkNumber forknum);
extern void smgrregistersync(SMgrRelation reln, ForkNumber forknum);
+extern int smgrfd(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off);
extern void AtEOXact_SMgr(void);
extern bool ProcessBarrierSmgrRelease(void);
@@ -127,4 +142,11 @@ smgrwrite(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
smgrwritev(reln, forknum, blocknum, &buffer, 1, skipFsync);
}
+extern void pgaio_io_set_target_smgr(struct PgAioHandle *ioh,
+ SMgrRelationData *smgr,
+ ForkNumber forknum,
+ BlockNumber blocknum,
+ int nblocks,
+ bool skip_fsync);
+
#endif /* SMGR_H */
diff --git a/src/backend/storage/aio/aio_callback.c b/src/backend/storage/aio/aio_callback.c
index 5629dc4cc94..adb8050eb58 100644
--- a/src/backend/storage/aio/aio_callback.c
+++ b/src/backend/storage/aio/aio_callback.c
@@ -18,6 +18,7 @@
#include "miscadmin.h"
#include "storage/aio.h"
#include "storage/aio_internal.h"
+#include "storage/md.h"
#include "utils/memutils.h"
@@ -38,6 +39,9 @@ typedef struct PgAioHandleCallbacksEntry
static const PgAioHandleCallbacksEntry aio_handle_cbs[] = {
#define CALLBACK_ENTRY(id, callback) [id] = {.cb = &callback, .name = #callback}
CALLBACK_ENTRY(PGAIO_HCB_INVALID, aio_invalid_cb),
+
+ CALLBACK_ENTRY(PGAIO_HCB_MD_READV, aio_md_readv_cb),
+ CALLBACK_ENTRY(PGAIO_HCB_MD_WRITEV, aio_md_writev_cb),
#undef CALLBACK_ENTRY
};
diff --git a/src/backend/storage/aio/aio_target.c b/src/backend/storage/aio/aio_target.c
index 15428968e58..a43edd89890 100644
--- a/src/backend/storage/aio/aio_target.c
+++ b/src/backend/storage/aio/aio_target.c
@@ -18,6 +18,7 @@
#include "storage/aio.h"
#include "storage/aio_internal.h"
+#include "storage/smgr.h"
/*
@@ -31,6 +32,7 @@ static const PgAioTargetInfo *pgaio_target_info[] = {
[PGAIO_TID_INVALID] = &(PgAioTargetInfo) {
.name = "invalid",
},
+ [PGAIO_TID_SMGR] = &aio_smgr_target_info,
};
diff --git a/src/backend/storage/file/fd.c b/src/backend/storage/file/fd.c
index e454db4c020..a9c90bc4e59 100644
--- a/src/backend/storage/file/fd.c
+++ b/src/backend/storage/file/fd.c
@@ -94,6 +94,7 @@
#include "miscadmin.h"
#include "pgstat.h"
#include "postmaster/startup.h"
+#include "storage/aio.h"
#include "storage/fd.h"
#include "storage/ipc.h"
#include "utils/guc.h"
@@ -1294,6 +1295,8 @@ LruDelete(File file)
vfdP = &VfdCache[file];
+ pgaio_closing_fd(vfdP->fd);
+
/*
* Close the file. We aren't expecting this to fail; if it does, better
* to leak the FD than to mess up our internal state.
@@ -1987,6 +1990,8 @@ FileClose(File file)
if (!FileIsNotOpen(file))
{
+ pgaio_closing_fd(vfdP->fd);
+
/* close the file */
if (close(vfdP->fd) != 0)
{
@@ -2210,6 +2215,32 @@ retry:
return returnCode;
}
+int
+FileStartReadV(struct PgAioHandle *ioh, File file,
+ int iovcnt, off_t offset,
+ uint32 wait_event_info)
+{
+ int returnCode;
+ Vfd *vfdP;
+
+ Assert(FileIsValid(file));
+
+ DO_DB(elog(LOG, "FileStartReadV: %d (%s) " INT64_FORMAT " %d",
+ file, VfdCache[file].fileName,
+ (int64) offset,
+ iovcnt));
+
+ returnCode = FileAccess(file);
+ if (returnCode < 0)
+ return returnCode;
+
+ vfdP = &VfdCache[file];
+
+ pgaio_io_prep_readv(ioh, vfdP->fd, iovcnt, offset);
+
+ return 0;
+}
+
ssize_t
FileWriteV(File file, const struct iovec *iov, int iovcnt, off_t offset,
uint32 wait_event_info)
@@ -2315,6 +2346,34 @@ retry:
return returnCode;
}
+int
+FileStartWriteV(struct PgAioHandle *ioh, File file,
+ int iovcnt, off_t offset,
+ uint32 wait_event_info)
+{
+ int returnCode;
+ Vfd *vfdP;
+
+ Assert(FileIsValid(file));
+
+ DO_DB(elog(LOG, "FileStartWriteV: %d (%s) " INT64_FORMAT " %d",
+ file, VfdCache[file].fileName,
+ (int64) offset,
+ iovcnt));
+
+ returnCode = FileAccess(file);
+ if (returnCode < 0)
+ return returnCode;
+
+ vfdP = &VfdCache[file];
+
+ /* FIXME: think about / reimplement temp_file_limit */
+
+ pgaio_io_prep_writev(ioh, vfdP->fd, iovcnt, offset);
+
+ return 0;
+}
+
int
FileSync(File file, uint32 wait_event_info)
{
@@ -2498,6 +2557,12 @@ FilePathName(File file)
int
FileGetRawDesc(File file)
{
+ int returnCode;
+
+ returnCode = FileAccess(file);
+ if (returnCode < 0)
+ return returnCode;
+
Assert(FileIsValid(file));
return VfdCache[file].fd;
}
@@ -2778,6 +2843,7 @@ FreeDesc(AllocateDesc *desc)
result = closedir(desc->desc.dir);
break;
case AllocateDescRawFD:
+ pgaio_closing_fd(desc->desc.fd);
result = close(desc->desc.fd);
break;
default:
@@ -2846,6 +2912,8 @@ CloseTransientFile(int fd)
/* Only get here if someone passes us a file not in allocatedDescs */
elog(WARNING, "fd passed to CloseTransientFile was not obtained from OpenTransientFile");
+ pgaio_closing_fd(fd);
+
return close(fd);
}
diff --git a/src/backend/storage/smgr/md.c b/src/backend/storage/smgr/md.c
index 7bf0b45e2c3..db508d63573 100644
--- a/src/backend/storage/smgr/md.c
+++ b/src/backend/storage/smgr/md.c
@@ -31,6 +31,7 @@
#include "miscadmin.h"
#include "pg_trace.h"
#include "pgstat.h"
+#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
#include "storage/md.h"
@@ -132,6 +133,22 @@ static MdfdVec *_mdfd_getseg(SMgrRelation reln, ForkNumber forknum,
static BlockNumber _mdnblocks(SMgrRelation reln, ForkNumber forknum,
MdfdVec *seg);
+static PgAioResult md_readv_complete(PgAioHandle *ioh, PgAioResult prior_result);
+static void md_readv_report(PgAioResult result, const PgAioTargetData *target_data, int elevel);
+static PgAioResult md_writev_complete(PgAioHandle *ioh, PgAioResult prior_result);
+static void md_writev_report(PgAioResult result, const PgAioTargetData *target_data, int elevel);
+
+const struct PgAioHandleCallbacks aio_md_readv_cb = {
+ .complete_shared = md_readv_complete,
+ .report = md_readv_report,
+};
+
+const struct PgAioHandleCallbacks aio_md_writev_cb = {
+ .complete_shared = md_writev_complete,
+ .report = md_writev_report,
+};
+
+
static inline int
_mdfd_open_flags(void)
{
@@ -927,6 +944,53 @@ mdreadv(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
}
}
+void
+mdstartreadv(PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks)
+{
+ off_t seekpos;
+ MdfdVec *v;
+ BlockNumber nblocks_this_segment;
+ struct iovec *iov;
+ int iovcnt;
+
+ v = _mdfd_getseg(reln, forknum, blocknum, false,
+ EXTENSION_FAIL | EXTENSION_CREATE_RECOVERY);
+
+ seekpos = (off_t) BLCKSZ * (blocknum % ((BlockNumber) RELSEG_SIZE));
+
+ Assert(seekpos < (off_t) BLCKSZ * RELSEG_SIZE);
+
+ nblocks_this_segment =
+ Min(nblocks,
+ RELSEG_SIZE - (blocknum % ((BlockNumber) RELSEG_SIZE)));
+
+ if (nblocks_this_segment != nblocks)
+ elog(ERROR, "read crossing segment boundary");
+
+ iovcnt = pgaio_io_get_iovec(ioh, &iov);
+
+ Assert(nblocks <= iovcnt);
+
+ iovcnt = buffers_to_iovec(iov, buffers, nblocks_this_segment);
+
+ Assert(iovcnt <= nblocks_this_segment);
+
+ if (!(io_direct_flags & IO_DIRECT_DATA))
+ pgaio_io_set_flag(ioh, PGAIO_HF_BUFFERED);
+
+ pgaio_io_set_target_smgr(ioh,
+ reln,
+ forknum,
+ blocknum,
+ nblocks,
+ false);
+ pgaio_io_register_callbacks(ioh, PGAIO_HCB_MD_READV);
+
+ FileStartReadV(ioh, v->mdfd_vfd, iovcnt, seekpos, WAIT_EVENT_DATA_FILE_READ);
+}
+
/*
* mdwritev() -- Write the supplied blocks at the appropriate location.
*
@@ -1032,6 +1096,53 @@ mdwritev(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
}
}
+void
+mdstartwritev(PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks, bool skipFsync)
+{
+ off_t seekpos;
+ MdfdVec *v;
+ BlockNumber nblocks_this_segment;
+ struct iovec *iov;
+ int iovcnt;
+
+ v = _mdfd_getseg(reln, forknum, blocknum, false,
+ EXTENSION_FAIL | EXTENSION_CREATE_RECOVERY);
+
+ seekpos = (off_t) BLCKSZ * (blocknum % ((BlockNumber) RELSEG_SIZE));
+
+ Assert(seekpos < (off_t) BLCKSZ * RELSEG_SIZE);
+
+ nblocks_this_segment =
+ Min(nblocks,
+ RELSEG_SIZE - (blocknum % ((BlockNumber) RELSEG_SIZE)));
+
+ if (nblocks_this_segment != nblocks)
+ elog(ERROR, "write crossing segment boundary");
+
+ iovcnt = pgaio_io_get_iovec(ioh, &iov);
+
+ Assert(nblocks <= iovcnt);
+
+ iovcnt = buffers_to_iovec(iov, unconstify(void **, buffers), nblocks_this_segment);
+
+ Assert(iovcnt <= nblocks_this_segment);
+
+ if (!(io_direct_flags & IO_DIRECT_DATA))
+ pgaio_io_set_flag(ioh, PGAIO_HF_BUFFERED);
+
+ pgaio_io_set_target_smgr(ioh,
+ reln,
+ forknum,
+ blocknum,
+ nblocks,
+ skipFsync);
+ pgaio_io_register_callbacks(ioh, PGAIO_HCB_MD_WRITEV);
+
+ FileStartWriteV(ioh, v->mdfd_vfd, iovcnt, seekpos, WAIT_EVENT_DATA_FILE_WRITE);
+}
+
/*
* mdwriteback() -- Tell the kernel to write pages back to storage.
@@ -1355,6 +1466,21 @@ mdimmedsync(SMgrRelation reln, ForkNumber forknum)
}
}
+int
+mdfd(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off)
+{
+ MdfdVec *v = mdopenfork(reln, forknum, EXTENSION_FAIL);
+
+ v = _mdfd_getseg(reln, forknum, blocknum, false,
+ EXTENSION_FAIL);
+
+ *off = (off_t) BLCKSZ * (blocknum % ((BlockNumber) RELSEG_SIZE));
+
+ Assert(*off < (off_t) BLCKSZ * RELSEG_SIZE);
+
+ return FileGetRawDesc(v->mdfd_vfd);
+}
+
/*
* register_dirty_segment() -- Mark a relation segment as needing fsync
*
@@ -1405,6 +1531,35 @@ register_dirty_segment(SMgrRelation reln, ForkNumber forknum, MdfdVec *seg)
}
}
+/*
+ * Like register_dirty_segment(), except for use by AIO. In the completion
+ * callback we don't have access to the MdfdVec (the completion callback might
+ * be executed in a different backend than the issuing backend), therefore we
+ * have to implement this slightly differently.
+ */
+static void
+register_dirty_segment_aio(RelFileLocator locator, ForkNumber forknum, uint64 segno)
+{
+ FileTag tag;
+
+ INIT_MD_FILETAG(tag, locator, forknum, segno);
+
+ if (!RegisterSyncRequest(&tag, SYNC_REQUEST, false /* retryOnError */ ))
+ {
+ char path[MAXPGPATH];
+
+ ereport(DEBUG1,
+ (errmsg_internal("could not forward fsync request because request queue is full")));
+
+ /* reuse mdsyncfiletag() to avoid duplicating code */
+ if (mdsyncfiletag(&tag, path))
+ ereport(data_sync_elevel(ERROR),
+ (errcode_for_file_access(),
+ errmsg("could not fsync file \"%s\": %m",
+ path)));
+ }
+}
+
/*
* register_unlink_segment() -- Schedule a file to be deleted after next checkpoint
*/
@@ -1838,3 +1993,210 @@ mdfiletagmatches(const FileTag *ftag, const FileTag *candidate)
*/
return ftag->rlocator.dbOid == candidate->rlocator.dbOid;
}
+
+/*
+ * AIO completion callback for mdstartreadv().
+ */
+static PgAioResult
+md_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ PgAioTargetData *td = pgaio_io_get_target_data(ioh);
+ PgAioResult result = prior_result;
+
+ if (prior_result.result < 0)
+ {
+ result.status = ARS_ERROR;
+ result.id = PGAIO_HCB_MD_READV;
+ /* For "hard" errors, track the error number in error_data */
+ result.error_data = -prior_result.result;
+ result.result = 0;
+
+ md_readv_report(result, td, LOG);
+
+ return result;
+ }
+
+ /*
+ * The smgr API operates in blocks, therefore convert the result from
+ * bytes to blocks.
+ */
+ result.result /= BLCKSZ;
+
+ if (result.result == 0)
+ {
+ /* consider 0 blocks read a failure */
+ result.status = ARS_ERROR;
+ result.id = PGAIO_HCB_MD_READV;
+ result.error_data = 0;
+
+ md_readv_report(result, td, LOG);
+
+ return result;
+ }
+
+ if (result.status != ARS_ERROR &&
+ result.result < td->smgr.nblocks)
+ {
+ /* partial reads should be retried at upper level */
+ result.status = ARS_PARTIAL;
+ result.id = PGAIO_HCB_MD_READV;
+ }
+
+ return result;
+}
+
+/*
+ * AIO error reporting callback for mdstartreadv().
+ */
+static void
+md_readv_report(PgAioResult result, const PgAioTargetData *td, int elevel)
+{
+ MemoryContext oldContext = CurrentMemoryContext;
+ char *path;
+
+ /* AFIXME: */
+ oldContext = MemoryContextSwitchTo(ErrorContext);
+
+ path = relpathbackend(td->smgr.rlocator,
+ td->smgr.is_temp ? MyProcNumber : INVALID_PROC_NUMBER,
+ td->smgr.forkNum);
+
+ if (result.error_data != 0)
+ {
+ errno = result.error_data; /* for errcode_for_file_access() */
+
+ ereport(elevel,
+ errcode_for_file_access(),
+ errmsg("could not read blocks %u..%u in file \"%s\": %m",
+ td->smgr.blockNum,
+ td->smgr.blockNum + td->smgr.nblocks,
+ path
+ )
+ );
+ }
+ else
+ {
+ /*
+ * NB: This will typically only be output in debug messages, while
+ * retrying a partial IO.
+ */
+ ereport(elevel,
+ errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("could not read blocks %u..%u in file \"%s\": read only %zu of %zu bytes",
+ td->smgr.blockNum,
+ td->smgr.blockNum + td->smgr.nblocks - 1,
+ path,
+ result.result * (size_t) BLCKSZ,
+ td->smgr.nblocks * (size_t) BLCKSZ
+ )
+ );
+ }
+
+ pfree(path);
+ MemoryContextSwitchTo(oldContext);
+}
+
+/*
+ * AIO completion callback for mdstartwritev().
+ */
+static PgAioResult
+md_writev_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ PgAioTargetData *td = pgaio_io_get_target_data(ioh);
+ PgAioResult result = prior_result;
+
+ if (prior_result.result < 0)
+ {
+ result.status = ARS_ERROR;
+ result.id = PGAIO_HCB_MD_WRITEV;
+ /* For "hard" errors, track the error number in error_data */
+ result.error_data = -prior_result.result;
+ result.result = 0;
+
+ md_writev_report(result, td, LOG);
+
+ return result;
+ }
+
+ /*
+ * The smgr API operates in blocks, therefore convert the result from
+ * bytes to blocks.
+ */
+ result.result /= BLCKSZ;
+
+ if (result.result == 0)
+ {
+ /* consider 0 blocks written a failure */
+ result.status = ARS_ERROR;
+ result.id = PGAIO_HCB_MD_WRITEV;
+ result.error_data = 0;
+
+ md_writev_report(result, td, LOG);
+
+ return result;
+ }
+
+ if (result.status != ARS_ERROR &&
+ result.result < td->smgr.nblocks)
+ {
+ /* partial writes should be retried at upper level */
+ result.status = ARS_PARTIAL;
+ result.id = PGAIO_HCB_MD_WRITEV;
+ }
+
+ if (!td->smgr.skip_fsync)
+ register_dirty_segment_aio(td->smgr.rlocator, td->smgr.forkNum,
+ td->smgr.blockNum / ((BlockNumber) RELSEG_SIZE));
+
+ return result;
+}
+
+/*
+ * AIO error reporting callback for mdstartwritev().
+ */
+static void
+md_writev_report(PgAioResult result, const PgAioTargetData *td, int elevel)
+{
+ MemoryContext oldContext = CurrentMemoryContext;
+ char *path;
+
+ /* AFIXME: */
+ oldContext = MemoryContextSwitchTo(ErrorContext);
+
+ path = relpathbackend(td->smgr.rlocator,
+ td->smgr.is_temp ? MyProcNumber : INVALID_PROC_NUMBER,
+ td->smgr.forkNum);
+
+ if (result.error_data != 0)
+ {
+ errno = result.error_data; /* for errcode_for_file_access() */
+
+ ereport(elevel,
+ errcode_for_file_access(),
+ errmsg("could not write blocks %u..%u in file \"%s\": %m",
+ td->smgr.blockNum,
+ td->smgr.blockNum + td->smgr.nblocks,
+ path)
+ );
+ }
+ else
+ {
+ /*
+ * NB: This will typically only be output in debug messages, while
+ * retrying a partial IO.
+ */
+ ereport(elevel,
+ errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("could not write blocks %u..%u in file \"%s\": wrote only %zu of %zu bytes",
+ td->smgr.blockNum,
+ td->smgr.blockNum + td->smgr.nblocks - 1,
+ path,
+ result.result * (size_t) BLCKSZ,
+ td->smgr.nblocks * (size_t) BLCKSZ
+ )
+ );
+ }
+
+ pfree(path);
+ MemoryContextSwitchTo(oldContext);
+}
diff --git a/src/backend/storage/smgr/smgr.c b/src/backend/storage/smgr/smgr.c
index ebe35c04de5..fb231e6ad48 100644
--- a/src/backend/storage/smgr/smgr.c
+++ b/src/backend/storage/smgr/smgr.c
@@ -53,6 +53,7 @@
#include "access/xlogutils.h"
#include "lib/ilist.h"
+#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
#include "storage/md.h"
@@ -93,10 +94,19 @@ typedef struct f_smgr
void (*smgr_readv) (SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
void **buffers, BlockNumber nblocks);
+ void (*smgr_startreadv) (struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks);
void (*smgr_writev) (SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
const void **buffers, BlockNumber nblocks,
bool skipFsync);
+ void (*smgr_startwritev) (struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks,
+ bool skipFsync);
void (*smgr_writeback) (SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum, BlockNumber nblocks);
BlockNumber (*smgr_nblocks) (SMgrRelation reln, ForkNumber forknum);
@@ -104,6 +114,7 @@ typedef struct f_smgr
BlockNumber old_blocks, BlockNumber nblocks);
void (*smgr_immedsync) (SMgrRelation reln, ForkNumber forknum);
void (*smgr_registersync) (SMgrRelation reln, ForkNumber forknum);
+ int (*smgr_fd) (SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off);
} f_smgr;
static const f_smgr smgrsw[] = {
@@ -121,12 +132,15 @@ static const f_smgr smgrsw[] = {
.smgr_prefetch = mdprefetch,
.smgr_maxcombine = mdmaxcombine,
.smgr_readv = mdreadv,
+ .smgr_startreadv = mdstartreadv,
.smgr_writev = mdwritev,
+ .smgr_startwritev = mdstartwritev,
.smgr_writeback = mdwriteback,
.smgr_nblocks = mdnblocks,
.smgr_truncate = mdtruncate,
.smgr_immedsync = mdimmedsync,
.smgr_registersync = mdregistersync,
+ .smgr_fd = mdfd,
}
};
@@ -145,6 +159,16 @@ static void smgrshutdown(int code, Datum arg);
static void smgrdestroy(SMgrRelation reln);
+static void smgr_aio_reopen(PgAioHandle *ioh);
+static char *smgr_aio_describe_identity(const PgAioTargetData *sd);
+
+const struct PgAioTargetInfo aio_smgr_target_info = {
+ .name = "smgr",
+ .reopen = smgr_aio_reopen,
+ .describe_identity = smgr_aio_describe_identity,
+};
+
+
/*
* smgrinit(), smgrshutdown() -- Initialize or shut down storage
* managers.
@@ -623,6 +647,19 @@ smgrreadv(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
nblocks);
}
+/*
+ * AFIXME: FILL ME IN
+ */
+void
+smgrstartreadv(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks)
+{
+ smgrsw[reln->smgr_which].smgr_startreadv(ioh,
+ reln, forknum, blocknum, buffers,
+ nblocks);
+}
+
/*
* smgrwritev() -- Write the supplied buffers out.
*
@@ -657,6 +694,19 @@ smgrwritev(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
buffers, nblocks, skipFsync);
}
+/*
+ * AFIXME: FILL ME IN
+ */
+void
+smgrstartwritev(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks, bool skipFsync)
+{
+ smgrsw[reln->smgr_which].smgr_startwritev(ioh,
+ reln, forknum, blocknum, buffers,
+ nblocks, skipFsync);
+}
+
/*
* smgrwriteback() -- Trigger kernel writeback for the supplied range of
* blocks.
@@ -819,6 +869,12 @@ smgrimmedsync(SMgrRelation reln, ForkNumber forknum)
smgrsw[reln->smgr_which].smgr_immedsync(reln, forknum);
}
+int
+smgrfd(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off)
+{
+ return smgrsw[reln->smgr_which].smgr_fd(reln, forknum, blocknum, off);
+}
+
/*
* AtEOXact_SMgr
*
@@ -847,3 +903,73 @@ ProcessBarrierSmgrRelease(void)
smgrreleaseall();
return true;
}
+
+void
+pgaio_io_set_target_smgr(PgAioHandle *ioh,
+ struct SMgrRelationData *smgr,
+ ForkNumber forknum,
+ BlockNumber blocknum,
+ int nblocks,
+ bool skip_fsync)
+{
+ PgAioTargetData *sd = pgaio_io_get_target_data(ioh);
+
+ pgaio_io_set_target(ioh, PGAIO_TID_SMGR);
+
+ /* backend is implied via IO owner */
+ sd->smgr.rlocator = smgr->smgr_rlocator.locator;
+ sd->smgr.forkNum = forknum;
+ sd->smgr.blockNum = blocknum;
+ sd->smgr.nblocks = nblocks;
+ sd->smgr.is_temp = SmgrIsTemp(smgr);
+ sd->smgr.release_lock = false;
+ /* Temp relations should never be fsync'd */
+ sd->smgr.skip_fsync = skip_fsync && !SmgrIsTemp(smgr);
+ sd->smgr.mode = RBM_NORMAL;
+}
+
+static void
+smgr_aio_reopen(PgAioHandle *ioh)
+{
+ PgAioTargetData *sd = pgaio_io_get_target_data(ioh);
+ PgAioOpData *od = pgaio_io_get_op_data(ioh);
+ SMgrRelation reln;
+ ProcNumber procno;
+ uint32 off;
+
+ if (sd->smgr.is_temp)
+ procno = pgaio_io_get_owner(ioh);
+ else
+ procno = INVALID_PROC_NUMBER;
+
+ reln = smgropen(sd->smgr.rlocator, procno);
+ od->read.fd = smgrfd(reln, sd->smgr.forkNum, sd->smgr.blockNum, &off);
+ Assert(off == od->read.offset);
+}
+
+static char *
+smgr_aio_describe_identity(const PgAioTargetData *sd)
+{
+ char *path;
+ char *desc;
+
+ path = relpathbackend(sd->smgr.rlocator,
+ sd->smgr.is_temp ? MyProcNumber : INVALID_PROC_NUMBER,
+ sd->smgr.forkNum);
+
+ if (sd->smgr.nblocks == 0)
+ desc = psprintf(_("file \"%s\""), path);
+ else if (sd->smgr.nblocks == 1)
+ desc = psprintf(_("block %u in file \"%s\""),
+ sd->smgr.blockNum,
+ path);
+ else
+ desc = psprintf(_("blocks %u..%u in file \"%s\""),
+ sd->smgr.blockNum,
+ sd->smgr.blockNum + sd->smgr.nblocks - 1,
+ path);
+
+ pfree(path);
+
+ return desc;
+}
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0017-aio-Add-pg_aios-view.patch (9.8K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/18-v2.4-0017-aio-Add-pg_aios-view.patch)
download | inline diff:
From 75eb063ade46283fe4ce29b969954db051bcd3cc Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 13:44:40 -0500
Subject: [PATCH v2.4 17/29] aio: Add pg_aios view
Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/catalog/pg_proc.dat | 10 ++
src/backend/catalog/system_views.sql | 3 +
src/backend/storage/aio/Makefile | 1 +
src/backend/storage/aio/aio_funcs.c | 222 +++++++++++++++++++++++++++
src/backend/storage/aio/meson.build | 1 +
src/test/regress/expected/rules.out | 17 ++
6 files changed, 254 insertions(+)
create mode 100644 src/backend/storage/aio/aio_funcs.c
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 9e803d610d7..fa7191e7f9a 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12464,4 +12464,14 @@
proargtypes => 'int4',
prosrc => 'gist_stratnum_common' },
+# AIO related functions
+{ oid => '9200', descr => 'information about in-progress asynchronous IOs',
+ proname => 'pg_get_aios', prorows => '100', proretset => 't',
+ provolatile => 'v', proparallel => 'r', prorettype => 'record', proargtypes => '',
+ proallargtypes => '{int4,int4,int8,text,text,int8,int8,text,int2,int4,text,text,text,bool,bool,bool}',
+ proargmodes => '{o,o,o,o,o,o,o,o,o,o,o,o,o,o,o,o}',
+ proargnames => '{pid,io_id,io_generation,state,operation,offset,length,target,handle_data_len,raw_result,result,error_desc,target_desc,f_sync,f_localmem,f_buffered}',
+ prosrc => 'pg_get_aios' },
+
+
]
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index eff0990957e..b4140f4c46e 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1394,3 +1394,6 @@ CREATE VIEW pg_stat_subscription_stats AS
CREATE VIEW pg_wait_events AS
SELECT * FROM pg_get_wait_events();
+
+CREATE VIEW pg_aios AS
+ SELECT * FROM pg_get_aios();
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index c06c50771e0..3f2469cc399 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -11,6 +11,7 @@ include $(top_builddir)/src/Makefile.global
OBJS = \
aio.o \
aio_callback.o \
+ aio_funcs.o \
aio_init.o \
aio_io.o \
aio_target.o \
diff --git a/src/backend/storage/aio/aio_funcs.c b/src/backend/storage/aio/aio_funcs.c
new file mode 100644
index 00000000000..cb4f3d29201
--- /dev/null
+++ b/src/backend/storage/aio/aio_funcs.c
@@ -0,0 +1,222 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_funcs.c
+ * AIO - SQL interface for AIO
+ *
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio_funcs.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "utils/builtins.h"
+#include "funcapi.h"
+#include "storage/proc.h"
+
+
+/*
+ * Byte length of an iovec.
+ */
+static size_t
+iov_byte_length(const struct iovec *iov, int cnt)
+{
+ size_t len = 0;
+
+ for (int i = 0; i < cnt; i++)
+ {
+ len += iov[i].iov_len;
+ }
+
+ return len;
+}
+
+Datum
+pg_get_aios(PG_FUNCTION_ARGS)
+{
+ ReturnSetInfo *rsinfo = (ReturnSetInfo *) fcinfo->resultinfo;
+
+ InitMaterializedSRF(fcinfo, 0);
+
+#define PG_GET_AIOS_COLS 16
+
+ for (uint64 i = 0; i < pgaio_ctl->io_handle_count; i++)
+ {
+ PgAioHandle *live_ioh = &pgaio_ctl->io_handles[i];
+ uint32 ioh_id = pgaio_io_get_id(live_ioh);
+ Datum values[PG_GET_AIOS_COLS] = {0};
+ bool nulls[PG_GET_AIOS_COLS] = {0};
+ ProcNumber owner;
+ PGPROC *owner_proc;
+ int32 owner_pid;
+ PgAioHandleState start_state;
+ uint64 start_generation;
+ PgAioHandle ioh_copy;
+ struct iovec iov_copy[PG_IOV_MAX];
+
+retry:
+
+ /*
+ * There is no lock that could prevent the state of the IO to advance
+ * concurrently - and we don't want to introduce one, as that would
+ * introduce atomics into a very common path. Instead we
+ *
+ * 1) determine the state + generation of the IO
+ *
+ * 2) copy the IO to local memory
+ *
+ * 3) check if state and generation of the IO changed
+ */
+
+ /* 1) from above */
+ start_generation = live_ioh->generation;
+ pg_read_barrier();
+ start_state = live_ioh->state;
+
+ if (start_state == PGAIO_HS_IDLE)
+ continue;
+
+ /* 2) from above */
+ memcpy(&ioh_copy, live_ioh, sizeof(PgAioHandle));
+
+ /*
+ * Safe to copy even if no iovec is used - we always reserve the
+ * required space.
+ */
+ memcpy(&iov_copy, &pgaio_ctl->iovecs[ioh_copy.iovec_off],
+ PG_IOV_MAX * sizeof(struct iovec));
+
+ /*
+ * Copy information about owner before 3) below, if the process exited
+ * it'd have to wait for the IO to finish first, which we would detect
+ * in 3).
+ */
+ owner = ioh_copy.owner_procno;
+ owner_proc = GetPGProcByNumber(owner);
+ owner_pid = owner_proc->pid;
+
+ /* 3) from above */
+ pg_read_barrier();
+
+ /*
+ * The IO completed and a new one was started with the same ID. Don't
+ * display it - it really started after this function was called.
+ * There be a risk of a livelock if we just retried endlessly, if IOs
+ * complete very quickly.
+ */
+ if (live_ioh->generation != start_generation)
+ continue;
+
+ /*
+ * The IOs state changed while we were "rendering" it. Just start from
+ * scratch. There's no risk of a livelock here, as an IO has a limited
+ * sets of states it can be in, and state changes go only in a single
+ * direction.
+ */
+ if (live_ioh->state != start_state)
+ goto retry;
+
+ /*
+ * Now that we have copied the IO into local memory and checked that
+ * it's still in the same state, we are not allowed to access "live"
+ * memory anymore. To make it slightly easier to catch such cases, set
+ * the "live" pointers to NULL.
+ */
+ live_ioh = NULL;
+ owner_proc = NULL;
+
+
+ /* column: owning pid */
+ if (owner_pid != 0)
+ values[0] = Int32GetDatum(owner_pid);
+ else
+ nulls[0] = false;
+
+ /* column: IO's id */
+ values[1] = ioh_id;
+
+ /* column: IO's generation */
+ values[2] = Int64GetDatum(start_generation);
+
+ /* column: IO's state */
+ values[3] = CStringGetTextDatum(pgaio_io_get_state_name(&ioh_copy));
+
+ /*
+ * If the IO is in PGAIO_HS_HANDED_OUT state, none of it's fields are
+ * valid yet (or are in the process of being set). Therefore we don't
+ * want to display any other columns.
+ */
+ if (start_state == PGAIO_HS_HANDED_OUT)
+ {
+ memset(nulls + 4, 1, (lengthof(nulls) - 4) * sizeof(bool));
+ goto display;
+ }
+
+ /* column: IO's operation */
+ values[4] = CStringGetTextDatum(pgaio_io_get_op_name(&ioh_copy));
+
+ /* columns: details about the IO's operation */
+ switch (ioh_copy.op)
+ {
+ case PGAIO_OP_INVALID:
+ nulls[5] = true;
+ nulls[6] = true;
+ break;
+ case PGAIO_OP_READV:
+ values[5] = Int64GetDatum(ioh_copy.op_data.read.offset);
+ values[6] =
+ Int64GetDatum(iov_byte_length(iov_copy, ioh_copy.op_data.read.iov_length));
+ break;
+ case PGAIO_OP_WRITEV:
+ values[5] = Int64GetDatum(ioh_copy.op_data.write.offset);
+ values[6] =
+ Int64GetDatum(iov_byte_length(iov_copy, ioh_copy.op_data.write.iov_length));
+ break;
+ }
+
+ /* column: IO's target */
+ values[7] = CStringGetTextDatum(pgaio_io_get_target_name(&ioh_copy));
+
+ /* column: length of IO's data array */
+ values[8] = Int16GetDatum(ioh_copy.handle_data_len);
+
+ /* column: raw result (i.e. some form of syscall return value) */
+ if (start_state == PGAIO_HS_COMPLETED_IO
+ || start_state == PGAIO_HS_COMPLETED_SHARED)
+ values[9] = Int32GetDatum(ioh_copy.result);
+ else
+ nulls[9] = true;
+
+ /*
+ * column: result in the higher level representation (unknown if not
+ * finished
+ */
+ values[10] =
+ CStringGetTextDatum(pgaio_result_status_string(ioh_copy.distilled_result.status));
+
+ /* column: error description */
+ /* AFIXME: implement */
+ nulls[11] = true;
+
+ /* column: target description */
+ values[12] = CStringGetTextDatum(pgaio_io_get_target_description(&ioh_copy));
+
+ /* columns: one for each flag */
+ values[13] = BoolGetDatum(ioh_copy.flags & PGAIO_HF_SYNCHRONOUS);
+ values[14] = BoolGetDatum(ioh_copy.flags & PGAIO_HF_REFERENCES_LOCAL);
+ values[15] = BoolGetDatum(ioh_copy.flags & PGAIO_HF_BUFFERED);
+
+display:
+
+ tuplestore_putvalues(rsinfo->setResult, rsinfo->setDesc, values, nulls);
+ }
+
+ return (Datum) 0;
+}
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 2f0f03d8071..da6df2d3654 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -3,6 +3,7 @@
backend_sources += files(
'aio.c',
'aio_callback.c',
+ 'aio_funcs.c',
'aio_init.c',
'aio_io.c',
'aio_target.c',
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 5baba8d39ff..918bf9efde3 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1286,6 +1286,23 @@ drop table cchild;
SELECT viewname, definition FROM pg_views
WHERE schemaname = 'pg_catalog'
ORDER BY viewname;
+pg_aios| SELECT pid,
+ io_id,
+ io_generation,
+ state,
+ operation,
+ "offset",
+ length,
+ target,
+ handle_data_len,
+ raw_result,
+ result,
+ error_desc,
+ target_desc,
+ f_sync,
+ f_localmem,
+ f_buffered
+ FROM pg_get_aios() pg_get_aios(pid, io_id, io_generation, state, operation, "offset", length, target, handle_data_len, raw_result, result, error_desc, target_desc, f_sync, f_localmem, f_buffered);
pg_available_extension_versions| SELECT e.name,
e.version,
(x.extname IS NOT NULL) AS installed,
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0018-WIP-localbuf-Track-pincount-in-BufferDesc-as-we.patch (7.7K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/19-v2.4-0018-WIP-localbuf-Track-pincount-in-BufferDesc-as-we.patch)
download | inline diff:
From b1e2c462aacb49c6c9d093d5d4fc578cf4001348 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 13:44:44 -0500
Subject: [PATCH v2.4 18/29] WIP: localbuf: Track pincount in BufferDesc as
well
For AIO on temp tables the AIO subsystem needs to be able to ensure a pin on a
buffer while AIO is going on, even if the IO issuing query errors out. To do
so, track the refcount in BufferDesc.state, not ust LocalRefCount.
Note that we still don't need locking, AIO completion callbacks for local
buffers are executed in the issuing session (nobody else has access to the
BufferDesc).
---
src/backend/storage/buffer/bufmgr.c | 30 ++++++--
src/backend/storage/buffer/localbuf.c | 99 +++++++++++++++++----------
2 files changed, 87 insertions(+), 42 deletions(-)
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 47e1c3442b4..ec308557179 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5444,8 +5444,20 @@ ConditionalLockBufferForCleanup(Buffer buffer)
Assert(refcount > 0);
if (refcount != 1)
return false;
- /* Nobody else to wait for */
- return true;
+
+ bufHdr = GetLocalBufferDescriptor(-buffer - 1);
+ buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ /*
+ * Check that the AIO subsystem doesn't have a pin. Likely not
+ * possible today, but better safe than sorry.
+ */
+ refcount = BUF_STATE_GET_REFCOUNT(buf_state);
+ Assert(refcount > 0);
+ if (refcount == 1)
+ return true;
+
+ return false;
}
/* There should be exactly one local pin */
@@ -5497,8 +5509,18 @@ IsBufferCleanupOK(Buffer buffer)
/* There should be exactly one pin */
if (LocalRefCount[-buffer - 1] != 1)
return false;
- /* Nobody else to wait for */
- return true;
+
+ bufHdr = GetLocalBufferDescriptor(-buffer - 1);
+ buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ /*
+ * Check that the AIO subsystem doesn't have a pin. Likely not
+ * possible today, but better safe than sorry.
+ */
+ if (BUF_STATE_GET_REFCOUNT(buf_state) == 1)
+ return true;
+
+ return false;
}
/* There should be exactly one local pin */
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 3c055f6ec8b..92c45611e0f 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -207,10 +207,19 @@ GetLocalVictimBuffer(void)
pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
trycounter = NLocBuffer;
}
+ else if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
+ {
+ /*
+ * This can be reached if the backend initiated AIO for this
+ * buffer and then errored out.
+ */
+ }
else
{
/* Found a usable buffer */
PinLocalBuffer(bufHdr, false);
+ /* the buf_state may be modified inside PinLocalBuffer */
+ buf_state = pg_atomic_read_u32(&bufHdr->state);
break;
}
}
@@ -491,6 +500,44 @@ MarkLocalBufferDirty(Buffer buffer)
pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
}
+static void
+InvalidateLocalBuffer(BufferDesc *bufHdr)
+{
+ Buffer buffer = BufferDescriptorGetBuffer(bufHdr);
+ int bufid = -buffer - 1;
+ uint32 buf_state;
+ LocalBufferLookupEnt *hresult;
+
+ buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ /*
+ * We need to test not just LocalRefCount[bufid] but also the BufferDesc
+ * itself, as the latter is used to represent a pin by the AIO subsystem.
+ * This can happen if AIO is initiated and then the query errors out.
+ */
+ if (LocalRefCount[bufid] != 0 ||
+ BUF_STATE_GET_REFCOUNT(buf_state) > 0)
+ elog(ERROR, "block %u of %s is still referenced (local %u)",
+ bufHdr->tag.blockNum,
+ relpathbackend(BufTagGetRelFileLocator(&bufHdr->tag),
+ MyProcNumber,
+ BufTagGetForkNum(&bufHdr->tag)),
+ LocalRefCount[bufid]);
+
+ /* Remove entry from hashtable */
+ hresult = (LocalBufferLookupEnt *)
+ hash_search(LocalBufHash, &bufHdr->tag, HASH_REMOVE, NULL);
+ if (!hresult) /* shouldn't happen */
+ elog(ERROR, "local buffer hash table corrupted");
+ /* Mark buffer invalid */
+ ClearBufferTag(&bufHdr->tag);
+
+ buf_state &= ~BUF_FLAG_MASK;
+ buf_state &= ~BUF_USAGECOUNT_MASK;
+ pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+
+}
+
/*
* DropRelationLocalBuffers
* This function removes from the buffer pool all the pages of the
@@ -511,7 +558,6 @@ DropRelationLocalBuffers(RelFileLocator rlocator, ForkNumber forkNum,
for (i = 0; i < NLocBuffer; i++)
{
BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
- LocalBufferLookupEnt *hresult;
uint32 buf_state;
buf_state = pg_atomic_read_u32(&bufHdr->state);
@@ -521,24 +567,7 @@ DropRelationLocalBuffers(RelFileLocator rlocator, ForkNumber forkNum,
BufTagGetForkNum(&bufHdr->tag) == forkNum &&
bufHdr->tag.blockNum >= firstDelBlock)
{
- if (LocalRefCount[i] != 0)
- elog(ERROR, "block %u of %s is still referenced (local %u)",
- bufHdr->tag.blockNum,
- relpathbackend(BufTagGetRelFileLocator(&bufHdr->tag),
- MyProcNumber,
- BufTagGetForkNum(&bufHdr->tag)),
- LocalRefCount[i]);
-
- /* Remove entry from hashtable */
- hresult = (LocalBufferLookupEnt *)
- hash_search(LocalBufHash, &bufHdr->tag, HASH_REMOVE, NULL);
- if (!hresult) /* shouldn't happen */
- elog(ERROR, "local buffer hash table corrupted");
- /* Mark buffer invalid */
- ClearBufferTag(&bufHdr->tag);
- buf_state &= ~BUF_FLAG_MASK;
- buf_state &= ~BUF_USAGECOUNT_MASK;
- pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+ InvalidateLocalBuffer(bufHdr);
}
}
}
@@ -558,7 +587,6 @@ DropRelationAllLocalBuffers(RelFileLocator rlocator)
for (i = 0; i < NLocBuffer; i++)
{
BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
- LocalBufferLookupEnt *hresult;
uint32 buf_state;
buf_state = pg_atomic_read_u32(&bufHdr->state);
@@ -566,23 +594,7 @@ DropRelationAllLocalBuffers(RelFileLocator rlocator)
if ((buf_state & BM_TAG_VALID) &&
BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
{
- if (LocalRefCount[i] != 0)
- elog(ERROR, "block %u of %s is still referenced (local %u)",
- bufHdr->tag.blockNum,
- relpathbackend(BufTagGetRelFileLocator(&bufHdr->tag),
- MyProcNumber,
- BufTagGetForkNum(&bufHdr->tag)),
- LocalRefCount[i]);
- /* Remove entry from hashtable */
- hresult = (LocalBufferLookupEnt *)
- hash_search(LocalBufHash, &bufHdr->tag, HASH_REMOVE, NULL);
- if (!hresult) /* shouldn't happen */
- elog(ERROR, "local buffer hash table corrupted");
- /* Mark buffer invalid */
- ClearBufferTag(&bufHdr->tag);
- buf_state &= ~BUF_FLAG_MASK;
- buf_state &= ~BUF_USAGECOUNT_MASK;
- pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+ InvalidateLocalBuffer(bufHdr);
}
}
}
@@ -680,12 +692,13 @@ PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
if (LocalRefCount[bufid] == 0)
{
NLocalPinnedBuffers++;
+ buf_state += BUF_REFCOUNT_ONE;
if (adjust_usagecount &&
BUF_STATE_GET_USAGECOUNT(buf_state) < BM_MAX_USAGE_COUNT)
{
buf_state += BUF_USAGECOUNT_ONE;
- pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
}
+ pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
}
LocalRefCount[bufid]++;
ResourceOwnerRememberBuffer(CurrentResourceOwner,
@@ -711,7 +724,17 @@ UnpinLocalBufferNoOwner(Buffer buffer)
Assert(NLocalPinnedBuffers > 0);
if (--LocalRefCount[buffid] == 0)
+ {
+ BufferDesc *buf_hdr = GetLocalBufferDescriptor(buffid);
+ uint32 buf_state;
+
NLocalPinnedBuffers--;
+
+ buf_state = pg_atomic_read_u32(&buf_hdr->state);
+ Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
+ buf_state -= BUF_REFCOUNT_ONE;
+ pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+ }
}
/*
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0019-bufmgr-Implement-AIO-read-support.patch (20.0K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/20-v2.4-0019-bufmgr-Implement-AIO-read-support.patch)
download | inline diff:
From a2f7a2d85772011077082cd06b7f8fa54c324035 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 16:08:58 -0500
Subject: [PATCH v2.4 19/29] bufmgr: Implement AIO read support
As of this commit there are no users of these AIO facilities, that'll come in
later commits.
Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/storage/aio.h | 4 +
src/include/storage/buf_internals.h | 6 +
src/include/storage/bufmgr.h | 8 +
src/backend/storage/aio/aio_callback.c | 5 +
src/backend/storage/buffer/buf_init.c | 3 +
src/backend/storage/buffer/bufmgr.c | 376 ++++++++++++++++++++++++-
src/backend/storage/buffer/localbuf.c | 77 +++++
7 files changed, 472 insertions(+), 7 deletions(-)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index 4c5fb7bcfce..6b34422607c 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -178,6 +178,10 @@ typedef enum PgAioHandleCallbackID
PGAIO_HCB_MD_READV,
PGAIO_HCB_MD_WRITEV,
+
+ PGAIO_HCB_SHARED_BUFFER_READV,
+
+ PGAIO_HCB_LOCAL_BUFFER_READV,
} PgAioHandleCallbackID;
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 1a65342177d..2a0c70c9998 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -17,6 +17,7 @@
#include "pgstat.h"
#include "port/atomics.h"
+#include "storage/aio_types.h"
#include "storage/buf.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
@@ -251,6 +252,8 @@ typedef struct BufferDesc
int wait_backend_pgprocno; /* backend of pin-count waiter */
int freeNext; /* link in freelist chain */
+
+ PgAioWaitRef io_wref;
LWLock content_lock; /* to lock access to buffer contents */
} BufferDesc;
@@ -464,4 +467,7 @@ extern void DropRelationLocalBuffers(RelFileLocator rlocator,
extern void DropRelationAllLocalBuffers(RelFileLocator rlocator);
extern void AtEOXact_LocalBuffers(bool isCommit);
+
+extern PgAioResult LocalBufferCompleteRead(int buf_off, Buffer buffer, int mode, bool failed);
+
#endif /* BUFMGR_INTERNALS_H */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 4a035f59a7d..efba4d88d7d 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -176,6 +176,12 @@ extern PGDLLIMPORT int NLocBuffer;
extern PGDLLIMPORT Block *LocalBufferBlockPointers;
extern PGDLLIMPORT int32 *LocalRefCount;
+
+struct PgAioHandleCallbacks;
+extern const struct PgAioHandleCallbacks aio_shared_buffer_readv_cb;
+extern const struct PgAioHandleCallbacks aio_local_buffer_readv_cb;
+
+
/* upper limit for effective_io_concurrency */
#define MAX_IO_CONCURRENCY 1000
@@ -193,6 +199,8 @@ extern PGDLLIMPORT int32 *LocalRefCount;
/*
* prototypes for functions in bufmgr.c
*/
+struct PgAioHandle;
+
extern PrefetchBufferResult PrefetchSharedBuffer(struct SMgrRelationData *smgr_reln,
ForkNumber forkNum,
BlockNumber blockNum);
diff --git a/src/backend/storage/aio/aio_callback.c b/src/backend/storage/aio/aio_callback.c
index adb8050eb58..6afdaaa434b 100644
--- a/src/backend/storage/aio/aio_callback.c
+++ b/src/backend/storage/aio/aio_callback.c
@@ -18,6 +18,7 @@
#include "miscadmin.h"
#include "storage/aio.h"
#include "storage/aio_internal.h"
+#include "storage/bufmgr.h"
#include "storage/md.h"
#include "utils/memutils.h"
@@ -42,6 +43,10 @@ static const PgAioHandleCallbacksEntry aio_handle_cbs[] = {
CALLBACK_ENTRY(PGAIO_HCB_MD_READV, aio_md_readv_cb),
CALLBACK_ENTRY(PGAIO_HCB_MD_WRITEV, aio_md_writev_cb),
+
+ CALLBACK_ENTRY(PGAIO_HCB_SHARED_BUFFER_READV, aio_shared_buffer_readv_cb),
+
+ CALLBACK_ENTRY(PGAIO_HCB_LOCAL_BUFFER_READV, aio_local_buffer_readv_cb),
#undef CALLBACK_ENTRY
};
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index ed1f8e03190..ed1dc488a42 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -14,6 +14,7 @@
*/
#include "postgres.h"
+#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
@@ -125,6 +126,8 @@ BufferManagerShmemInit(void)
buf->buf_id = i;
+ pgaio_wref_clear(&buf->io_wref);
+
/*
* Initially link all the buffers together as unused. Subsequent
* management of this list is done by freelist.c.
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index ec308557179..96b54f7abdf 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -48,6 +48,7 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "postmaster/bgwriter.h"
+#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
@@ -58,6 +59,7 @@
#include "storage/smgr.h"
#include "storage/standby.h"
#include "utils/memdebug.h"
+#include "utils/memutils.h"
#include "utils/ps_status.h"
#include "utils/rel.h"
#include "utils/resowner.h"
@@ -516,7 +518,8 @@ static int SyncOneBuffer(int buf_id, bool skip_recently_used,
static void WaitIO(BufferDesc *buf);
static bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
static void TerminateBufferIO(BufferDesc *buf, bool clear_dirty,
- uint32 set_flag_bits, bool forget_owner);
+ uint32 set_flag_bits, bool forget_owner,
+ bool syncio);
static void AbortBufferIO(Buffer buffer);
static void shared_buffer_write_error_callback(void *arg);
static void local_buffer_write_error_callback(void *arg);
@@ -1083,7 +1086,7 @@ ZeroAndLockBuffer(Buffer buffer, ReadBufferMode mode, bool already_valid)
else
{
/* Set BM_VALID, terminate IO, and wake up any waiters */
- TerminateBufferIO(bufHdr, false, BM_VALID, true);
+ TerminateBufferIO(bufHdr, false, BM_VALID, true, true);
}
}
else if (!isLocalBuf)
@@ -1619,7 +1622,7 @@ WaitReadBuffers(ReadBuffersOperation *operation)
else
{
/* Set BM_VALID, terminate IO, and wake up any waiters */
- TerminateBufferIO(bufHdr, false, BM_VALID, true);
+ TerminateBufferIO(bufHdr, false, BM_VALID, true, true);
}
/* Report I/Os as completing individually. */
@@ -2530,7 +2533,7 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
if (lock)
LWLockAcquire(BufferDescriptorGetContentLock(buf_hdr), LW_EXCLUSIVE);
- TerminateBufferIO(buf_hdr, false, BM_VALID, true);
+ TerminateBufferIO(buf_hdr, false, BM_VALID, true, true);
}
pgBufferUsage.shared_blks_written += extend_by;
@@ -3989,7 +3992,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
* Mark the buffer as clean (unless BM_JUST_DIRTIED has become set) and
* end the BM_IO_IN_PROGRESS state.
*/
- TerminateBufferIO(buf, true, 0, true);
+ TerminateBufferIO(buf, true, 0, true, true);
TRACE_POSTGRESQL_BUFFER_FLUSH_DONE(BufTagGetForkNum(&buf->tag),
buf->tag.blockNum,
@@ -5569,6 +5572,7 @@ WaitIO(BufferDesc *buf)
for (;;)
{
uint32 buf_state;
+ PgAioWaitRef iow;
/*
* It may not be necessary to acquire the spinlock to check the flag
@@ -5576,10 +5580,19 @@ WaitIO(BufferDesc *buf)
* play it safe.
*/
buf_state = LockBufHdr(buf);
+ iow = buf->io_wref;
UnlockBufHdr(buf, buf_state);
if (!(buf_state & BM_IO_IN_PROGRESS))
break;
+
+ if (pgaio_wref_valid(&iow))
+ {
+ pgaio_wref_wait(&iow);
+ ConditionVariablePrepareToSleep(cv);
+ continue;
+ }
+
ConditionVariableSleep(cv, WAIT_EVENT_BUFFER_IO);
}
ConditionVariableCancelSleep();
@@ -5668,7 +5681,7 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
*/
static void
TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
- bool forget_owner)
+ bool forget_owner, bool syncio)
{
uint32 buf_state;
@@ -5680,6 +5693,13 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
if (clear_dirty && !(buf_state & BM_JUST_DIRTIED))
buf_state &= ~(BM_DIRTY | BM_CHECKPOINT_NEEDED);
+ if (!syncio)
+ {
+ /* release ownership by the AIO subsystem */
+ buf_state -= BUF_REFCOUNT_ONE;
+ pgaio_wref_clear(&buf->io_wref);
+ }
+
buf_state |= set_flag_bits;
UnlockBufHdr(buf, buf_state);
@@ -5688,6 +5708,40 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
BufferDescriptorGetBuffer(buf));
ConditionVariableBroadcast(BufferDescriptorGetIOCV(buf));
+
+ /*
+ * If we just released a pin, need to do BM_PIN_COUNT_WAITER handling.
+ * Most of the time the current backend will hold another pin preventing
+ * that from happening, but that's e.g. not the case when completing an IO
+ * another backend started.
+ *
+ * AFIXME: Deduplicate with UnpinBufferNoOwner() or just replace
+ * BM_PIN_COUNT_WAITER with something saner.
+ */
+ /* Support LockBufferForCleanup() */
+ if (buf_state & BM_PIN_COUNT_WAITER)
+ {
+ /*
+ * Acquire the buffer header lock, re-check that there's a waiter.
+ * Another backend could have unpinned this buffer, and already woken
+ * up the waiter. There's no danger of the buffer being replaced
+ * after we unpinned it above, as it's pinned by the waiter.
+ */
+ buf_state = LockBufHdr(buf);
+
+ if ((buf_state & BM_PIN_COUNT_WAITER) &&
+ BUF_STATE_GET_REFCOUNT(buf_state) == 1)
+ {
+ /* we just released the last pin other than the waiter's */
+ int wait_backend_pgprocno = buf->wait_backend_pgprocno;
+
+ buf_state &= ~BM_PIN_COUNT_WAITER;
+ UnlockBufHdr(buf, buf_state);
+ ProcSendSignal(wait_backend_pgprocno);
+ }
+ else
+ UnlockBufHdr(buf, buf_state);
+ }
}
/*
@@ -5739,7 +5793,7 @@ AbortBufferIO(Buffer buffer)
}
}
- TerminateBufferIO(buf_hdr, false, BM_IO_ERROR, false);
+ TerminateBufferIO(buf_hdr, false, BM_IO_ERROR, false, true);
}
/*
@@ -6198,3 +6252,311 @@ EvictUnpinnedBuffer(Buffer buf)
return result;
}
+
+static PgAioResult
+SharedBufferCompleteRead(int buf_off, Buffer buffer, int mode, bool failed)
+{
+ BufferDesc *bufHdr = GetBufferDescriptor(buffer - 1);
+ BufferTag tag = bufHdr->tag;
+ char *bufdata = BufferGetBlock(buffer);
+ PgAioResult result;
+
+ Assert(BufferIsValid(buffer));
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ uint32 buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ Assert(buf_state & BM_TAG_VALID);
+ Assert(!(buf_state & BM_VALID));
+ Assert(buf_state & BM_IO_IN_PROGRESS);
+ Assert(!(buf_state & BM_DIRTY));
+ }
+#endif
+
+ result.status = ARS_OK;
+
+ /* check for garbage data */
+ if (!failed &&
+ !PageIsVerifiedExtended((Page) bufdata, tag.blockNum,
+ PIV_LOG_WARNING | PIV_REPORT_STAT))
+ {
+ RelFileLocator rlocator = BufTagGetRelFileLocator(&tag);
+
+ /* AFIXME: relpathperm allocates memory */
+ MemoryContextSwitchTo(ErrorContext);
+ if (mode == READ_BUFFERS_ZERO_ON_ERROR || zero_damaged_pages)
+ {
+ ereport(LOG,
+ (errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("invalid page in block %u of relation %s; zeroing out page",
+ tag.blockNum,
+ relpathperm(rlocator, tag.forkNum))));
+ memset(bufdata, 0, BLCKSZ);
+ }
+ else
+ {
+ /* mark buffer as having failed */
+ failed = true;
+
+ /* encode error for buffer_readv_report */
+ result.status = ARS_ERROR;
+ result.id = PGAIO_HCB_SHARED_BUFFER_READV;
+ result.error_data = buf_off;
+ }
+ }
+
+ /* Terminate I/O and set BM_VALID. */
+ TerminateBufferIO(bufHdr, false,
+ failed ? BM_IO_ERROR : BM_VALID,
+ false, false);
+
+ TRACE_POSTGRESQL_BUFFER_READ_DONE(tag.forkNum,
+ tag.blockNum,
+ tag.spcOid,
+ tag.dbOid,
+ tag.relNumber,
+ INVALID_PROC_NUMBER,
+ false);
+
+ return result;
+}
+
+/*
+ * Helper to prepare IO on shared buffers for execution, shared between reads
+ * and writes.
+ */
+static void
+shared_buffer_stage_common(PgAioHandle *ioh, bool is_write)
+{
+ uint64 *io_data;
+ uint8 handle_data_len;
+ PgAioWaitRef io_ref;
+ BufferTag first PG_USED_FOR_ASSERTS_ONLY = {0};
+
+ io_data = pgaio_io_get_handle_data(ioh, &handle_data_len);
+
+ pgaio_io_get_wref(ioh, &io_ref);
+
+ for (int i = 0; i < handle_data_len; i++)
+ {
+ Buffer buf = (Buffer) io_data[i];
+ BufferDesc *bufHdr;
+ uint32 buf_state;
+
+ bufHdr = GetBufferDescriptor(buf - 1);
+
+ if (i == 0)
+ first = bufHdr->tag;
+ else
+ {
+ Assert(bufHdr->tag.relNumber == first.relNumber);
+ Assert(bufHdr->tag.blockNum == first.blockNum + i);
+ }
+
+
+ buf_state = LockBufHdr(bufHdr);
+
+ Assert(buf_state & BM_TAG_VALID);
+ if (is_write)
+ {
+ Assert(buf_state & BM_VALID);
+ Assert(buf_state & BM_DIRTY);
+ }
+ else
+ Assert(!(buf_state & BM_VALID));
+
+ Assert(buf_state & BM_IO_IN_PROGRESS);
+ Assert(BUF_STATE_GET_REFCOUNT(buf_state) >= 1);
+
+ buf_state += BUF_REFCOUNT_ONE;
+ bufHdr->io_wref = io_ref;
+
+ UnlockBufHdr(bufHdr, buf_state);
+
+ if (is_write)
+ {
+ LWLock *content_lock;
+
+ content_lock = BufferDescriptorGetContentLock(bufHdr);
+
+ Assert(LWLockHeldByMe(content_lock));
+
+ /*
+ * Lock is now owned by AIO subsystem.
+ */
+ LWLockDisown(content_lock);
+ }
+
+ /*
+ * Stop tracking this buffer via the resowner - the AIO system now
+ * keeps track.
+ */
+ ResourceOwnerForgetBufferIO(CurrentResourceOwner, buf);
+ }
+}
+
+static void
+shared_buffer_readv_stage(PgAioHandle *ioh)
+{
+ shared_buffer_stage_common(ioh, false);
+}
+
+static void
+buffer_readv_report(PgAioResult result, const PgAioTargetData *target_data, int elevel)
+{
+ MemoryContext oldContext = CurrentMemoryContext;
+ ProcNumber errProc;
+
+ if (target_data->smgr.is_temp)
+ errProc = MyProcNumber;
+ else
+ errProc = INVALID_PROC_NUMBER;
+
+ /*
+ * AFIXME: need infrastructure to allow memory allocation for error
+ * reporting
+ */
+ oldContext = MemoryContextSwitchTo(ErrorContext);
+
+ ereport(elevel,
+ errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("invalid page in block %u of relation %s",
+ target_data->smgr.blockNum + result.error_data,
+ relpathbackend(target_data->smgr.rlocator, errProc, target_data->smgr.forkNum)
+ )
+ );
+ MemoryContextSwitchTo(oldContext);
+}
+
+/*
+ * Perform completion handling of a single AIO read. This read may cover
+ * multiple blocks / buffers.
+ *
+ * Shared between shared and local buffers, to reduce code duplication.
+ */
+static PgAioResult
+buffer_readv_complete_common(PgAioHandle *ioh, PgAioResult prior_result, bool is_temp)
+{
+ PgAioResult result = prior_result;
+ PgAioTargetData *td = pgaio_io_get_target_data(ioh);
+ int mode = td->smgr.mode;
+ uint64 *io_data;
+ uint8 handle_data_len;
+
+ if (is_temp)
+ {
+ Assert(td->smgr.is_temp);
+ Assert(pgaio_io_get_owner(ioh) == MyProcNumber);
+ }
+ else
+ Assert(!td->smgr.is_temp);
+
+ /*
+ * Iterate over all the buffers affected by this IO and call appropriate
+ * per-buffer completion function for each buffer.
+ */
+ io_data = pgaio_io_get_handle_data(ioh, &handle_data_len);
+ for (int buf_off = 0; buf_off < handle_data_len; buf_off++)
+ {
+ Buffer buf = io_data[buf_off];
+ PgAioResult buf_result;
+ bool failed;
+
+ /*
+ * If the entire failed on a lower-level, each buffer needs to be
+ * marked as failed. In case of a partial read, some buffers may be
+ * ok.
+ */
+ failed =
+ prior_result.status == ARS_ERROR
+ || prior_result.result <= buf_off;
+
+ if (is_temp)
+ buf_result = LocalBufferCompleteRead(buf_off, buf, mode, failed);
+ else
+ buf_result = SharedBufferCompleteRead(buf_off, buf, mode, failed);
+
+ /*
+ * If there wasn't any prior error and the IO for this page failed in
+ * some form, set the whole IO's to the page's result.
+ */
+ if (result.status != ARS_ERROR && buf_result.status != ARS_OK)
+ {
+ buffer_readv_report(result, td, LOG);
+ result = buf_result;
+ }
+ }
+
+ return result;
+}
+
+static PgAioResult
+shared_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ return buffer_readv_complete_common(ioh, prior_result, false);
+}
+
+/*
+ * Helper to stage IO on local buffers for execution, shared between reads
+ * and writes.
+ */
+static void
+local_buffer_readv_stage(PgAioHandle *ioh)
+{
+ uint64 *io_data;
+ uint8 handle_data_len;
+ PgAioWaitRef io_wref;
+
+ io_data = pgaio_io_get_handle_data(ioh, &handle_data_len);
+
+ pgaio_io_get_wref(ioh, &io_wref);
+
+ for (int i = 0; i < handle_data_len; i++)
+ {
+ Buffer buf = (Buffer) io_data[i];
+ BufferDesc *bufHdr;
+ uint32 buf_state;
+
+ bufHdr = GetLocalBufferDescriptor(-buf - 1);
+
+ buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ bufHdr->io_wref = io_wref;
+
+ /*
+ * Track pin by AIO subsystem in BufferDesc, not in LocalRefCount as
+ * one might initially think. This is necessary to handle this backend
+ * erroring out while AIO is still in progress.
+ */
+ buf_state += BUF_REFCOUNT_ONE;
+
+ pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+ }
+}
+
+static PgAioResult
+local_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ return buffer_readv_complete_common(ioh, prior_result, true);
+
+}
+
+
+const struct PgAioHandleCallbacks aio_shared_buffer_readv_cb = {
+ .stage = shared_buffer_readv_stage,
+ .complete_shared = shared_buffer_readv_complete,
+ .report = buffer_readv_report,
+};
+const struct PgAioHandleCallbacks aio_local_buffer_readv_cb = {
+ .stage = local_buffer_readv_stage,
+
+ /*
+ * Note that this, in contrast to the shared_buffers case, uses
+ * complete_local, as only the issuing backend has access to the required
+ * datastructures. This is important in case the IO completion may be
+ * consumed incidentally by another backend.
+ */
+ .complete_local = local_buffer_readv_complete,
+ .report = buffer_readv_report,
+};
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 92c45611e0f..d997d8e8632 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -17,7 +17,9 @@
#include "access/parallel.h"
#include "executor/instrument.h"
+#include "pg_trace.h"
#include "pgstat.h"
+#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
@@ -649,6 +651,8 @@ InitLocalBuffers(void)
*/
buf->buf_id = -i - 2;
+ pgaio_wref_clear(&buf->io_wref);
+
/*
* Intentionally do not initialize the buffer's atomic variable
* (besides zeroing the underlying memory above). That way we get
@@ -876,3 +880,76 @@ AtProcExit_LocalBuffers(void)
*/
CheckForLocalBufferLeaks();
}
+
+PgAioResult
+LocalBufferCompleteRead(int buf_off, Buffer buffer, int mode, bool failed)
+{
+ BufferDesc *bufHdr = GetLocalBufferDescriptor(-buffer - 1);
+ BufferTag tag = bufHdr->tag;
+ char *bufdata = BufferGetBlock(buffer);
+ PgAioResult result;
+
+ Assert(BufferIsValid(buffer));
+
+ result.status = ARS_OK;
+
+ /* check for garbage data */
+ if (!failed &&
+ !PageIsVerifiedExtended((Page) bufdata, tag.blockNum,
+ PIV_LOG_WARNING | PIV_REPORT_STAT))
+ {
+ RelFileLocator rlocator = BufTagGetRelFileLocator(&tag);
+ BlockNumber forkNum = tag.forkNum;
+
+ MemoryContextSwitchTo(ErrorContext);
+
+ if (mode == READ_BUFFERS_ZERO_ON_ERROR || zero_damaged_pages)
+ {
+
+ ereport(LOG,
+ (errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("invalid page in block %u of relation %s; zeroing out page",
+ tag.blockNum,
+ relpathbackend(rlocator, MyProcNumber, forkNum))));
+ memset(bufdata, 0, BLCKSZ);
+ }
+ else
+ {
+ /* mark buffer as having failed */
+ failed = true;
+
+ /* encode error for buffer_readv_report */
+ result.status = ARS_ERROR;
+ result.id = PGAIO_HCB_LOCAL_BUFFER_READV;
+ result.error_data = buf_off;
+ }
+ }
+
+ /* Terminate I/O and set BM_VALID. */
+ pgaio_wref_clear(&bufHdr->io_wref);
+
+ {
+ uint32 buf_state;
+
+ buf_state = pg_atomic_read_u32(&bufHdr->state);
+ buf_state |= BM_VALID;
+
+ /*
+ * Release pin held by IO subsystem, see also
+ * local_buffer_readv_prepare().
+ */
+ buf_state -= BUF_REFCOUNT_ONE;
+ pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+ }
+
+
+ TRACE_POSTGRESQL_BUFFER_READ_DONE(tag.forkNum,
+ tag.blockNum,
+ tag.spcOid,
+ tag.dbOid,
+ tag.relNumber,
+ INVALID_PROC_NUMBER,
+ false);
+
+ return result;
+}
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0020-bufmgr-Use-aio-for-StartReadBuffers.patch (18.8K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/21-v2.4-0020-bufmgr-Use-aio-for-StartReadBuffers.patch)
download | inline diff:
From 517e55c26298decd26eab0dec4da220aa084ad35 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 12 Feb 2025 14:19:20 -0500
Subject: [PATCH v2.4 20/29] bufmgr: Use aio for StartReadBuffers()
Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/storage/bufmgr.h | 7 +
src/backend/storage/buffer/bufmgr.c | 411 +++++++++++++++++++++-------
2 files changed, 317 insertions(+), 101 deletions(-)
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index efba4d88d7d..dc8fe197d6f 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -15,6 +15,7 @@
#define BUFMGR_H
#include "port/pg_iovec.h"
+#include "storage/aio_types.h"
#include "storage/block.h"
#include "storage/buf.h"
#include "storage/bufpage.h"
@@ -111,6 +112,9 @@ typedef struct BufferManagerRelation
#define READ_BUFFERS_ZERO_ON_ERROR (1 << 0)
/* Call smgrprefetch() if I/O necessary. */
#define READ_BUFFERS_ISSUE_ADVICE (1 << 1)
+/* IO will immediately be waited for */
+#define READ_BUFFERS_SYNCHRONOUSLY (1 << 2)
+
struct ReadBuffersOperation
{
@@ -130,6 +134,9 @@ struct ReadBuffersOperation
BlockNumber blocknum;
int flags;
int16 nblocks;
+
+ PgAioWaitRef io_wref;
+ PgAioReturn io_return;
};
typedef struct ReadBuffersOperation ReadBuffersOperation;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 96b54f7abdf..ee9a9f70167 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -529,6 +529,8 @@ static inline BufferDesc *BufferAlloc(SMgrRelation smgr,
BlockNumber blockNum,
BufferAccessStrategy strategy,
bool *foundPtr, IOContext io_context);
+static bool AsyncReadBuffers(ReadBuffersOperation *operation,
+ int *nblocks);
static Buffer GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context);
static void FlushBuffer(BufferDesc *buf, SMgrRelation reln,
IOObject io_object, IOContext io_context);
@@ -1237,10 +1239,9 @@ ReadBuffer_common(Relation rel, SMgrRelation smgr, char smgr_persistence,
return buffer;
}
+ flags = READ_BUFFERS_SYNCHRONOUSLY;
if (mode == RBM_ZERO_ON_ERROR)
- flags = READ_BUFFERS_ZERO_ON_ERROR;
- else
- flags = 0;
+ flags |= READ_BUFFERS_ZERO_ON_ERROR;
operation.smgr = smgr;
operation.rel = rel;
operation.persistence = persistence;
@@ -1268,6 +1269,7 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
Assert(*nblocks > 0);
Assert(*nblocks <= MAX_IO_COMBINE_LIMIT);
+ Assert(*nblocks == 1 || allow_forwarding);
for (int i = 0; i < actual_nblocks; ++i)
{
@@ -1307,6 +1309,11 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
else
bufHdr = GetBufferDescriptor(buffers[i] - 1);
found = pg_atomic_read_u32(&bufHdr->state) & BM_VALID;
+
+ ereport(DEBUG3,
+ errmsg("found forwarded buffer %d",
+ buffers[i]),
+ errhidestmt(true), errhidecontext(true));
}
else
{
@@ -1372,25 +1379,59 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
operation->blocknum = blockNum;
operation->flags = flags;
operation->nblocks = actual_nblocks;
+ pgaio_wref_clear(&operation->io_wref);
- if (flags & READ_BUFFERS_ISSUE_ADVICE)
+ /*
+ * When using AIO, start the IO in the background. If not, issue prefetch
+ * requests if desired by the caller.
+ *
+ * The reason we have a dedicated path for IOMETHOD_SYNC here is to
+ * de-risk the introduction of AIO somewhat. It's a large architectural
+ * change, with lots of chances for unanticipated performance effects.
+ *
+ * Use of IOMETHOD_SYNC already leads to not actually performing IO
+ * asynchronously, but without the check here we'd execute IO earlier than
+ * we used to. Eventually this IOMETHOD_SYNC specific path should go away.
+ */
+ if (io_method != IOMETHOD_SYNC)
{
/*
- * In theory we should only do this if PinBufferForBlock() had to
- * allocate new buffers above. That way, if two calls to
- * StartReadBuffers() were made for the same blocks before
- * WaitReadBuffers(), only the first would issue the advice. That'd be
- * a better simulation of true asynchronous I/O, which would only
- * start the I/O once, but isn't done here for simplicity.
+ * Try to start IO asynchronously. It's possible that no IO needs to
+ * be started, if another backend already performed the IO.
+ *
+ * Note that if an IO is started, it might not cover the entire
+ * requested range, e.g. because an intermediary block has been read
+ * in by another backend. In that case any "trailing" buffers we
+ * already pinned above will be "forwarded" by read_stream.c to the
+ * next call to StartReadBuffers(). This is signalled to the caller by
+ * decrementing *nblocks.
*/
- smgrprefetch(operation->smgr,
- operation->forknum,
- blockNum,
- actual_nblocks);
+ return AsyncReadBuffers(operation, nblocks);
}
+ else
+ {
+ operation->flags |= READ_BUFFERS_SYNCHRONOUSLY;
- /* Indicate that WaitReadBuffers() should be called. */
- return true;
+ if (flags & READ_BUFFERS_ISSUE_ADVICE)
+ {
+ /*
+ * In theory we should only do this if PinBufferForBlock() had to
+ * allocate new buffers above. That way, if two calls to
+ * StartReadBuffers() were made for the same blocks before
+ * WaitReadBuffers(), only the first would issue the advice.
+ * That'd be a better simulation of true asynchronous I/O, which
+ * would only start the I/O once, but isn't done here for
+ * simplicity.
+ */
+ smgrprefetch(operation->smgr,
+ operation->forknum,
+ blockNum,
+ actual_nblocks);
+ }
+
+ /* Indicate that WaitReadBuffers() should be called. */
+ return true;
+ }
}
/*
@@ -1458,12 +1499,31 @@ StartReadBuffer(ReadBuffersOperation *operation,
}
static inline bool
-WaitReadBuffersCanStartIO(Buffer buffer, bool nowait)
+ReadBuffersCanStartIO(Buffer buffer, bool nowait)
{
if (BufferIsLocal(buffer))
{
BufferDesc *bufHdr = GetLocalBufferDescriptor(-buffer - 1);
+ /*
+ * The buffer could have IO in progress by another scan. Right now
+ * localbuf.c doesn't use IO_IN_PROGRESS, which is why we need this
+ * hack.
+ *
+ * TODO: localbuf.c should use IO_IN_PROGRESS / have an equivalent of
+ * StartBufferIO().
+ */
+ if (pgaio_wref_valid(&bufHdr->io_wref))
+ {
+ PgAioWaitRef iow = bufHdr->io_wref;
+
+ ereport(DEBUG3,
+ errmsg("waiting for temp buffer IO in CSIO"),
+ errhidestmt(true), errhidecontext(true));
+ pgaio_wref_wait(&iow);
+ return false;
+ }
+
return (pg_atomic_read_u32(&bufHdr->state) & BM_VALID) == 0;
}
else
@@ -1473,28 +1533,163 @@ WaitReadBuffersCanStartIO(Buffer buffer, bool nowait)
void
WaitReadBuffers(ReadBuffersOperation *operation)
{
- Buffer *buffers;
+ IOContext io_context;
+ IOObject io_object;
int nblocks;
- BlockNumber blocknum;
- ForkNumber forknum;
- IOContext io_context;
- IOObject io_object;
- char persistence;
+ PgAioReturn *aio_ret;
+
+ /*
+ * If we get here without an IO operation having been issued, io_method ==
+ * IOMETHOD_SYNC path must have been used. In that case, we start - as we
+ * used to before - the IO now, just before waiting.
+ *
+ * This path is expected to eventually go away.
+ */
+ if (!pgaio_wref_valid(&operation->io_wref))
+ {
+ Assert(io_method == IOMETHOD_SYNC);
+
+ while (true)
+ {
+ nblocks = operation->nblocks;
+
+ if (!AsyncReadBuffers(operation, &nblocks))
+ {
+ /* all blocks were already read in concurrently */
+ Assert(nblocks == operation->nblocks);
+ return;
+ }
+
+ Assert(nblocks > 0 && nblocks <= operation->nblocks);
+
+ if (nblocks == operation->nblocks)
+ {
+ /* will wait below as if this had been normal AIO */
+ break;
+ }
+
+ /*
+ * It's unlikely, but possible, that AsyncReadBuffers() wasn't
+ * able to initiate IO for all the relevant buffers. In that case
+ * we need to wait for the prior IO before issuing more IO.
+ */
+ WaitReadBuffers(operation);
+ }
+ }
+
+ if (operation->persistence == RELPERSISTENCE_TEMP)
+ {
+ io_context = IOCONTEXT_NORMAL;
+ io_object = IOOBJECT_TEMP_RELATION;
+ }
+ else
+ {
+ io_context = IOContextForStrategy(operation->strategy);
+ io_object = IOOBJECT_RELATION;
+ }
+
+restart:
/* Find the range of the physical read we need to perform. */
nblocks = operation->nblocks;
- buffers = &operation->buffers[0];
- blocknum = operation->blocknum;
- forknum = operation->forknum;
- persistence = operation->persistence;
-
Assert(nblocks > 0);
Assert(nblocks <= MAX_IO_COMBINE_LIMIT);
+ aio_ret = &operation->io_return;
+
+ /*
+ * For IO timing we just count the time spent waiting for the IO.
+ *
+ * XXX: We probably should track the IO operation, rather than its time,
+ * separately, when initiating the IO. But right now that's not quite
+ * allowed by the interface.
+ */
+
+ /*
+ * Tracking a wait even if we don't actually need to wait
+ *
+ * a) is not cheap
+ *
+ * b) reports some time as waiting, even if we never waited.
+ */
+ if (aio_ret->result.status == ARS_UNKNOWN &&
+ !pgaio_wref_check_done(&operation->io_wref))
+ {
+ instr_time io_start = pgstat_prepare_io_time(track_io_timing);
+
+ pgaio_wref_wait(&operation->io_wref);
+
+ /*
+ * The IO operation itself was already counted earlier, in
+ * AsyncReadBuffers().
+ */
+ pgstat_count_io_op_time(io_object, io_context, IOOP_READ,
+ io_start, 0, 0);
+ }
+ else
+ {
+ Assert(pgaio_wref_check_done(&operation->io_wref));
+ }
+
+ if (aio_ret->result.status == ARS_PARTIAL)
+ {
+ /*
+ * We'll retry below, so we just emit a debug message the server log
+ * (or not even that in prod scenarios).
+ */
+ pgaio_result_report(aio_ret->result, &aio_ret->target_data, DEBUG1);
+
+ /*
+ * Try to perform the rest of the IO. Buffers for which IO has
+ * completed successfully will be discovered as such and not retried.
+ */
+ nblocks = operation->nblocks;
+
+ elog(DEBUG3, "retrying IO after partial failure");
+ CHECK_FOR_INTERRUPTS();
+ AsyncReadBuffers(operation, &nblocks);
+ goto restart;
+ }
+ else if (aio_ret->result.status != ARS_OK)
+ pgaio_result_report(aio_ret->result, &aio_ret->target_data, ERROR);
+
+ if (VacuumCostActive)
+ VacuumCostBalance += VacuumCostPageMiss * nblocks;
+
+ /* NB: READ_DONE tracepoint is executed in IO completion callback */
+}
+
+/*
+ * Initiate IO for the ReadBuffersOperation. If IO is only initiated for a
+ * subset of the blocks, *nblocks is updated to reflect that.
+ *
+ * Returns true if IO was initiated, false if no IO was necessary.
+ */
+static bool
+AsyncReadBuffers(ReadBuffersOperation *operation,
+ int *nblocks)
+{
+ int io_buffers_len = 0;
+ Buffer *buffers = &operation->buffers[0];
+ int flags = operation->flags;
+ BlockNumber blocknum = operation->blocknum;
+ ForkNumber forknum = operation->forknum;
+ bool did_start_io = false;
+ PgAioHandle *ioh = NULL;
+ uint32 ioh_flags = 0;
+ IOContext io_context;
+ IOObject io_object;
+ char persistence;
+
+ persistence = operation->rel
+ ? operation->rel->rd_rel->relpersistence
+ : RELPERSISTENCE_PERMANENT;
+
if (persistence == RELPERSISTENCE_TEMP)
{
io_context = IOCONTEXT_NORMAL;
io_object = IOOBJECT_TEMP_RELATION;
+ ioh_flags |= PGAIO_HF_REFERENCES_LOCAL;
}
else
{
@@ -1502,6 +1697,14 @@ WaitReadBuffers(ReadBuffersOperation *operation)
io_object = IOOBJECT_RELATION;
}
+ /*
+ * When this IO is executed synchronously, either because the caller will
+ * immediately block waiting for the IO or because IOMETHOD_SYNC is used,
+ * the AIO subsystem needs to know.
+ */
+ if (flags & READ_BUFFERS_SYNCHRONOUSLY)
+ ioh_flags |= PGAIO_HF_SYNCHRONOUS;
+
/*
* We count all these blocks as read by this backend. This is traditional
* behavior, but might turn out to be not true if we find that someone
@@ -1511,25 +1714,53 @@ WaitReadBuffers(ReadBuffersOperation *operation)
* but another backend completed the read".
*/
if (persistence == RELPERSISTENCE_TEMP)
- pgBufferUsage.local_blks_read += nblocks;
+ pgBufferUsage.local_blks_read += *nblocks;
else
- pgBufferUsage.shared_blks_read += nblocks;
+ pgBufferUsage.shared_blks_read += *nblocks;
- for (int i = 0; i < nblocks; ++i)
+ pgaio_wref_clear(&operation->io_wref);
+
+ /*
+ * Loop until we have started one IO or we discover that all buffers are
+ * already valid.
+ */
+ for (int i = 0; i < *nblocks; ++i)
{
- int io_buffers_len;
Buffer io_buffers[MAX_IO_COMBINE_LIMIT];
void *io_pages[MAX_IO_COMBINE_LIMIT];
- instr_time io_start;
BlockNumber io_first_block;
/*
- * Skip this block if someone else has already completed it. If an
- * I/O is already in progress in another backend, this will wait for
- * the outcome: either done, or something went wrong and we will
- * retry.
+ * Get IO before ReadBuffersCanStartIO, as pgaio_io_acquire() might
+ * block, which we don't want after setting IO_IN_PROGRESS.
+ *
+ * XXX: Should we attribute the time spent in here to the IO? If there
+ * already are a lot of IO operations in progress, getting an IO
+ * handle will block waiting for some other IO operation to finish.
+ *
+ * In most cases it'll be free to get the IO, so a timer would be
+ * overhead. Perhaps we should use pgaio_io_acquire_nb() and only
+ * account IO time when pgaio_io_acquire_nb() returned false?
*/
- if (!WaitReadBuffersCanStartIO(buffers[i], false))
+ if (likely(!ioh))
+ ioh = pgaio_io_acquire(CurrentResourceOwner,
+ &operation->io_return);
+
+ /*
+ * Skip this block if someone else has already completed it.
+ *
+ * If an I/O is already in progress in another backend, this will wait
+ * for the outcome: either done, or something went wrong and we will
+ * retry. But don't wait if we have staged, but haven't issued,
+ * another IO.
+ *
+ * It's safe to start IO while we have unsubmitted IO, but it'd be
+ * better to first submit it. But right now the boolean return value
+ * from ReadBuffersCanStartIO()/StartBufferIO() doesn't allow to
+ * distinguish between nowait=true trigger failure and the buffer
+ * already being valid.
+ */
+ if (!ReadBuffersCanStartIO(buffers[i], false))
{
/*
* Report this as a 'hit' for this backend, even though it must
@@ -1541,6 +1772,11 @@ WaitReadBuffers(ReadBuffersOperation *operation)
operation->smgr->smgr_rlocator.locator.relNumber,
operation->smgr->smgr_rlocator.backend,
true);
+
+ ereport(DEBUG3,
+ errmsg("can't start io for first buffer %u: %s",
+ buffers[i], DebugPrintBufferRefcount(buffers[i])),
+ errhidestmt(true), errhidecontext(true));
continue;
}
@@ -1550,6 +1786,11 @@ WaitReadBuffers(ReadBuffersOperation *operation)
io_first_block = blocknum + i;
io_buffers_len = 1;
+ ereport(DEBUG5,
+ errmsg("first prepped for io: %s, offset %d",
+ DebugPrintBufferRefcount(io_buffers[0]), i),
+ errhidestmt(true), errhidecontext(true));
+
/*
* How many neighboring-on-disk blocks can we scatter-read into other
* buffers at the same time? In this case we don't wait if we see an
@@ -1557,86 +1798,54 @@ WaitReadBuffers(ReadBuffersOperation *operation)
* head block, so we should get on with that I/O as soon as possible.
* We'll come back to this block again, above.
*/
- while ((i + 1) < nblocks &&
- WaitReadBuffersCanStartIO(buffers[i + 1], true))
+ while ((i + 1) < *nblocks &&
+ ReadBuffersCanStartIO(buffers[i + 1], true))
{
/* Must be consecutive block numbers. */
Assert(BufferGetBlockNumber(buffers[i + 1]) ==
BufferGetBlockNumber(buffers[i]) + 1);
+ ereport(DEBUG5,
+ errmsg("seq prepped for io: %s, offset %d",
+ DebugPrintBufferRefcount(buffers[i + 1]),
+ i + 1),
+ errhidestmt(true), errhidecontext(true));
+
io_buffers[io_buffers_len] = buffers[++i];
io_pages[io_buffers_len++] = BufferGetBlock(buffers[i]);
}
- io_start = pgstat_prepare_io_time(track_io_timing);
- smgrreadv(operation->smgr, forknum, io_first_block, io_pages, io_buffers_len);
- pgstat_count_io_op_time(io_object, io_context, IOOP_READ, io_start,
- 1, io_buffers_len * BLCKSZ);
+ pgaio_io_get_wref(ioh, &operation->io_wref);
- /* Verify each block we read, and terminate the I/O. */
- for (int j = 0; j < io_buffers_len; ++j)
- {
- BufferDesc *bufHdr;
- Block bufBlock;
+ pgaio_io_set_handle_data_32(ioh, (uint32 *) io_buffers, io_buffers_len);
- if (persistence == RELPERSISTENCE_TEMP)
- {
- bufHdr = GetLocalBufferDescriptor(-io_buffers[j] - 1);
- bufBlock = LocalBufHdrGetBlock(bufHdr);
- }
- else
- {
- bufHdr = GetBufferDescriptor(io_buffers[j] - 1);
- bufBlock = BufHdrGetBlock(bufHdr);
- }
+ if (persistence == RELPERSISTENCE_TEMP)
+ pgaio_io_register_callbacks(ioh, PGAIO_HCB_LOCAL_BUFFER_READV);
+ else
+ pgaio_io_register_callbacks(ioh, PGAIO_HCB_SHARED_BUFFER_READV);
- /* check for garbage data */
- if (!PageIsVerifiedExtended((Page) bufBlock, io_first_block + j,
- PIV_LOG_WARNING | PIV_REPORT_STAT))
- {
- if ((operation->flags & READ_BUFFERS_ZERO_ON_ERROR) || zero_damaged_pages)
- {
- ereport(WARNING,
- (errcode(ERRCODE_DATA_CORRUPTED),
- errmsg("invalid page in block %u of relation %s; zeroing out page",
- io_first_block + j,
- relpath(operation->smgr->smgr_rlocator, forknum))));
- memset(bufBlock, 0, BLCKSZ);
- }
- else
- ereport(ERROR,
- (errcode(ERRCODE_DATA_CORRUPTED),
- errmsg("invalid page in block %u of relation %s",
- io_first_block + j,
- relpath(operation->smgr->smgr_rlocator, forknum))));
- }
+ pgaio_io_set_flag(ioh, ioh_flags);
- /* Terminate I/O and set BM_VALID. */
- if (persistence == RELPERSISTENCE_TEMP)
- {
- uint32 buf_state = pg_atomic_read_u32(&bufHdr->state);
+ did_start_io = true;
+ smgrstartreadv(ioh, operation->smgr, forknum, io_first_block,
+ io_pages, io_buffers_len);
+ ioh = NULL;
- buf_state |= BM_VALID;
- pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
- }
- else
- {
- /* Set BM_VALID, terminate IO, and wake up any waiters */
- TerminateBufferIO(bufHdr, false, BM_VALID, true, true);
- }
+ /* not obvious what we'd use for time */
+ pgstat_count_io_op(io_object, io_context, IOOP_READ,
+ 1, io_buffers_len * BLCKSZ);
- /* Report I/Os as completing individually. */
- TRACE_POSTGRESQL_BUFFER_READ_DONE(forknum, io_first_block + j,
- operation->smgr->smgr_rlocator.locator.spcOid,
- operation->smgr->smgr_rlocator.locator.dbOid,
- operation->smgr->smgr_rlocator.locator.relNumber,
- operation->smgr->smgr_rlocator.backend,
- false);
- }
+ *nblocks = io_buffers_len;
+ break;
+ }
- if (VacuumCostActive)
- VacuumCostBalance += VacuumCostPageMiss * io_buffers_len;
+ if (ioh)
+ {
+ pgaio_io_release(ioh);
+ ioh = NULL;
}
+
+ return did_start_io;
}
/*
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0021-WIP-aio-read_stream.c-adjustments-for-real-AIO.patch (3.7K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/22-v2.4-0021-WIP-aio-read_stream.c-adjustments-for-real-AIO.patch)
download | inline diff:
From e5d682eedac10c0b703a81ffed501152e46db1a3 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 11 Feb 2025 18:21:47 -0500
Subject: [PATCH v2.4 21/29] WIP: aio: read_stream.c adjustments for real AIO
Comments need to be fixed.
The batching logic probably needs to be adjusted.
---
src/backend/storage/aio/read_stream.c | 31 +++++++++++++++++++++++++--
1 file changed, 29 insertions(+), 2 deletions(-)
diff --git a/src/backend/storage/aio/read_stream.c b/src/backend/storage/aio/read_stream.c
index 32e5def29f8..7d1d308f13f 100644
--- a/src/backend/storage/aio/read_stream.c
+++ b/src/backend/storage/aio/read_stream.c
@@ -32,10 +32,15 @@
* calls. Looking further ahead would pin many buffers and perform
* speculative work for no benefit.
*
+ * FIXME: This only applies to io_method == sync, otherwise this path is not
+ * used.
+ *
* C) I/O is necessary, it appears to be random, and this system supports
* read-ahead advice. We'll look further ahead in order to reach the
* configured level of I/O concurrency.
*
+ * FIXME: restriction to random only applies to io_method == sync
+ *
* The distance increases rapidly and decays slowly, so that it moves towards
* those levels as different I/O patterns are discovered. For example, a
* sequential scan of fully cached data doesn't bother looking ahead, but a
@@ -90,6 +95,7 @@
#include "postgres.h"
#include "miscadmin.h"
+#include "storage/aio.h"
#include "storage/fd.h"
#include "storage/smgr.h"
#include "storage/read_stream.h"
@@ -116,6 +122,7 @@ struct ReadStream
int16 pinned_buffers;
int16 distance;
int16 initialized_buffers;
+ bool sync_mode;
bool advice_enabled;
bool temporary;
@@ -458,6 +465,19 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
{
int16 buffer_limit;
+ if (stream->distance > (io_combine_limit * 8))
+ {
+ if (stream->pinned_buffers + stream->pending_read_nblocks > ((stream->distance * 3) / 4))
+ {
+ return;
+ }
+ }
+
+ /*
+ * Try to amortize cost of submitting IOs over multiple IOs.
+ */
+ pgaio_enter_batchmode();
+
/*
* Check how many pins we could acquire now. We do this here rather than
* pushing it down into read_stream_start_pending_read(), because it
@@ -524,6 +544,7 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
{
/* And we've hit a limit. Rewind, and stop here. */
read_stream_unget_block(stream, blocknum);
+ pgaio_exit_batchmode();
return;
}
}
@@ -549,6 +570,8 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
stream->distance == 0) &&
stream->ios_in_progress < stream->max_ios)
read_stream_start_pending_read(stream, buffer_limit, suppress_advice);
+
+ pgaio_exit_batchmode();
}
/*
@@ -668,6 +691,8 @@ read_stream_begin_impl(int flags,
stream->per_buffer_data = (void *)
MAXALIGN(&stream->ios[Max(1, max_ios)]);
+ stream->sync_mode = io_method == IOMETHOD_SYNC;
+
#ifdef USE_PREFETCH
/*
@@ -676,7 +701,8 @@ read_stream_begin_impl(int flags,
* (overriding our detection heuristics), and max_ios hasn't been set to
* zero.
*/
- if ((io_direct_flags & IO_DIRECT_DATA) == 0 &&
+ if (stream->sync_mode &&
+ (io_direct_flags & IO_DIRECT_DATA) == 0 &&
(flags & READ_STREAM_SEQUENTIAL) == 0 &&
max_ios > 0)
stream->advice_enabled = true;
@@ -919,7 +945,8 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
if (++stream->oldest_io_index == stream->max_ios)
stream->oldest_io_index = 0;
- if (stream->ios[io_index].op.flags & READ_BUFFERS_ISSUE_ADVICE)
+ if (!stream->sync_mode ||
+ stream->ios[io_index].op.flags & READ_BUFFERS_ISSUE_ADVICE)
{
/* Distance ramps up fast (behavior C). */
distance = stream->distance * 2;
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0022-aio-Add-test_aio-module.patch (31.9K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/23-v2.4-0022-aio-Add-test_aio-module.patch)
download | inline diff:
From 51185432a85dea16d55a49537d4bcef813cf0b40 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 13:44:54 -0500
Subject: [PATCH v2.4 22/29] aio: Add test_aio module
---
src/include/storage/buf_internals.h | 4 +
src/backend/storage/buffer/bufmgr.c | 3 +-
src/test/modules/Makefile | 1 +
src/test/modules/meson.build | 1 +
src/test/modules/test_aio/.gitignore | 2 +
src/test/modules/test_aio/Makefile | 27 +
src/test/modules/test_aio/meson.build | 37 ++
src/test/modules/test_aio/t/001_aio.pl | 401 +++++++++++++++
src/test/modules/test_aio/test_aio--1.0.sql | 84 ++++
src/test/modules/test_aio/test_aio.c | 518 ++++++++++++++++++++
src/test/modules/test_aio/test_aio.control | 3 +
11 files changed, 1079 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_aio/.gitignore
create mode 100644 src/test/modules/test_aio/Makefile
create mode 100644 src/test/modules/test_aio/meson.build
create mode 100644 src/test/modules/test_aio/t/001_aio.pl
create mode 100644 src/test/modules/test_aio/test_aio--1.0.sql
create mode 100644 src/test/modules/test_aio/test_aio.c
create mode 100644 src/test/modules/test_aio/test_aio.control
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 2a0c70c9998..396642415bc 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -421,6 +421,10 @@ extern void IssuePendingWritebacks(WritebackContext *wb_context, IOContext io_co
extern void ScheduleBufferTagForWriteback(WritebackContext *wb_context,
IOContext io_context, BufferTag *tag);
+/* solely to make it easier to write tests */
+extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
+
+
/* freelist.c */
extern IOContext IOContextForStrategy(BufferAccessStrategy strategy);
extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index ee9a9f70167..b641bc3982b 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -516,7 +516,6 @@ static uint32 WaitBufHdrUnlocked(BufferDesc *buf);
static int SyncOneBuffer(int buf_id, bool skip_recently_used,
WritebackContext *wb_context);
static void WaitIO(BufferDesc *buf);
-static bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
static void TerminateBufferIO(BufferDesc *buf, bool clear_dirty,
uint32 set_flag_bits, bool forget_owner,
bool syncio);
@@ -5831,7 +5830,7 @@ WaitIO(BufferDesc *buf)
* find out if they can perform the I/O as part of a larger operation, without
* waiting for the answer or distinguishing the reasons why not.
*/
-static bool
+bool
StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
{
uint32 buf_state;
diff --git a/src/test/modules/Makefile b/src/test/modules/Makefile
index 89e78b7d114..0ae2d4b6669 100644
--- a/src/test/modules/Makefile
+++ b/src/test/modules/Makefile
@@ -13,6 +13,7 @@ SUBDIRS = \
libpq_pipeline \
plsample \
spgist_name_ops \
+ test_aio \
test_bloomfilter \
test_copy_callbacks \
test_custom_rmgrs \
diff --git a/src/test/modules/meson.build b/src/test/modules/meson.build
index a57077b682e..94a9ebfdf60 100644
--- a/src/test/modules/meson.build
+++ b/src/test/modules/meson.build
@@ -1,5 +1,6 @@
# Copyright (c) 2022-2025, PostgreSQL Global Development Group
+subdir('test_aio')
subdir('brin')
subdir('commit_ts')
subdir('delay_execution')
diff --git a/src/test/modules/test_aio/.gitignore b/src/test/modules/test_aio/.gitignore
new file mode 100644
index 00000000000..716e17f5a2a
--- /dev/null
+++ b/src/test/modules/test_aio/.gitignore
@@ -0,0 +1,2 @@
+# Generated subdirectories
+/tmp_check/
diff --git a/src/test/modules/test_aio/Makefile b/src/test/modules/test_aio/Makefile
new file mode 100644
index 00000000000..87d5315ba00
--- /dev/null
+++ b/src/test/modules/test_aio/Makefile
@@ -0,0 +1,27 @@
+# src/test/modules/delay_execution/Makefile
+
+PGFILEDESC = "test_aio - test code for AIO"
+
+MODULE_big = test_aio
+OBJS = \
+ $(WIN32RES) \
+ test_aio.o
+
+EXTENSION = test_aio
+DATA = test_aio--1.0.sql
+
+TAP_TESTS = 1
+
+export enable_injection_points
+export with_liburing
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = src/test/modules/test_aio
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/src/test/modules/test_aio/meson.build b/src/test/modules/test_aio/meson.build
new file mode 100644
index 00000000000..ac846b2b6f3
--- /dev/null
+++ b/src/test/modules/test_aio/meson.build
@@ -0,0 +1,37 @@
+# Copyright (c) 2022-2024, PostgreSQL Global Development Group
+
+test_aio_sources = files(
+ 'test_aio.c',
+)
+
+if host_system == 'windows'
+ test_aio_sources += rc_lib_gen.process(win32ver_rc, extra_args: [
+ '--NAME', 'test_aio',
+ '--FILEDESC', 'test_aio - test code for AIO',])
+endif
+
+test_aio = shared_module('test_aio',
+ test_aio_sources,
+ kwargs: pg_test_mod_args,
+)
+test_install_libs += test_aio
+
+test_install_data += files(
+ 'test_aio.control',
+ 'test_aio--1.0.sql',
+)
+
+tests += {
+ 'name': 'test_aio',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'tap': {
+ 'env': {
+ 'enable_injection_points': get_option('injection_points') ? 'yes' : 'no',
+ 'with_liburing': liburing.found() ? 'yes' : 'no',
+ },
+ 'tests': [
+ 't/001_aio.pl',
+ ],
+ },
+}
diff --git a/src/test/modules/test_aio/t/001_aio.pl b/src/test/modules/test_aio/t/001_aio.pl
new file mode 100644
index 00000000000..225dce08e82
--- /dev/null
+++ b/src/test/modules/test_aio/t/001_aio.pl
@@ -0,0 +1,401 @@
+# Copyright (c) 2025, PostgreSQL Global Development Group
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+
+###
+# Test io_method=worker
+###
+my $node_worker = create_node('worker');
+$node_worker->start();
+
+run_generic_test('worker', $node_worker);
+SKIP:
+{
+ skip 'Injection points not supported by this build', 1
+ unless $ENV{enable_injection_points} eq 'yes';
+ test_inject_worker('worker', $node_worker);
+}
+
+$node_worker->stop();
+
+
+###
+# Test io_method=io_uring
+###
+
+if ($ENV{with_liburing} eq 'yes')
+{
+ my $node_uring = create_node('io_uring');
+ $node_uring->start();
+ run_generic_test('io_uring', $node_uring);
+ $node_uring->stop();
+}
+
+
+###
+# Test io_method=sync
+###
+
+my $node_sync = create_node('sync');
+
+# just to have one test not use the default auto-tuning
+
+$node_sync->append_conf('postgresql.conf', qq(
+io_max_concurrency=4
+));
+
+$node_sync->start();
+run_generic_test('sync', $node_sync);
+$node_sync->stop();
+
+done_testing();
+
+
+###
+# Test Helpers
+###
+
+
+sub create_node
+{
+ my $io_method = shift;
+
+ my $node = PostgreSQL::Test::Cluster->new($io_method);
+
+ # Want to test initdb for each IO method, otherwise we could just reuse
+ # the cluster.
+ $node->init(extra => ['-c', "io_method=$io_method"]);
+
+ $node->append_conf('postgresql.conf', qq(
+shared_preload_libraries=test_aio
+log_min_messages = 'DEBUG3'
+log_statement=all
+restart_after_crash=false
+));
+
+ ok(1, "$io_method: initdb");
+
+ return $node;
+}
+
+
+sub psql_like
+{
+ my $io_method = shift;
+ my $psql = shift;
+ my $name = shift;
+ my $sql = shift;
+ my $expected_stdout = shift;
+ my $expected_stderr = shift;
+ my ($cmdret, $output);
+
+ ($output, $cmdret) = $psql->query($sql);
+
+ like($output, $expected_stdout, "$io_method: $name: expected stdout");
+ like($psql->{stderr}, $expected_stderr, "$io_method: $name: expected stderr");
+ $psql->{stderr} = '';
+}
+
+
+sub test_handle
+{
+ my $io_method = shift;
+ my $node = shift;
+
+ my $psql = $node->background_psql('postgres', on_error_stop => 0);
+
+ # leak warning: implicit xact
+ psql_like($io_method, $psql,
+ "handle_get() leak in implicit xact",
+ qq(SELECT handle_get()),
+ qr/^$/,
+ qr/leaked AIO handle/, "$io_method: leaky handle_get() warns");
+
+ # leak warning: explicit xact
+ psql_like($io_method, $psql,
+ "handle_get() leak in explicit xact",
+ qq(BEGIN; SELECT handle_get(); COMMIT),
+ qr/^$/,
+ qr/leaked AIO handle/);
+
+
+ # leak warning: explicit xact, rollback
+ psql_like($io_method, $psql,
+ "handle_get() leak in explicit xact, rollback",
+ qq(BEGIN; SELECT handle_get(); ROLLBACK;),
+ qr/^$/,
+ qr/leaked AIO handle/);
+
+ # leak warning: subtrans
+ psql_like($io_method, $psql,
+ "handle_get() leak in subxact",
+ qq(BEGIN; SAVEPOINT foo; SELECT handle_get(); COMMIT;),
+ qr/^$/,
+ qr/leaked AIO handle/);
+
+ # leak warning + error: released in different command (thus resowner)
+ psql_like($io_method, $psql,
+ "handle_release() in different command",
+ qq(BEGIN; SELECT handle_get(); SELECT handle_release_last(); COMMIT;),
+ qr/^$/,
+ qr/leaked AIO handle.*release in unexpected state/ms);
+
+ # no leak, release in same command
+ psql_like($io_method, $psql,
+ "handle_release() in same command",
+ qq(BEGIN; SELECT handle_get() UNION ALL SELECT handle_release_last(); COMMIT;),
+ qr/^$/,
+ qr/^$/);
+
+ # normal handle use
+ psql_like($io_method, $psql,
+ "handle_get_release()",
+ qq(SELECT handle_get_release()),
+ qr/^$/,
+ qr/^$/);
+
+ # should error out, API violation
+ psql_like($io_method, $psql,
+ "handle_get_twice()",
+ qq(SELECT handle_get_release()),
+ qr/^$/,
+ qr/^$/);
+
+ # recover after error in implicit xact
+ psql_like($io_method, $psql,
+ "handle error recovery in implicit xact",
+ qq(SELECT handle_get_and_error(); SELECT 'ok', handle_get_release()),
+ qr/^|ok$/,
+ qr/ERROR.*as you command/);
+
+ # recover after error in implicit xact
+ psql_like($io_method, $psql,
+ "handle error recovery in explicit xact",
+ qq(BEGIN; SELECT handle_get_and_error(); SELECT handle_get_release(), 'ok'; COMMIT;),
+ qr/^|ok$/,
+ qr/ERROR.*as you command/);
+
+ # recover after error in subtrans
+ psql_like($io_method, $psql,
+ "handle error recovery in explicit subxact",
+ qq(BEGIN; SAVEPOINT foo; SELECT handle_get_and_error(); ROLLBACK TO SAVEPOINT foo; SELECT handle_get_release(); ROLLBACK;),
+ qr/^|ok$/,
+ qr/ERROR.*as you command/);
+
+ $psql->quit();
+}
+
+
+sub test_batch
+{
+ my $io_method = shift;
+ my $node = shift;
+
+ my $psql = $node->background_psql('postgres', on_error_stop => 0);
+
+ # leak warning & recovery: implicit xact
+ psql_like($io_method, $psql,
+ "batch_start() leak & cleanup in implicit xact",
+ qq(SELECT batch_start()),
+ qr/^$/,
+ qr/open AIO batch at end/, "$io_method: leaky batch_start() warns");
+
+ # leak warning & recovery: explicit xact
+ psql_like($io_method, $psql,
+ "batch_start() leak & cleanup in explicit xact",
+ qq(BEGIN; SELECT batch_start(); COMMIT;),
+ qr/^$/,
+ qr/open AIO batch at end/, "$io_method: leaky batch_start() warns");
+
+
+ # leak warning & recovery: explicit xact, rollback
+ #
+ # FIXME: This doesn't fail right now, due to not getting a chance to do
+ # something at transaction command commit. That's not a correctness issue,
+ # it just means it's a bit harder to find buggy code.
+ #psql_like($io_method, $psql,
+ # "batch_start() leak & cleanup after abort",
+ # qq(BEGIN; SELECT batch_start(); ROLLBACK;),
+ # qr/^$/,
+ # qr/open AIO batch at end/, "$io_method: leaky batch_start() warns");
+
+ # no warning, batch closed in same command
+ psql_like($io_method, $psql,
+ "batch_start(), batch_end() works",
+ qq(SELECT batch_start() UNION ALL SELECT batch_end()),
+ qr/^$/,
+ qr/^$/, "$io_method: batch_start(), batch_end()");
+
+ $psql->quit();
+}
+
+sub test_io_error
+{
+ my $io_method = shift;
+ my $node = shift;
+ my ($ret, $output);
+
+ my $psql = $node->background_psql('postgres', on_error_stop => 0);
+
+ # verify the error is reported in custom C code
+ ($output, $ret) = $psql->query(qq(SELECT read_corrupt_rel_block('tbl_a', 1);));
+ is($ret, 1, "$io_method: read_corrupt_rel_block() fails");
+ like($psql->{stderr}, qr/invalid page in block 1 of relation base\/.*/,
+ "$io_method: read_corrupt_rel_block() reports error");
+ $psql->{stderr} = '';
+
+ # verify the error is reported for bufmgr reads
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_a WHERE ctid = '(1, 1)'));
+ is($ret, 1, "$io_method: tid scan reading corrupt block fails");
+ like($psql->{stderr}, qr/invalid page in block 1 of relation base\/.*/,
+ "$io_method: tid scan reading corrupt block reports error");
+ $psql->{stderr} = '';
+
+ # verify the error is reported for bufmgr reads
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_a WHERE ctid = '(1, 1)'));
+ is($ret, 1, "$io_method: sequential scan reading corrupt block fails");
+ like($psql->{stderr}, qr/invalid page in block 1 of relation base\/.*/,
+ "$io_method: sequential scan reading corrupt block reports error");
+ $psql->{stderr} = '';
+
+ $psql->quit();
+}
+
+
+sub test_inject
+{
+ my $io_method = shift;
+ my $node = shift;
+ my ($ret, $output);
+
+ my $psql = $node->background_psql('postgres', on_error_stop => 0);
+
+ # injected what we'd expect
+ $psql->query_safe(qq(SELECT inj_io_short_read_attach(8192);));
+ $psql->query_safe(qq(SELECT invalidate_rel_block('tbl_b', 2);));
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';));
+ is($ret, 0, "$io_method: injection point not triggering failure succeeds");
+
+ # injected a read shorter than a single block, expecting error
+ $psql->query_safe(qq(SELECT inj_io_short_read_attach(17);));
+ $psql->query_safe(qq(SELECT invalidate_rel_block('tbl_b', 2);));
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';));
+ is($ret, 1, "$io_method: single block short read fails");
+ like($psql->{stderr}, qr/ERROR:.*could not read blocks 2\.\.2 in file "base\/.*": read only 0 of 8192 bytes/,
+ "$io_method: single block short read reports error");
+ $psql->{stderr} = '';
+
+ # shorten multi-block read to a single block, should retry
+ $psql->query_safe(qq(
+SELECT invalidate_rel_block('tbl_b', 0);
+SELECT invalidate_rel_block('tbl_b', 1);
+SELECT invalidate_rel_block('tbl_b', 2);
+SELECT inj_io_short_read_attach(8192);
+ ));
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_b;));
+ is($ret, 0, "$io_method: multi block short read is retried");
+
+ # verify that page verification errors are detected even as part of a
+ # shortened multi-block read (tbl_a, block 1 is corrupted)
+ $psql->query_safe(qq(
+SELECT invalidate_rel_block('tbl_a', 0);
+SELECT invalidate_rel_block('tbl_a', 1);
+SELECT invalidate_rel_block('tbl_a', 2);
+SELECT inj_io_short_read_attach(8192);
+ ));
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_a WHERE ctid < '(2, 1)'));
+ is($ret, 1, "$io_method: shortened multi-block read detects invalid page");
+ like($psql->{stderr}, qr/ERROR:.*invalid page in block 1 of relation base\/.*/,
+ "$io_method: shortened multi-block reads reports invalid page");
+ $psql->{stderr} = '';
+
+ # trigger a hard error, should error out
+ $psql->query_safe(qq(
+SELECT inj_io_short_read_attach(-errno_from_string('EIO'));
+SELECT invalidate_rel_block('tbl_b', 2);
+ ));
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_b; SELECT 1;));
+ is($ret, 1, "$io_method: hard IO error is detected");
+ like($psql->{stderr}, qr/ERROR:.*could not read blocks 2..3 in file \"base\/.*\": Input\/output error/,
+ "$io_method: hard IO error is reported");
+ $psql->{stderr} = '';
+
+ $psql->query_safe(qq(
+SELECT inj_io_short_read_detach();
+ ));
+
+ $psql->quit();
+}
+
+
+sub test_inject_worker
+{
+ my $io_method = shift;
+ my $node = shift;
+ my ($ret, $output);
+
+ my $psql = $node->background_psql('postgres', on_error_stop => 0);
+
+ # trigger a failure to reopen, should error out, but should recover
+ $psql->query_safe(qq(
+SELECT inj_io_reopen_attach();
+SELECT invalidate_rel_block('tbl_b', 1);
+ ));
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_b;));
+ is($ret, 1, "$io_method: failure to open is detected");
+ like($psql->{stderr}, qr/ERROR:.*could not read blocks 1..2 in file "base\/.*": No such file or directory/,
+ "$io_method: failure to open is reported");
+ $psql->{stderr} = '';
+
+ $psql->query_safe(qq(
+SELECT inj_io_reopen_detach();
+ ));
+
+ # check that we indeed recover
+ ($output, $ret) = $psql->query(qq(SELECT count(*) FROM tbl_b;));
+ is($ret, 0, "$io_method: recovers from failure to open ");
+
+
+ $psql->quit();
+}
+
+
+sub run_generic_test
+{
+ my $io_method = shift;
+ my $node = shift;
+
+ is($node->safe_psql('postgres', 'SHOW io_method'),
+ $io_method,
+ "$io_method: io_method set correctly");
+
+ $node->safe_psql('postgres', qq(
+CREATE EXTENSION test_aio;
+CREATE TABLE tbl_a(data int not null) WITH (AUTOVACUUM_ENABLED = false);
+CREATE TABLE tbl_b(data int not null) WITH (AUTOVACUUM_ENABLED = false);
+
+INSERT INTO tbl_a SELECT generate_series(1, 10000);
+INSERT INTO tbl_b SELECT generate_series(1, 10000);
+SELECT grow_rel('tbl_a', 500);
+SELECT grow_rel('tbl_b', 550);
+
+SELECT corrupt_rel_block('tbl_a', 1);
+));
+
+ test_handle($io_method, $node);
+ test_io_error($io_method, $node);
+ test_batch($io_method, $node);
+
+ SKIP:
+ {
+ skip 'Injection points not supported by this build', 1
+ unless $ENV{enable_injection_points} eq 'yes';
+ test_inject($io_method, $node);
+ }
+}
diff --git a/src/test/modules/test_aio/test_aio--1.0.sql b/src/test/modules/test_aio/test_aio--1.0.sql
new file mode 100644
index 00000000000..e7c7c6a6db6
--- /dev/null
+++ b/src/test/modules/test_aio/test_aio--1.0.sql
@@ -0,0 +1,84 @@
+/* src/test/modules/test_aio/test_aio--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION test_aio" to load this file. \quit
+
+
+CREATE FUNCTION errno_from_string(sym text)
+RETURNS pg_catalog.int4 STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+CREATE FUNCTION grow_rel(rel regclass, nblocks int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+CREATE FUNCTION corrupt_rel_block(rel regclass, blockno int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION read_corrupt_rel_block(rel regclass, blockno int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION invalidate_rel_block(rel regclass, blockno int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+/*
+ * Handle related functions
+ */
+CREATE FUNCTION handle_get_and_error()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_get_twice()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_get()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_get_release()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_release_last()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+/*
+ * Batchmode related functions
+ */
+CREATE FUNCTION batch_start()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION batch_end()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+
+/*
+ * Injection point related functions
+ */
+CREATE FUNCTION inj_io_short_read_attach(result int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION inj_io_short_read_detach()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION inj_io_reopen_attach()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION inj_io_reopen_detach()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
new file mode 100644
index 00000000000..15851565853
--- /dev/null
+++ b/src/test/modules/test_aio/test_aio.c
@@ -0,0 +1,518 @@
+/*-------------------------------------------------------------------------
+ *
+ * delay_execution.c
+ * Test module to allow delay between parsing and execution of a query.
+ *
+ * The delay is implemented by taking and immediately releasing a specified
+ * advisory lock. If another process has previously taken that lock, the
+ * current process will be blocked until the lock is released; otherwise,
+ * there's no effect. This allows an isolationtester script to reliably
+ * test behaviors where some specified action happens in another backend
+ * between parsing and execution of any desired query.
+ *
+ * Copyright (c) 2020-2025, PostgreSQL Global Development Group
+ *
+ * IDENTIFICATION
+ * src/test/modules/delay_execution/delay_execution.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "access/relation.h"
+#include "fmgr.h"
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "storage/buf_internals.h"
+#include "storage/bufmgr.h"
+#include "storage/ipc.h"
+#include "storage/lwlock.h"
+#include "utils/builtins.h"
+#include "utils/injection_point.h"
+#include "utils/rel.h"
+
+
+PG_MODULE_MAGIC;
+
+
+typedef struct InjIoErrorState
+{
+ bool enabled_short_read;
+ bool enabled_reopen;
+
+ bool short_read_result_set;
+ int short_read_result;
+} InjIoErrorState;
+
+static InjIoErrorState * inj_io_error_state;
+
+/* Shared memory init callbacks */
+static shmem_request_hook_type prev_shmem_request_hook = NULL;
+static shmem_startup_hook_type prev_shmem_startup_hook = NULL;
+
+
+static PgAioHandle *last_handle;
+
+
+
+static void
+test_aio_shmem_request(void)
+{
+ if (prev_shmem_request_hook)
+ prev_shmem_request_hook();
+
+ RequestAddinShmemSpace(sizeof(InjIoErrorState));
+}
+
+static void
+test_aio_shmem_startup(void)
+{
+ bool found;
+
+ if (prev_shmem_startup_hook)
+ prev_shmem_startup_hook();
+
+ /* Create or attach to the shared memory state */
+ LWLockAcquire(AddinShmemInitLock, LW_EXCLUSIVE);
+
+ inj_io_error_state = ShmemInitStruct("injection_points",
+ sizeof(InjIoErrorState),
+ &found);
+
+ if (!found)
+ {
+ /*
+ * First time through, so initialize. This is shared with the dynamic
+ * initialization using a DSM.
+ */
+ inj_io_error_state->enabled_short_read = false;
+ inj_io_error_state->enabled_reopen = false;
+
+#ifdef USE_INJECTION_POINTS
+ InjectionPointAttach("AIO_PROCESS_COMPLETION_BEFORE_SHARED",
+ "test_aio",
+ "inj_io_short_read",
+ NULL,
+ 0);
+ InjectionPointLoad("AIO_PROCESS_COMPLETION_BEFORE_SHARED");
+
+ InjectionPointAttach("AIO_WORKER_AFTER_REOPEN",
+ "test_aio",
+ "inj_io_reopen",
+ NULL,
+ 0);
+ InjectionPointLoad("AIO_WORKER_AFTER_REOPEN");
+
+#endif
+ }
+ else
+ {
+#ifdef USE_INJECTION_POINTS
+ InjectionPointLoad("AIO_PROCESS_COMPLETION_BEFORE_SHARED");
+ InjectionPointLoad("AIO_WORKER_AFTER_REOPEN");
+ elog(LOG, "injection point loaded");
+#endif
+ }
+
+ LWLockRelease(AddinShmemInitLock);
+}
+
+void
+_PG_init(void)
+{
+ if (!process_shared_preload_libraries_in_progress)
+ return;
+
+ /* Shared memory initialization */
+ prev_shmem_request_hook = shmem_request_hook;
+ shmem_request_hook = test_aio_shmem_request;
+ prev_shmem_startup_hook = shmem_startup_hook;
+ shmem_startup_hook = test_aio_shmem_startup;
+}
+
+
+PG_FUNCTION_INFO_V1(errno_from_string);
+Datum
+errno_from_string(PG_FUNCTION_ARGS)
+{
+ const char *sym = text_to_cstring(PG_GETARG_TEXT_PP(0));
+
+ if (strcmp(sym, "EIO") == 0)
+ PG_RETURN_INT32(EIO);
+ else if (strcmp(sym, "EAGAIN") == 0)
+ PG_RETURN_INT32(EAGAIN);
+ else if (strcmp(sym, "EINTR") == 0)
+ PG_RETURN_INT32(EINTR);
+ else if (strcmp(sym, "ENOSPC") == 0)
+ PG_RETURN_INT32(ENOSPC);
+ else if (strcmp(sym, "EROFS") == 0)
+ PG_RETURN_INT32(EROFS);
+
+ ereport(ERROR,
+ errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg_internal("%s is not a supported errno value", sym));
+ PG_RETURN_INT32(0);
+}
+
+
+PG_FUNCTION_INFO_V1(grow_rel);
+Datum
+grow_rel(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 nblocks = PG_GETARG_UINT32(1);
+ Relation rel;
+#define MAX_BUFFERS_TO_EXTEND_BY 64
+ Buffer victim_buffers[MAX_BUFFERS_TO_EXTEND_BY];
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ while (nblocks > 0)
+ {
+ uint32 extend_by_pages;
+
+ extend_by_pages = Min(nblocks, MAX_BUFFERS_TO_EXTEND_BY);
+
+ ExtendBufferedRelBy(BMR_REL(rel),
+ MAIN_FORKNUM,
+ NULL,
+ 0,
+ extend_by_pages,
+ victim_buffers,
+ &extend_by_pages);
+
+ nblocks -= extend_by_pages;
+
+ for (uint32 i = 0; i < extend_by_pages; i++)
+ {
+ ReleaseBuffer(victim_buffers[i]);
+ }
+ }
+
+ relation_close(rel, NoLock);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(corrupt_rel_block);
+Datum
+corrupt_rel_block(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 block = PG_GETARG_UINT32(1);
+ Relation rel;
+ Buffer buf;
+ Page page;
+ PageHeader ph;
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ buf = ReadBuffer(rel, block);
+ page = BufferGetPage(buf);
+
+ LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
+
+ MarkBufferDirty(buf);
+
+ PageInit(page, BufferGetPageSize(buf), 0);
+
+ ph = (PageHeader) page;
+ ph->pd_special = BLCKSZ + 1;
+
+ FlushOneBuffer(buf);
+
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
+ ReleaseBuffer(buf);
+
+ EvictUnpinnedBuffer(buf);
+
+ relation_close(rel, NoLock);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(read_corrupt_rel_block);
+Datum
+read_corrupt_rel_block(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 block = PG_GETARG_UINT32(1);
+ Relation rel;
+ Buffer buf;
+ BufferDesc *buf_hdr;
+ Page page;
+ PgAioReturn ior;
+ PgAioHandle *ioh;
+ PgAioWaitRef iow;
+ SMgrRelation smgr;
+ uint32 buf_state;
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ /* read buffer without erroring out */
+ buf = ReadBufferExtended(rel, MAIN_FORKNUM, block, RBM_ZERO_AND_LOCK, NULL);
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
+ page = BufferGetBlock(buf);
+
+ ioh = pgaio_io_acquire(CurrentResourceOwner, &ior);
+ pgaio_io_get_wref(ioh, &iow);
+
+ buf_hdr = GetBufferDescriptor(buf - 1);
+ smgr = RelationGetSmgr(rel);
+
+ /* FIXME: even if just a test, we should verify nobody else uses this */
+ buf_state = LockBufHdr(buf_hdr);
+ buf_state &= ~(BM_VALID | BM_DIRTY);
+ UnlockBufHdr(buf_hdr, buf_state);
+
+ StartBufferIO(buf_hdr, true, false);
+
+ pgaio_io_set_handle_data_32(ioh, (uint32 *) &buf, 1);
+ pgaio_io_register_callbacks(ioh, PGAIO_HCB_SHARED_BUFFER_READV);
+
+ smgrstartreadv(ioh, smgr, MAIN_FORKNUM, block,
+ (void *) &page, 1);
+
+ ReleaseBuffer(buf);
+
+ pgaio_wref_wait(&iow);
+
+ if (ior.result.status != ARS_OK)
+ pgaio_result_report(ior.result, &ior.target_data,
+ ior.result.status == ARS_PARTIAL ? WARNING : ERROR);
+
+ relation_close(rel, NoLock);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(invalidate_rel_block);
+Datum
+invalidate_rel_block(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 block = PG_GETARG_UINT32(1);
+ Relation rel;
+ PrefetchBufferResult pr;
+ Buffer buf;
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ /* this is a gross hack, but there's no good API exposed */
+ pr = PrefetchBuffer(rel, MAIN_FORKNUM, block);
+ buf = pr.recent_buffer;
+ elog(LOG, "recent: %d", buf);
+ if (BufferIsValid(buf))
+ {
+ /* if the buffer contents aren't valid, this'll return false */
+ if (ReadRecentBuffer(rel->rd_locator, MAIN_FORKNUM, block, buf))
+ {
+ LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
+ FlushOneBuffer(buf);
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+ ReleaseBuffer(buf);
+
+ if (!EvictUnpinnedBuffer(buf))
+ elog(ERROR, "couldn't unpin");
+ }
+ }
+
+ relation_close(rel, AccessExclusiveLock);
+
+ PG_RETURN_VOID();
+}
+
+#if 0
+PG_FUNCTION_INFO_V1(test_unsubmitted_vs_close);
+Datum
+test_unsubmitted_vs_close(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 block = PG_GETARG_UINT32(1);
+ Relation rel;
+ Buffer buf;
+ Page page;
+ PageHeader ph;
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ buf = ReadBufferExtended(rel, MAIN_FORKNUM, block, RBM_ZERO_AND_LOCK, NULL);
+
+ buf = ReadBuffer(rel, block);
+ page = BufferGetPage(buf);
+
+ EvictUnpinnedBuffer(buf);
+
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
+
+ MarkBufferDirty(buf);
+ ph->pd_special = BLCKSZ + 1;
+
+ /* last_handle = pgaio_io_acquire(); */
+
+ PG_RETURN_VOID();
+}
+#endif
+
+PG_FUNCTION_INFO_V1(handle_get);
+Datum
+handle_get(PG_FUNCTION_ARGS)
+{
+ last_handle = pgaio_io_acquire(CurrentResourceOwner, NULL);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(handle_release_last);
+Datum
+handle_release_last(PG_FUNCTION_ARGS)
+{
+ if (!last_handle)
+ elog(ERROR, "no handle");
+
+ pgaio_io_release(last_handle);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(handle_get_and_error);
+Datum
+handle_get_and_error(PG_FUNCTION_ARGS)
+{
+ pgaio_io_acquire(CurrentResourceOwner, NULL);
+
+ elog(ERROR, "as you command");
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(handle_get_twice);
+Datum
+handle_get_twice(PG_FUNCTION_ARGS)
+{
+ pgaio_io_acquire(CurrentResourceOwner, NULL);
+ pgaio_io_acquire(CurrentResourceOwner, NULL);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(handle_get_release);
+Datum
+handle_get_release(PG_FUNCTION_ARGS)
+{
+ PgAioHandle *handle;
+
+ handle = pgaio_io_acquire(CurrentResourceOwner, NULL);
+ pgaio_io_release(handle);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(batch_start);
+Datum
+batch_start(PG_FUNCTION_ARGS)
+{
+ pgaio_enter_batchmode();
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(batch_end);
+Datum
+batch_end(PG_FUNCTION_ARGS)
+{
+ pgaio_exit_batchmode();
+ PG_RETURN_VOID();
+}
+
+#ifdef USE_INJECTION_POINTS
+extern PGDLLEXPORT void inj_io_short_read(const char *name, const void *private_data);
+extern PGDLLEXPORT void inj_io_reopen(const char *name, const void *private_data);
+
+void
+inj_io_short_read(const char *name, const void *private_data)
+{
+ PgAioHandle *ioh;
+
+ elog(LOG, "short read called: %d", inj_io_error_state->enabled_short_read);
+
+ if (inj_io_error_state->enabled_short_read)
+ {
+ ioh = pgaio_inj_io_get();
+
+ if (inj_io_error_state->short_read_result_set)
+ {
+ elog(LOG, "short read, changing result from %d to %d",
+ ioh->result, inj_io_error_state->short_read_result);
+
+ ioh->result = inj_io_error_state->short_read_result;
+ }
+ }
+}
+
+void
+inj_io_reopen(const char *name, const void *private_data)
+{
+ elog(LOG, "reopen called: %d", inj_io_error_state->enabled_reopen);
+
+ if (inj_io_error_state->enabled_reopen)
+ {
+ elog(ERROR, "injection point triggering failure to reopen ");
+ }
+}
+#endif
+
+PG_FUNCTION_INFO_V1(inj_io_short_read_attach);
+Datum
+inj_io_short_read_attach(PG_FUNCTION_ARGS)
+{
+#ifdef USE_INJECTION_POINTS
+ inj_io_error_state->enabled_short_read = true;
+ inj_io_error_state->short_read_result_set = !PG_ARGISNULL(0);
+ if (inj_io_error_state->short_read_result_set)
+ inj_io_error_state->short_read_result = PG_GETARG_INT32(0);
+#else
+ elog(ERROR, "injection points not supported");
+#endif
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(inj_io_short_read_detach);
+Datum
+inj_io_short_read_detach(PG_FUNCTION_ARGS)
+{
+#ifdef USE_INJECTION_POINTS
+ inj_io_error_state->enabled_short_read = false;
+#else
+ elog(ERROR, "injection points not supported");
+#endif
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(inj_io_reopen_attach);
+Datum
+inj_io_reopen_attach(PG_FUNCTION_ARGS)
+{
+#ifdef USE_INJECTION_POINTS
+ inj_io_error_state->enabled_reopen = true;
+#else
+ elog(ERROR, "injection points not supported");
+#endif
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(inj_io_reopen_detach);
+Datum
+inj_io_reopen_detach(PG_FUNCTION_ARGS)
+{
+#ifdef USE_INJECTION_POINTS
+ inj_io_error_state->enabled_reopen = false;
+#else
+ elog(ERROR, "injection points not supported");
+#endif
+ PG_RETURN_VOID();
+}
diff --git a/src/test/modules/test_aio/test_aio.control b/src/test/modules/test_aio/test_aio.control
new file mode 100644
index 00000000000..cd91c3ed16b
--- /dev/null
+++ b/src/test/modules/test_aio/test_aio.control
@@ -0,0 +1,3 @@
+comment = 'Test code for AIO'
+default_version = '1.0'
+module_pathname = '$libdir/test_aio'
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0023-aio-Add-bounce-buffers.patch (22.8K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/24-v2.4-0023-aio-Add-bounce-buffers.patch)
download | inline diff:
From 6ead714154be99719ce2cf032ef3547b6663ca95 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 17 Feb 2025 20:03:15 -0500
Subject: [PATCH v2.4 23/29] aio: Add bounce buffers
---
src/include/storage/aio.h | 19 ++
src/include/storage/aio_internal.h | 33 ++++
src/include/utils/resowner.h | 2 +
src/backend/storage/aio/README.md | 27 +++
src/backend/storage/aio/aio.c | 178 ++++++++++++++++++
src/backend/storage/aio/aio_init.c | 123 ++++++++++++
src/backend/utils/misc/guc_tables.c | 13 ++
src/backend/utils/misc/postgresql.conf.sample | 2 +
src/backend/utils/resowner/resowner.c | 25 ++-
src/test/modules/test_aio/test_aio--1.0.sql | 21 +++
src/test/modules/test_aio/test_aio.c | 55 ++++++
src/tools/pgindent/typedefs.list | 1 +
12 files changed, 497 insertions(+), 2 deletions(-)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index 6b34422607c..2a50683adc5 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -247,6 +247,10 @@ typedef struct PgAioHandleCallbacks
+typedef struct PgAioBounceBuffer PgAioBounceBuffer;
+
+
+
/* AIO API */
@@ -332,6 +336,20 @@ extern bool pgaio_have_staged(void);
+/* --------------------------------------------------------------------------------
+ * Bounce Buffers
+ * --------------------------------------------------------------------------------
+ */
+
+extern PgAioBounceBuffer *pgaio_bounce_buffer_get(void);
+extern void pgaio_io_assoc_bounce_buffer(PgAioHandle *ioh, PgAioBounceBuffer *bb);
+extern uint32 pgaio_bounce_buffer_id(PgAioBounceBuffer *bb);
+extern void pgaio_bounce_buffer_release(PgAioBounceBuffer *bb);
+extern char *pgaio_bounce_buffer_buffer(PgAioBounceBuffer *bb);
+extern void pgaio_bounce_buffer_release_resowner(struct dlist_node *bb_node, bool on_error);
+
+
+
/* --------------------------------------------------------------------------------
* Other
* --------------------------------------------------------------------------------
@@ -344,6 +362,7 @@ extern void pgaio_closing_fd(int fd);
/* GUCs */
extern PGDLLIMPORT int io_method;
extern PGDLLIMPORT int io_max_concurrency;
+extern PGDLLIMPORT int io_bounce_buffers;
#endif /* AIO_H */
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index ac26aff80b6..9c82e01ed17 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -99,6 +99,12 @@ struct PgAioHandle
*/
uint32 iovec_off;
+ /*
+ * List of bounce_buffers owned by IO. It would suffice to use an index
+ * based linked list here.
+ */
+ slist_head bounce_buffers;
+
/**
* In which list the handle is registered, depends on the state:
* - IDLE, in per-backend list
@@ -135,11 +141,23 @@ struct PgAioHandle
};
+struct PgAioBounceBuffer
+{
+ slist_node node;
+ struct ResourceOwnerData *resowner;
+ dlist_node resowner_node;
+ char *buffer;
+};
+
+
typedef struct PgAioBackend
{
/* index into PgAioCtl->io_handles */
uint32 io_handle_off;
+ /* index into PgAioCtl->bounce_buffers */
+ uint32 bounce_buffers_off;
+
/* IO Handles that currently are not used */
dclist_head idle_ios;
@@ -170,6 +188,12 @@ typedef struct PgAioBackend
* IOs being appended at the end.
*/
dclist_head in_flight_ios;
+
+ /* Bounce Buffers that currently are not used */
+ slist_head idle_bbs;
+
+ /* see handed_out_io */
+ PgAioBounceBuffer *handed_out_bb;
} PgAioBackend;
@@ -195,6 +219,15 @@ typedef struct PgAioCtl
*/
uint64 *handle_data;
+ /*
+ * To perform AIO on buffers that are not located in shared memory (either
+ * because they are not in shared memory or because we need to operate on
+ * a copy, as e.g. the case for writes when checksums are in use)
+ */
+ uint64 bounce_buffers_count;
+ PgAioBounceBuffer *bounce_buffers;
+ char *bounce_buffers_data;
+
uint64 io_handle_count;
PgAioHandle *io_handles;
} PgAioCtl;
diff --git a/src/include/utils/resowner.h b/src/include/utils/resowner.h
index aede4bfc820..7e2ec224169 100644
--- a/src/include/utils/resowner.h
+++ b/src/include/utils/resowner.h
@@ -168,5 +168,7 @@ extern void ResourceOwnerForgetLock(ResourceOwner owner, struct LOCALLOCK *local
struct dlist_node;
extern void ResourceOwnerRememberAioHandle(ResourceOwner owner, struct dlist_node *ioh_node);
extern void ResourceOwnerForgetAioHandle(ResourceOwner owner, struct dlist_node *ioh_node);
+extern void ResourceOwnerRememberAioBounceBuffer(ResourceOwner owner, struct dlist_node *bb_node);
+extern void ResourceOwnerForgetAioBounceBuffer(ResourceOwner owner, struct dlist_node *bb_node);
#endif /* RESOWNER_H */
diff --git a/src/backend/storage/aio/README.md b/src/backend/storage/aio/README.md
index 55e64194ded..191fb21e6a7 100644
--- a/src/backend/storage/aio/README.md
+++ b/src/backend/storage/aio/README.md
@@ -404,6 +404,33 @@ shared memory no less!), completion callbacks instead have to encode errors in
a more compact format that can be converted into an error message.
+### AIO Bounce Buffers
+
+For some uses of AIO there is no convenient memory location as the source /
+destination of an AIO. E.g. when data checksums are enabled, writes from
+shared buffers currently cannot be done directly from shared buffers, as a
+shared buffer lock still allows some modification, e.g., for hint bits(see
+`FlushBuffer()`). If the write were done in-place, such modifications can
+cause the checksum to fail.
+
+For synchronous IO this is solved by copying the buffer to separate memory
+before computing the checksum and using that copy as the source buffer for the
+AIO.
+
+However, for AIO that is not a workable solution:
+- Instead of a single buffer many buffers are required, as many IOs might be
+ in flight
+- When using the [worker method](#worker), the source/target of IO needs to be
+ in shared memory, otherwise the workers won't be able to access the memory.
+
+The AIO subsystem addresses this by providing a limited number of bounce
+buffers that can be used as the source / target for IO. A bounce buffer be
+acquired with `pgaio_bounce_buffer_get()` and multiple bounce buffers can be
+associated with an AIO Handle with `pgaio_io_assoc_bounce_buffer()`.
+
+Bounce buffers are automatically released when the IO completes.
+
+
## Helpers
Using the low-level AIO API introduces too much complexity to do so all over
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index f2d763180d1..fc82908c338 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -61,6 +61,8 @@ static PgAioHandle *pgaio_io_from_wref(PgAioWaitRef *iow, uint64 *ref_generation
static const char *pgaio_io_state_get_name(PgAioHandleState s);
static void pgaio_io_wait(PgAioHandle *ioh, uint64 ref_generation);
+static void pgaio_bounce_buffer_wait_for_free(void);
+
/* Options for io_method. */
const struct config_enum_entry io_method_options[] = {
@@ -75,6 +77,7 @@ const struct config_enum_entry io_method_options[] = {
/* GUCs */
int io_method = DEFAULT_IO_METHOD;
int io_max_concurrency = -1;
+int io_bounce_buffers = -1;
/* global control for AIO */
PgAioCtl *pgaio_ctl;
@@ -642,6 +645,21 @@ pgaio_io_reclaim(PgAioHandle *ioh)
}
}
+ /* reclaim all associated bounce buffers */
+ if (!slist_is_empty(&ioh->bounce_buffers))
+ {
+ slist_mutable_iter it;
+
+ slist_foreach_modify(it, &ioh->bounce_buffers)
+ {
+ PgAioBounceBuffer *bb = slist_container(PgAioBounceBuffer, node, it.cur);
+
+ slist_delete_current(&it);
+
+ slist_push_head(&pgaio_my_backend->idle_bbs, &bb->node);
+ }
+ }
+
if (ioh->resowner)
{
ResourceOwnerForgetAioHandle(ioh->resowner, &ioh->resowner_node);
@@ -1013,6 +1031,166 @@ pgaio_submit_staged(void)
+/* --------------------------------------------------------------------------------
+ * Functions primarily related to PgAioBounceBuffer
+ * --------------------------------------------------------------------------------
+ */
+
+PgAioBounceBuffer *
+pgaio_bounce_buffer_get(void)
+{
+ PgAioBounceBuffer *bb = NULL;
+ slist_node *node;
+
+ if (pgaio_my_backend->handed_out_bb != NULL)
+ elog(ERROR, "can only hand out one BB");
+
+ /*
+ * XXX: It probably is not a good idea to have bounce buffers be per
+ * backend, that's a fair bit of memory.
+ */
+ if (slist_is_empty(&pgaio_my_backend->idle_bbs))
+ {
+ pgaio_bounce_buffer_wait_for_free();
+ }
+
+ node = slist_pop_head_node(&pgaio_my_backend->idle_bbs);
+ bb = slist_container(PgAioBounceBuffer, node, node);
+
+ pgaio_my_backend->handed_out_bb = bb;
+
+ bb->resowner = CurrentResourceOwner;
+ ResourceOwnerRememberAioBounceBuffer(bb->resowner, &bb->resowner_node);
+
+ return bb;
+}
+
+void
+pgaio_io_assoc_bounce_buffer(PgAioHandle *ioh, PgAioBounceBuffer *bb)
+{
+ if (pgaio_my_backend->handed_out_bb != bb)
+ elog(ERROR, "can only assign handed out BB");
+ pgaio_my_backend->handed_out_bb = NULL;
+
+ /*
+ * There can be many bounce buffers assigned in case of vectorized IOs.
+ */
+ slist_push_head(&ioh->bounce_buffers, &bb->node);
+
+ /* once associated with an IO, the IO has ownership */
+ ResourceOwnerForgetAioBounceBuffer(bb->resowner, &bb->resowner_node);
+ bb->resowner = NULL;
+}
+
+uint32
+pgaio_bounce_buffer_id(PgAioBounceBuffer *bb)
+{
+ return bb - pgaio_ctl->bounce_buffers;
+}
+
+void
+pgaio_bounce_buffer_release(PgAioBounceBuffer *bb)
+{
+ if (pgaio_my_backend->handed_out_bb != bb)
+ elog(ERROR, "can only release handed out BB");
+
+ slist_push_head(&pgaio_my_backend->idle_bbs, &bb->node);
+ pgaio_my_backend->handed_out_bb = NULL;
+
+ ResourceOwnerForgetAioBounceBuffer(bb->resowner, &bb->resowner_node);
+ bb->resowner = NULL;
+}
+
+void
+pgaio_bounce_buffer_release_resowner(dlist_node *bb_node, bool on_error)
+{
+ PgAioBounceBuffer *bb = dlist_container(PgAioBounceBuffer, resowner_node, bb_node);
+
+ Assert(bb->resowner);
+
+ if (!on_error)
+ elog(WARNING, "leaked AIO bounce buffer");
+
+ pgaio_bounce_buffer_release(bb);
+}
+
+char *
+pgaio_bounce_buffer_buffer(PgAioBounceBuffer *bb)
+{
+ return bb->buffer;
+}
+
+static void
+pgaio_bounce_buffer_wait_for_free(void)
+{
+ static uint32 lastpos = 0;
+
+ if (pgaio_my_backend->num_staged_ios > 0)
+ {
+ pgaio_debug(DEBUG2, "submitting %d, while acquiring free bb",
+ pgaio_my_backend->num_staged_ios);
+ pgaio_submit_staged();
+ }
+
+ for (uint32 i = lastpos; i < lastpos + io_max_concurrency; i++)
+ {
+ uint32 thisoff = pgaio_my_backend->io_handle_off + (i % io_max_concurrency);
+ PgAioHandle *ioh = &pgaio_ctl->io_handles[thisoff];
+
+ switch (ioh->state)
+ {
+ case PGAIO_HS_IDLE:
+ case PGAIO_HS_HANDED_OUT:
+ continue;
+ case PGAIO_HS_DEFINED: /* should have been submitted above */
+ case PGAIO_HS_STAGED:
+ elog(ERROR, "shouldn't get here with io:%d in state %d",
+ pgaio_io_get_id(ioh), ioh->state);
+ break;
+ case PGAIO_HS_COMPLETED_IO:
+ case PGAIO_HS_SUBMITTED:
+ if (!slist_is_empty(&ioh->bounce_buffers))
+ {
+ pgaio_debug_io(DEBUG2, ioh,
+ "waiting for IO to reclaim BB with %d in flight",
+ dclist_count(&pgaio_my_backend->in_flight_ios));
+
+ /* see comment in pgaio_io_wait_for_free() about raciness */
+ pgaio_io_wait(ioh, ioh->generation);
+
+ if (slist_is_empty(&pgaio_my_backend->idle_bbs))
+ elog(WARNING, "empty after wait");
+
+ if (!slist_is_empty(&pgaio_my_backend->idle_bbs))
+ {
+ lastpos = i;
+ return;
+ }
+ }
+ break;
+ case PGAIO_HS_COMPLETED_SHARED:
+ case PGAIO_HS_COMPLETED_LOCAL:
+ /* reclaim */
+ pgaio_io_reclaim(ioh);
+
+ if (!slist_is_empty(&pgaio_my_backend->idle_bbs))
+ {
+ lastpos = i;
+ return;
+ }
+ break;
+ }
+ }
+
+ /*
+ * The submission above could have caused the IO to complete at any time.
+ */
+ if (slist_is_empty(&pgaio_my_backend->idle_bbs))
+ elog(PANIC, "no more bbs");
+}
+
+
+
/* --------------------------------------------------------------------------------
* Other
* --------------------------------------------------------------------------------
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index 87eac5e961c..c6b29a7b134 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -82,6 +82,32 @@ AioHandleDataShmemSize(void)
io_max_concurrency));
}
+static Size
+AioBounceBufferDescShmemSize(void)
+{
+ Size sz;
+
+ /* PgAioBounceBuffer itself */
+ sz = mul_size(sizeof(PgAioBounceBuffer),
+ mul_size(AioProcs(), io_bounce_buffers));
+
+ return sz;
+}
+
+static Size
+AioBounceBufferDataShmemSize(void)
+{
+ Size sz;
+
+ /* and the associated buffer */
+ sz = mul_size(BLCKSZ,
+ mul_size(io_bounce_buffers, AioProcs()));
+ /* memory for alignment */
+ sz += BLCKSZ;
+
+ return sz;
+}
+
/*
* Choose a suitable value for io_max_concurrency.
*
@@ -107,6 +133,33 @@ AioChooseMaxConcurrency(void)
return Min(max_proportional_pins, 64);
}
+/*
+ * Choose a suitable value for io_bounce_buffers.
+ *
+ * It's very unlikely to be useful to allocate more bounce buffers for each
+ * backend than the backend is allowed to pin. Additionally, bounce buffers
+ * currently are used for writes, it seems very uncommon for more than 10% of
+ * shared_buffers to be written out concurrently.
+ *
+ * XXX: This quickly can take up significant amounts of memory, the logic
+ * should probably fine tuned.
+ */
+static int
+AioChooseBounceBuffers(void)
+{
+ uint32 max_backends;
+ int max_proportional_pins;
+
+ /* Similar logic to LimitAdditionalPins() */
+ max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
+ max_proportional_pins = (NBuffers / 10) / max_backends;
+
+ max_proportional_pins = Max(max_proportional_pins, 1);
+
+ /* apply upper limit */
+ return Min(max_proportional_pins, 256);
+}
+
Size
AioShmemSize(void)
{
@@ -130,11 +183,31 @@ AioShmemSize(void)
PGC_S_OVERRIDE);
}
+
+ /*
+ * If io_bounce_buffers is -1, we automatically choose a suitable value.
+ *
+ * See also comment above.
+ */
+ if (io_bounce_buffers == -1)
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", AioChooseBounceBuffers());
+ SetConfigOption("io_bounce_buffers", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+ if (io_bounce_buffers == -1) /* failed to apply it? */
+ SetConfigOption("io_bounce_buffers", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+
sz = add_size(sz, AioCtlShmemSize());
sz = add_size(sz, AioBackendShmemSize());
sz = add_size(sz, AioHandleShmemSize());
sz = add_size(sz, AioHandleIOVShmemSize());
sz = add_size(sz, AioHandleDataShmemSize());
+ sz = add_size(sz, AioBounceBufferDescShmemSize());
+ sz = add_size(sz, AioBounceBufferDataShmemSize());
if (pgaio_method_ops->shmem_size)
sz = add_size(sz, pgaio_method_ops->shmem_size());
@@ -149,6 +222,9 @@ AioShmemInit(void)
uint32 io_handle_off = 0;
uint32 iovec_off = 0;
uint32 per_backend_iovecs = io_max_concurrency * PG_IOV_MAX;
+ uint32 bounce_buffers_off = 0;
+ uint32 per_backend_bb = io_bounce_buffers;
+ char *bounce_buffers_data;
pgaio_ctl = (PgAioCtl *)
ShmemInitStruct("AioCtl", AioCtlShmemSize(), &found);
@@ -160,6 +236,7 @@ AioShmemInit(void)
pgaio_ctl->io_handle_count = AioProcs() * io_max_concurrency;
pgaio_ctl->iovec_count = AioProcs() * per_backend_iovecs;
+ pgaio_ctl->bounce_buffers_count = AioProcs() * per_backend_bb;
pgaio_ctl->backend_state = (PgAioBackend *)
ShmemInitStruct("AioBackend", AioBackendShmemSize(), &found);
@@ -172,6 +249,40 @@ AioShmemInit(void)
pgaio_ctl->handle_data = (uint64 *)
ShmemInitStruct("AioHandleData", AioHandleDataShmemSize(), &found);
+ pgaio_ctl->bounce_buffers = (PgAioBounceBuffer *)
+ ShmemInitStruct("AioBounceBufferDesc", AioBounceBufferDescShmemSize(),
+ &found);
+
+ bounce_buffers_data =
+ ShmemInitStruct("AioBounceBufferData", AioBounceBufferDataShmemSize(),
+ &found);
+ bounce_buffers_data =
+ (char *) TYPEALIGN(BLCKSZ, (uintptr_t) bounce_buffers_data);
+ pgaio_ctl->bounce_buffers_data = bounce_buffers_data;
+
+
+ /* Initialize IO handles. */
+ for (uint64 i = 0; i < pgaio_ctl->io_handle_count; i++)
+ {
+ PgAioHandle *ioh = &pgaio_ctl->io_handles[i];
+
+ ioh->op = PGAIO_OP_INVALID;
+ ioh->target = PGAIO_TID_INVALID;
+ ioh->state = PGAIO_HS_IDLE;
+
+ slist_init(&ioh->bounce_buffers);
+ }
+
+ /* Initialize Bounce Buffers. */
+ for (uint64 i = 0; i < pgaio_ctl->bounce_buffers_count; i++)
+ {
+ PgAioBounceBuffer *bb = &pgaio_ctl->bounce_buffers[i];
+
+ bb->buffer = bounce_buffers_data;
+ bounce_buffers_data += BLCKSZ;
+ }
+
+
for (int procno = 0; procno < AioProcs(); procno++)
{
PgAioBackend *bs = &pgaio_ctl->backend_state[procno];
@@ -179,9 +290,13 @@ AioShmemInit(void)
bs->io_handle_off = io_handle_off;
io_handle_off += io_max_concurrency;
+ bs->bounce_buffers_off = bounce_buffers_off;
+ bounce_buffers_off += per_backend_bb;
+
dclist_init(&bs->idle_ios);
memset(bs->staged_ios, 0, sizeof(PgAioHandle *) * PGAIO_SUBMIT_BATCH_SIZE);
dclist_init(&bs->in_flight_ios);
+ slist_init(&bs->idle_bbs);
/* initialize per-backend IOs */
for (int i = 0; i < io_max_concurrency; i++)
@@ -203,6 +318,14 @@ AioShmemInit(void)
dclist_push_tail(&bs->idle_ios, &ioh->node);
iovec_off += PG_IOV_MAX;
}
+
+ /* initialize per-backend bounce buffers */
+ for (int i = 0; i < per_backend_bb; i++)
+ {
+ PgAioBounceBuffer *bb = &pgaio_ctl->bounce_buffers[bs->bounce_buffers_off + i];
+
+ slist_push_head(&bs->idle_bbs, &bb->node);
+ }
}
out:
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 15954f42d4e..58120b7add2 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -3266,6 +3266,19 @@ struct config_int ConfigureNamesInt[] =
check_io_max_concurrency, NULL, NULL
},
+ {
+ {"io_bounce_buffers",
+ PGC_POSTMASTER,
+ RESOURCES_IO,
+ gettext_noop("Number of IO Bounce Buffers reserved for each backend."),
+ NULL,
+ GUC_UNIT_BLOCKS
+ },
+ &io_bounce_buffers,
+ -1, -1, 4096,
+ NULL, NULL, NULL
+ },
+
{
{"io_workers",
PGC_SIGHUP,
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index 50fde0ba2c3..2214846a0b1 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -206,6 +206,8 @@
# -1 sets based on shared_buffers
# (change requires restart)
#io_workers = 3 # 1-32;
+#io_bounce_buffers = -1 # -1 sets based on shared_buffers
+ # (change requires restart)
# - Worker Processes -
diff --git a/src/backend/utils/resowner/resowner.c b/src/backend/utils/resowner/resowner.c
index 76b9cec1e26..de00346b549 100644
--- a/src/backend/utils/resowner/resowner.c
+++ b/src/backend/utils/resowner/resowner.c
@@ -159,10 +159,11 @@ struct ResourceOwnerData
LOCALLOCK *locks[MAX_RESOWNER_LOCKS]; /* list of owned locks */
/*
- * AIO handles need be registered in critical sections and therefore
- * cannot use the normal ResoureElem mechanism.
+ * AIO handles & bounce buffers need be registered in critical sections
+ * and therefore cannot use the normal ResoureElem mechanism.
*/
dlist_head aio_handles;
+ dlist_head aio_bounce_buffers;
};
@@ -434,6 +435,7 @@ ResourceOwnerCreate(ResourceOwner parent, const char *name)
}
dlist_init(&owner->aio_handles);
+ dlist_init(&owner->aio_bounce_buffers);
return owner;
}
@@ -742,6 +744,13 @@ ResourceOwnerReleaseInternal(ResourceOwner owner,
pgaio_io_release_resowner(node, !isCommit);
}
+
+ while (!dlist_is_empty(&owner->aio_bounce_buffers))
+ {
+ dlist_node *node = dlist_head_node(&owner->aio_bounce_buffers);
+
+ pgaio_bounce_buffer_release_resowner(node, !isCommit);
+ }
}
else if (phase == RESOURCE_RELEASE_LOCKS)
{
@@ -1111,3 +1120,15 @@ ResourceOwnerForgetAioHandle(ResourceOwner owner, struct dlist_node *ioh_node)
{
dlist_delete_from(&owner->aio_handles, ioh_node);
}
+
+void
+ResourceOwnerRememberAioBounceBuffer(ResourceOwner owner, struct dlist_node *ioh_node)
+{
+ dlist_push_tail(&owner->aio_bounce_buffers, ioh_node);
+}
+
+void
+ResourceOwnerForgetAioBounceBuffer(ResourceOwner owner, struct dlist_node *ioh_node)
+{
+ dlist_delete_from(&owner->aio_bounce_buffers, ioh_node);
+}
diff --git a/src/test/modules/test_aio/test_aio--1.0.sql b/src/test/modules/test_aio/test_aio--1.0.sql
index e7c7c6a6db6..822211f5dd4 100644
--- a/src/test/modules/test_aio/test_aio--1.0.sql
+++ b/src/test/modules/test_aio/test_aio--1.0.sql
@@ -63,6 +63,27 @@ RETURNS pg_catalog.void STRICT
AS 'MODULE_PATHNAME' LANGUAGE C;
+CREATE FUNCTION bb_get_and_error()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION bb_get_twice()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION bb_get()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION bb_get_release()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION bb_release_last()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
/*
* Injection point related functions
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index 15851565853..15a548cff1a 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -53,6 +53,7 @@ static shmem_startup_hook_type prev_shmem_startup_hook = NULL;
static PgAioHandle *last_handle;
+static PgAioBounceBuffer *last_bb;
@@ -427,6 +428,60 @@ batch_end(PG_FUNCTION_ARGS)
PG_RETURN_VOID();
}
+PG_FUNCTION_INFO_V1(bb_get);
+Datum
+bb_get(PG_FUNCTION_ARGS)
+{
+ last_bb = pgaio_bounce_buffer_get();
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(bb_release_last);
+Datum
+bb_release_last(PG_FUNCTION_ARGS)
+{
+ if (!last_bb)
+ elog(ERROR, "no bb");
+
+ pgaio_bounce_buffer_release(last_bb);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(bb_get_and_error);
+Datum
+bb_get_and_error(PG_FUNCTION_ARGS)
+{
+ pgaio_bounce_buffer_get();
+
+ elog(ERROR, "as you command");
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(bb_get_twice);
+Datum
+bb_get_twice(PG_FUNCTION_ARGS)
+{
+ pgaio_bounce_buffer_get();
+ pgaio_bounce_buffer_get();
+
+ PG_RETURN_VOID();
+}
+
+
+PG_FUNCTION_INFO_V1(bb_get_release);
+Datum
+bb_get_release(PG_FUNCTION_ARGS)
+{
+ PgAioBounceBuffer *bb;
+
+ bb = pgaio_bounce_buffer_get();
+ pgaio_bounce_buffer_release(bb);
+
+ PG_RETURN_VOID();
+}
+
#ifdef USE_INJECTION_POINTS
extern PGDLLEXPORT void inj_io_short_read(const char *name, const void *private_data);
extern PGDLLEXPORT void inj_io_reopen(const char *name, const void *private_data);
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index d4734b85c0d..d216785c3c8 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2113,6 +2113,7 @@ PermutationStep
PermutationStepBlocker
PermutationStepBlockerType
PgAioBackend
+PgAioBounceBuffer
PgAioCtl
PgAioHandle
PgAioHandleCallbackID
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0024-bufmgr-Implement-AIO-write-support.patch (6.0K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/25-v2.4-0024-bufmgr-Implement-AIO-write-support.patch)
download | inline diff:
From 3240ca639291ef0fa86654a8fd411081ea6f2723 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 16:09:51 -0500
Subject: [PATCH v2.4 24/29] bufmgr: Implement AIO write support
As of this commit there are no users of these AIO facilities, that'll come in
later commits.
Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/storage/aio.h | 2 +
src/include/storage/bufmgr.h | 2 +
src/backend/storage/aio/aio_callback.c | 2 +
src/backend/storage/buffer/bufmgr.c | 85 ++++++++++++++++++++++++++
4 files changed, 91 insertions(+)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index 2a50683adc5..1901b839aff 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -180,8 +180,10 @@ typedef enum PgAioHandleCallbackID
PGAIO_HCB_MD_WRITEV,
PGAIO_HCB_SHARED_BUFFER_READV,
+ PGAIO_HCB_SHARED_BUFFER_WRITEV,
PGAIO_HCB_LOCAL_BUFFER_READV,
+ PGAIO_HCB_LOCAL_BUFFER_WRITEV,
} PgAioHandleCallbackID;
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index dc8fe197d6f..655885ff2d0 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -186,7 +186,9 @@ extern PGDLLIMPORT int32 *LocalRefCount;
struct PgAioHandleCallbacks;
extern const struct PgAioHandleCallbacks aio_shared_buffer_readv_cb;
+extern const struct PgAioHandleCallbacks aio_shared_buffer_writev_cb;
extern const struct PgAioHandleCallbacks aio_local_buffer_readv_cb;
+extern const struct PgAioHandleCallbacks aio_local_buffer_writev_cb;
/* upper limit for effective_io_concurrency */
diff --git a/src/backend/storage/aio/aio_callback.c b/src/backend/storage/aio/aio_callback.c
index 6afdaaa434b..2e2cab305c0 100644
--- a/src/backend/storage/aio/aio_callback.c
+++ b/src/backend/storage/aio/aio_callback.c
@@ -45,8 +45,10 @@ static const PgAioHandleCallbacksEntry aio_handle_cbs[] = {
CALLBACK_ENTRY(PGAIO_HCB_MD_WRITEV, aio_md_writev_cb),
CALLBACK_ENTRY(PGAIO_HCB_SHARED_BUFFER_READV, aio_shared_buffer_readv_cb),
+ CALLBACK_ENTRY(PGAIO_HCB_SHARED_BUFFER_WRITEV, aio_shared_buffer_writev_cb),
CALLBACK_ENTRY(PGAIO_HCB_LOCAL_BUFFER_READV, aio_local_buffer_readv_cb),
+ CALLBACK_ENTRY(PGAIO_HCB_LOCAL_BUFFER_WRITEV, aio_local_buffer_writev_cb),
#undef CALLBACK_ENTRY
};
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b641bc3982b..65acc891ef1 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -6530,6 +6530,42 @@ SharedBufferCompleteRead(int buf_off, Buffer buffer, int mode, bool failed)
return result;
}
+static uint64
+BufferCompleteWriteShared(Buffer buffer, bool release_lock, bool failed)
+{
+ BufferDesc *bufHdr;
+ bool result = false;
+
+ Assert(BufferIsValid(buffer));
+
+ bufHdr = GetBufferDescriptor(buffer - 1);
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ uint32 buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ Assert(buf_state & BM_VALID);
+ Assert(buf_state & BM_TAG_VALID);
+ Assert(buf_state & BM_IO_IN_PROGRESS);
+ Assert(buf_state & BM_DIRTY);
+ }
+#endif
+
+ TerminateBufferIO(bufHdr, /* clear_dirty = */ true,
+ failed ? BM_IO_ERROR : 0,
+ /* forget_owner = */ false,
+ /* syncio = */ false);
+
+ /*
+ * The initiator of IO is not managing the lock (i.e. called
+ * LWLockDisown()), we are.
+ */
+ if (release_lock)
+ LWLockReleaseDisowned(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
+
+ return result;
+}
+
/*
* Helper to prepare IO on shared buffers for execution, shared between reads
* and writes.
@@ -6610,6 +6646,12 @@ shared_buffer_readv_stage(PgAioHandle *ioh)
shared_buffer_stage_common(ioh, false);
}
+static void
+shared_buffer_writev_stage(PgAioHandle *ioh)
+{
+ shared_buffer_stage_common(ioh, true);
+}
+
static void
buffer_readv_report(PgAioResult result, const PgAioTargetData *target_data, int elevel)
{
@@ -6705,6 +6747,33 @@ shared_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
return buffer_readv_complete_common(ioh, prior_result, false);
}
+static PgAioResult
+shared_buffer_writev_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ PgAioResult result = prior_result;
+ uint64 *io_data;
+ uint8 handle_data_len;
+
+ ereport(DEBUG5,
+ errmsg("%s: %d %d", __func__, prior_result.status, prior_result.result),
+ errhidestmt(true), errhidecontext(true));
+
+ io_data = pgaio_io_get_handle_data(ioh, &handle_data_len);
+
+ /* FIXME: handle outright errors */
+
+ for (int io_data_off = 0; io_data_off < handle_data_len; io_data_off++)
+ {
+ Buffer buf = io_data[io_data_off];
+
+ /* FIXME: handle short writes / failures */
+ /* FIXME: ioh->target_data.shared_buffer.release_lock */
+ BufferCompleteWriteShared(buf, true, false);
+ }
+
+ return result;
+}
+
/*
* Helper to stage IO on local buffers for execution, shared between reads
* and writes.
@@ -6747,7 +6816,16 @@ static PgAioResult
local_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
{
return buffer_readv_complete_common(ioh, prior_result, true);
+}
+static void
+local_buffer_writev_stage(PgAioHandle *ioh)
+{
+ /*
+ * Currently this is unreachable as the only write support is for
+ * checkpointer / bgwriter, which don't deal with local buffers.
+ */
+ elog(ERROR, "not yet");
}
@@ -6756,6 +6834,10 @@ const struct PgAioHandleCallbacks aio_shared_buffer_readv_cb = {
.complete_shared = shared_buffer_readv_complete,
.report = buffer_readv_report,
};
+const struct PgAioHandleCallbacks aio_shared_buffer_writev_cb = {
+ .stage = shared_buffer_writev_stage,
+ .complete_shared = shared_buffer_writev_complete,
+};
const struct PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.stage = local_buffer_readv_stage,
@@ -6768,3 +6850,6 @@ const struct PgAioHandleCallbacks aio_local_buffer_readv_cb = {
.complete_local = local_buffer_readv_complete,
.report = buffer_readv_report,
};
+const struct PgAioHandleCallbacks aio_local_buffer_writev_cb = {
+ .stage = local_buffer_writev_stage,
+};
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0025-aio-Add-IO-queue-helper.patch (7.2K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/26-v2.4-0025-aio-Add-IO-queue-helper.patch)
download | inline diff:
From 7cfb9661bd6accaba0b4dc218c811975af96235b Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 13:44:50 -0500
Subject: [PATCH v2.4 25/29] aio: Add IO queue helper
This is likely never going to anywhere - Thomas Munro is working on something
more complete. But I needed a way to exercise aio for checkpointer / bgwriter.
---
src/include/storage/io_queue.h | 31 +++++
src/backend/storage/aio/Makefile | 1 +
src/backend/storage/aio/io_queue.c | 198 ++++++++++++++++++++++++++++
src/backend/storage/aio/meson.build | 1 +
src/tools/pgindent/typedefs.list | 2 +
5 files changed, 233 insertions(+)
create mode 100644 src/include/storage/io_queue.h
create mode 100644 src/backend/storage/aio/io_queue.c
diff --git a/src/include/storage/io_queue.h b/src/include/storage/io_queue.h
new file mode 100644
index 00000000000..f5e1bc07ff3
--- /dev/null
+++ b/src/include/storage/io_queue.h
@@ -0,0 +1,31 @@
+/*-------------------------------------------------------------------------
+ *
+ * io_queue.h
+ * Mechanism for tracking many IOs
+ *
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/io_queue.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef IO_QUEUE_H
+#define IO_QUEUE_H
+
+struct IOQueue;
+typedef struct IOQueue IOQueue;
+
+struct PgAioWaitRef;
+
+extern IOQueue *io_queue_create(int depth, int flags);
+extern void io_queue_track(IOQueue *ioq, const struct PgAioWaitRef *iow);
+extern void io_queue_wait_one(IOQueue *ioq);
+extern void io_queue_wait_all(IOQueue *ioq);
+extern bool io_queue_is_empty(IOQueue *ioq);
+extern void io_queue_reserve(IOQueue *ioq);
+extern struct PgAioHandle *io_queue_acquire_io(IOQueue *ioq);
+extern void io_queue_free(IOQueue *ioq);
+
+#endif /* IO_QUEUE_H */
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index 3f2469cc399..86fa4276fda 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -15,6 +15,7 @@ OBJS = \
aio_init.o \
aio_io.o \
aio_target.o \
+ io_queue.o \
method_io_uring.o \
method_sync.o \
method_worker.o \
diff --git a/src/backend/storage/aio/io_queue.c b/src/backend/storage/aio/io_queue.c
new file mode 100644
index 00000000000..62ad06c8bfe
--- /dev/null
+++ b/src/backend/storage/aio/io_queue.c
@@ -0,0 +1,198 @@
+/*-------------------------------------------------------------------------
+ *
+ * io_queue.c
+ * AIO - Mechanism for tracking many IOs
+ *
+ * Portions Copyright (c) 1996-2025, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/io_queue.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "lib/ilist.h"
+#include "storage/aio.h"
+#include "storage/io_queue.h"
+#include "utils/resowner.h"
+
+
+
+typedef struct TrackedIO
+{
+ PgAioWaitRef iow;
+ dlist_node node;
+} TrackedIO;
+
+struct IOQueue
+{
+ int depth;
+ int unsubmitted;
+
+ bool has_reserved;
+
+ dclist_head idle;
+ dclist_head in_progress;
+
+ TrackedIO tracked_ios[FLEXIBLE_ARRAY_MEMBER];
+};
+
+
+IOQueue *
+io_queue_create(int depth, int flags)
+{
+ size_t sz;
+ IOQueue *ioq;
+
+ sz = offsetof(IOQueue, tracked_ios)
+ + sizeof(TrackedIO) * depth;
+
+ ioq = palloc0(sz);
+
+ ioq->depth = 0;
+
+ for (int i = 0; i < depth; i++)
+ {
+ TrackedIO *tio = &ioq->tracked_ios[i];
+
+ pgaio_wref_clear(&tio->iow);
+ dclist_push_tail(&ioq->idle, &tio->node);
+ }
+
+ return ioq;
+}
+
+void
+io_queue_wait_one(IOQueue *ioq)
+{
+ while (!dclist_is_empty(&ioq->in_progress))
+ {
+ /* FIXME: Should we really pop here already? */
+ dlist_node *node = dclist_pop_head_node(&ioq->in_progress);
+ TrackedIO *tio = dclist_container(TrackedIO, node, node);
+
+ pgaio_wref_wait(&tio->iow);
+ dclist_push_head(&ioq->idle, &tio->node);
+ }
+}
+
+void
+io_queue_reserve(IOQueue *ioq)
+{
+ if (ioq->has_reserved)
+ return;
+
+ if (dclist_is_empty(&ioq->idle))
+ io_queue_wait_one(ioq);
+
+ Assert(!dclist_is_empty(&ioq->idle));
+
+ ioq->has_reserved = true;
+}
+
+PgAioHandle *
+io_queue_acquire_io(IOQueue *ioq)
+{
+ PgAioHandle *ioh;
+
+ io_queue_reserve(ioq);
+
+ Assert(!dclist_is_empty(&ioq->idle));
+
+ if (!io_queue_is_empty(ioq))
+ {
+ ioh = pgaio_io_acquire_nb(CurrentResourceOwner, NULL);
+ if (ioh == NULL)
+ {
+ /*
+ * Need to wait for all IOs, blocking might not be legal in the
+ * context.
+ *
+ * XXX: This doesn't make a whole lot of sense, we're also
+ * blocking here. What was I smoking when I wrote the above?
+ */
+ io_queue_wait_all(ioq);
+ ioh = pgaio_io_acquire(CurrentResourceOwner, NULL);
+ }
+ }
+ else
+ {
+ ioh = pgaio_io_acquire(CurrentResourceOwner, NULL);
+ }
+
+ return ioh;
+}
+
+void
+io_queue_track(IOQueue *ioq, const struct PgAioWaitRef *iow)
+{
+ dlist_node *node;
+ TrackedIO *tio;
+
+ Assert(ioq->has_reserved);
+ ioq->has_reserved = false;
+
+ Assert(!dclist_is_empty(&ioq->idle));
+
+ node = dclist_pop_head_node(&ioq->idle);
+ tio = dclist_container(TrackedIO, node, node);
+
+ tio->iow = *iow;
+
+ dclist_push_tail(&ioq->in_progress, &tio->node);
+
+ ioq->unsubmitted++;
+
+ /*
+ * XXX: Should have some smarter logic here. We don't want to wait too
+ * long to submit, that'll mean we're more likely to block. But we also
+ * don't want to have the overhead of submitting every IO individually.
+ */
+ if (ioq->unsubmitted >= 4)
+ {
+ pgaio_submit_staged();
+ ioq->unsubmitted = 0;
+ }
+}
+
+void
+io_queue_wait_all(IOQueue *ioq)
+{
+ while (!dclist_is_empty(&ioq->in_progress))
+ {
+ /* wait for the last IO to minimize unnecessary wakeups */
+ dlist_node *node = dclist_tail_node(&ioq->in_progress);
+ TrackedIO *tio = dclist_container(TrackedIO, node, node);
+
+ if (!pgaio_wref_check_done(&tio->iow))
+ {
+ ereport(DEBUG3,
+ errmsg("io_queue_wait_all for io:%d",
+ pgaio_wref_get_id(&tio->iow)),
+ errhidestmt(true),
+ errhidecontext(true));
+
+ pgaio_wref_wait(&tio->iow);
+ }
+
+ dclist_delete_from(&ioq->in_progress, &tio->node);
+ dclist_push_head(&ioq->idle, &tio->node);
+ }
+}
+
+bool
+io_queue_is_empty(IOQueue *ioq)
+{
+ return dclist_is_empty(&ioq->in_progress);
+}
+
+void
+io_queue_free(IOQueue *ioq)
+{
+ io_queue_wait_all(ioq);
+
+ pfree(ioq);
+}
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index da6df2d3654..270c4a64428 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -7,6 +7,7 @@ backend_sources += files(
'aio_init.c',
'aio_io.c',
'aio_target.c',
+ 'io_queue.c',
'method_io_uring.c',
'method_sync.c',
'method_worker.c',
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index d216785c3c8..d084c476ec8 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1180,6 +1180,7 @@ IOContext
IOFuncSelector
IOObject
IOOp
+IOQueue
IO_STATUS_BLOCK
IPCompareMethod
ITEM
@@ -2993,6 +2994,7 @@ TocEntry
TokenAuxData
TokenizedAuthLine
TrackItem
+TrackedIO
TransApplyAction
TransInvalidationInfo
TransState
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0026-bufmgr-use-AIO-in-checkpointer-bgwriter.patch (31.1K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/27-v2.4-0026-bufmgr-use-AIO-in-checkpointer-bgwriter.patch)
download | inline diff:
From a08a59674ab0297ec89e8790adc4f7c65a0ab6f0 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 22 Jan 2025 13:44:52 -0500
Subject: [PATCH v2.4 26/29] bufmgr: use AIO in checkpointer, bgwriter
This is far from ready - just included to be able to exercise AIO writes and
get some preliminary numbers. In all likelihood this will instead be based
ontop of work by Thomas Munro instead of the preceding commit.
---
src/include/postmaster/bgwriter.h | 3 +-
src/include/storage/buf_internals.h | 2 +
src/include/storage/bufmgr.h | 3 +-
src/include/storage/bufpage.h | 1 +
src/backend/postmaster/bgwriter.c | 19 +-
src/backend/postmaster/checkpointer.c | 11 +-
src/backend/storage/buffer/bufmgr.c | 588 +++++++++++++++++++++++---
src/backend/storage/page/bufpage.c | 10 +
src/tools/pgindent/typedefs.list | 1 +
9 files changed, 580 insertions(+), 58 deletions(-)
diff --git a/src/include/postmaster/bgwriter.h b/src/include/postmaster/bgwriter.h
index 2d5854e6879..517c40cd804 100644
--- a/src/include/postmaster/bgwriter.h
+++ b/src/include/postmaster/bgwriter.h
@@ -31,7 +31,8 @@ extern void BackgroundWriterMain(char *startup_data, size_t startup_data_len) pg
extern void CheckpointerMain(char *startup_data, size_t startup_data_len) pg_attribute_noreturn();
extern void RequestCheckpoint(int flags);
-extern void CheckpointWriteDelay(int flags, double progress);
+struct IOQueue;
+extern void CheckpointWriteDelay(struct IOQueue *ioq, int flags, double progress);
extern bool ForwardSyncRequest(const FileTag *ftag, SyncRequestType type);
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 396642415bc..1d3a936837b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -21,6 +21,8 @@
#include "storage/buf.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
+#include "storage/io_queue.h"
+#include "storage/latch.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
#include "storage/smgr.h"
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 655885ff2d0..2cc7d47661f 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -304,7 +304,8 @@ extern bool ConditionalLockBufferForCleanup(Buffer buffer);
extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
-extern bool BgBufferSync(struct WritebackContext *wb_context);
+struct IOQueue;
+extern bool BgBufferSync(struct IOQueue *ioq, struct WritebackContext *wb_context);
extern uint32 GetSoftPinLimit(void);
extern uint32 GetSoftLocalPinLimit(void);
diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index 6646b6f6371..9c045e81857 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -509,5 +509,6 @@ extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
Item newtup, Size newsize);
extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern bool PageNeedsChecksumCopy(Page page);
#endif /* BUFPAGE_H */
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index ec1225c433f..1208926c0c9 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -38,11 +38,13 @@
#include "postmaster/auxprocess.h"
#include "postmaster/bgwriter.h"
#include "postmaster/interrupt.h"
+#include "storage/aio.h"
#include "storage/aio_subsys.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/fd.h"
+#include "storage/io_queue.h"
#include "storage/lwlock.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -90,6 +92,7 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
sigjmp_buf local_sigjmp_buf;
MemoryContext bgwriter_context;
bool prev_hibernate;
+ IOQueue *ioq;
WritebackContext wb_context;
Assert(startup_data_len == 0);
@@ -131,6 +134,7 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
ALLOCSET_DEFAULT_SIZES);
MemoryContextSwitchTo(bgwriter_context);
+ ioq = io_queue_create(128, 0);
WritebackContextInit(&wb_context, &bgwriter_flush_after);
/*
@@ -228,12 +232,22 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
/* Clear any already-pending wakeups */
ResetLatch(MyLatch);
+ /*
+ * FIXME: this is theoretically racy, but I didn't want to copy
+ * HandleMainLoopInterrupts() remaining body here.
+ */
+ if (ShutdownRequestPending)
+ {
+ io_queue_wait_all(ioq);
+ io_queue_free(ioq);
+ }
+
HandleMainLoopInterrupts();
/*
* Do one cycle of dirty-buffer writing.
*/
- can_hibernate = BgBufferSync(&wb_context);
+ can_hibernate = BgBufferSync(ioq, &wb_context);
/* Report pending statistics to the cumulative stats system */
pgstat_report_bgwriter();
@@ -250,6 +264,9 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
smgrdestroyall();
}
+ /* finish IO before sleeping, to avoid blocking other backends */
+ io_queue_wait_all(ioq);
+
/*
* Log a new xl_running_xacts every now and then so replication can
* get into a consistent state faster (think of suboverflowed
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index d254b2d1587..89baec0dd69 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -49,10 +49,12 @@
#include "postmaster/bgwriter.h"
#include "postmaster/interrupt.h"
#include "replication/syncrep.h"
+#include "storage/aio.h"
#include "storage/aio_subsys.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/fd.h"
+#include "storage/io_queue.h"
#include "storage/ipc.h"
#include "storage/lwlock.h"
#include "storage/pmsignal.h"
@@ -766,7 +768,7 @@ ImmediateCheckpointRequested(void)
* fraction between 0.0 meaning none, and 1.0 meaning all done.
*/
void
-CheckpointWriteDelay(int flags, double progress)
+CheckpointWriteDelay(IOQueue *ioq, int flags, double progress)
{
static int absorb_counter = WRITES_PER_ABSORB;
@@ -800,6 +802,13 @@ CheckpointWriteDelay(int flags, double progress)
/* Report interim statistics to the cumulative stats system */
pgstat_report_checkpointer();
+ /*
+ * Ensure all pending IO is submitted to avoid unnecessary delays for
+ * other processes.
+ */
+ io_queue_wait_all(ioq);
+
+
/*
* This sleep used to be connected to bgwriter_delay, typically 200ms.
* That resulted in more frequent wakeups if not much work to do.
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 65acc891ef1..8b835d962d9 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -52,6 +52,7 @@
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
+#include "storage/io_queue.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
#include "storage/proc.h"
@@ -77,6 +78,7 @@
/* Bits in SyncOneBuffer's return value */
#define BUF_WRITTEN 0x01
#define BUF_REUSABLE 0x02
+#define BUF_CANT_MERGE 0x04
#define RELS_BSEARCH_THRESHOLD 20
@@ -513,8 +515,6 @@ static void UnpinBuffer(BufferDesc *buf);
static void UnpinBufferNoOwner(BufferDesc *buf);
static void BufferSync(int flags);
static uint32 WaitBufHdrUnlocked(BufferDesc *buf);
-static int SyncOneBuffer(int buf_id, bool skip_recently_used,
- WritebackContext *wb_context);
static void WaitIO(BufferDesc *buf);
static void TerminateBufferIO(BufferDesc *buf, bool clear_dirty,
uint32 set_flag_bits, bool forget_owner,
@@ -533,6 +533,7 @@ static bool AsyncReadBuffers(ReadBuffersOperation *operation,
static Buffer GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context);
static void FlushBuffer(BufferDesc *buf, SMgrRelation reln,
IOObject io_object, IOContext io_context);
+
static void FindAndDropRelationBuffers(RelFileLocator rlocator,
ForkNumber forkNum,
BlockNumber nForkBlock,
@@ -3182,6 +3183,57 @@ UnpinBufferNoOwner(BufferDesc *buf)
}
}
+typedef struct BuffersToWrite
+{
+ int nbuffers;
+ BufferTag start_at_tag;
+ uint32 max_combine;
+
+ XLogRecPtr max_lsn;
+
+ PgAioHandle *ioh;
+ PgAioWaitRef iow;
+
+ uint64 total_writes;
+
+ Buffer buffers[IOV_MAX];
+ PgAioBounceBuffer *bounce_buffers[IOV_MAX];
+ const void *data_ptrs[IOV_MAX];
+} BuffersToWrite;
+
+static int PrepareToWriteBuffer(BuffersToWrite *to_write, Buffer buf,
+ bool skip_recently_used,
+ IOQueue *ioq, WritebackContext *wb_context);
+
+static void WriteBuffers(BuffersToWrite *to_write,
+ IOQueue *ioq, WritebackContext *wb_context);
+
+static void
+BuffersToWriteInit(BuffersToWrite *to_write,
+ IOQueue *ioq, WritebackContext *wb_context)
+{
+ to_write->total_writes = 0;
+ to_write->nbuffers = 0;
+ to_write->ioh = NULL;
+ pgaio_wref_clear(&to_write->iow);
+ to_write->max_lsn = InvalidXLogRecPtr;
+
+ pgaio_enter_batchmode();
+}
+
+static void
+BuffersToWriteEnd(BuffersToWrite *to_write)
+{
+ if (to_write->ioh != NULL)
+ {
+ pgaio_io_release(to_write->ioh);
+ to_write->ioh = NULL;
+ }
+
+ pgaio_exit_batchmode();
+}
+
+
#define ST_SORT sort_checkpoint_bufferids
#define ST_ELEMENT_TYPE CkptSortItem
#define ST_COMPARE(a, b) ckpt_buforder_comparator(a, b)
@@ -3213,7 +3265,10 @@ BufferSync(int flags)
binaryheap *ts_heap;
int i;
int mask = BM_DIRTY;
+ IOQueue *ioq;
WritebackContext wb_context;
+ BuffersToWrite to_write;
+ int max_combine;
/*
* Unless this is a shutdown checkpoint or we have been explicitly told,
@@ -3275,7 +3330,9 @@ BufferSync(int flags)
if (num_to_scan == 0)
return; /* nothing to do */
+ ioq = io_queue_create(512, 0);
WritebackContextInit(&wb_context, &checkpoint_flush_after);
+ max_combine = Min(io_bounce_buffers, io_combine_limit);
TRACE_POSTGRESQL_BUFFER_SYNC_START(NBuffers, num_to_scan);
@@ -3383,48 +3440,91 @@ BufferSync(int flags)
*/
num_processed = 0;
num_written = 0;
+
+ BuffersToWriteInit(&to_write, ioq, &wb_context);
+
while (!binaryheap_empty(ts_heap))
{
BufferDesc *bufHdr = NULL;
CkptTsStatus *ts_stat = (CkptTsStatus *)
DatumGetPointer(binaryheap_first(ts_heap));
+ bool batch_continue = true;
- buf_id = CkptBufferIds[ts_stat->index].buf_id;
- Assert(buf_id != -1);
-
- bufHdr = GetBufferDescriptor(buf_id);
-
- num_processed++;
+ Assert(ts_stat->num_scanned <= ts_stat->num_to_scan);
/*
- * We don't need to acquire the lock here, because we're only looking
- * at a single bit. It's possible that someone else writes the buffer
- * and clears the flag right after we check, but that doesn't matter
- * since SyncOneBuffer will then do nothing. However, there is a
- * further race condition: it's conceivable that between the time we
- * examine the bit here and the time SyncOneBuffer acquires the lock,
- * someone else not only wrote the buffer but replaced it with another
- * page and dirtied it. In that improbable case, SyncOneBuffer will
- * write the buffer though we didn't need to. It doesn't seem worth
- * guarding against this, though.
+ * Collect a batch of buffers to write out from the current
+ * tablespace. That causes some imbalance between the tablespaces, but
+ * that's more than outweighed by the efficiency gain due to batching.
*/
- if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+ while (batch_continue &&
+ to_write.nbuffers < max_combine &&
+ ts_stat->num_scanned < ts_stat->num_to_scan)
{
- if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
+ buf_id = CkptBufferIds[ts_stat->index].buf_id;
+ Assert(buf_id != -1);
+
+ bufHdr = GetBufferDescriptor(buf_id);
+
+ num_processed++;
+
+ /*
+ * We don't need to acquire the lock here, because we're only
+ * looking at a single bit. It's possible that someone else writes
+ * the buffer and clears the flag right after we check, but that
+ * doesn't matter since SyncOneBuffer will then do nothing.
+ * However, there is a further race condition: it's conceivable
+ * that between the time we examine the bit here and the time
+ * SyncOneBuffer acquires the lock, someone else not only wrote
+ * the buffer but replaced it with another page and dirtied it. In
+ * that improbable case, SyncOneBuffer will write the buffer
+ * though we didn't need to. It doesn't seem worth guarding
+ * against this, though.
+ */
+ if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
{
- TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id);
- PendingCheckpointerStats.buffers_written++;
- num_written++;
+ int result = PrepareToWriteBuffer(&to_write, buf_id + 1, false,
+ ioq, &wb_context);
+
+ if (result & BUF_CANT_MERGE)
+ {
+ Assert(to_write.nbuffers > 0);
+ WriteBuffers(&to_write, ioq, &wb_context);
+
+ result = PrepareToWriteBuffer(&to_write, buf_id + 1, false,
+ ioq, &wb_context);
+ Assert(result != BUF_CANT_MERGE);
+ }
+
+ if (result & BUF_WRITTEN)
+ {
+ TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id);
+ PendingCheckpointerStats.buffers_written++;
+ num_written++;
+ }
+ else
+ {
+ batch_continue = false;
+ }
}
+ else
+ {
+ if (to_write.nbuffers > 0)
+ WriteBuffers(&to_write, ioq, &wb_context);
+ }
+
+ /*
+ * Measure progress independent of actually having to flush the
+ * buffer - otherwise writing become unbalanced.
+ */
+ ts_stat->progress += ts_stat->progress_slice;
+ ts_stat->num_scanned++;
+ ts_stat->index++;
}
- /*
- * Measure progress independent of actually having to flush the buffer
- * - otherwise writing become unbalanced.
- */
- ts_stat->progress += ts_stat->progress_slice;
- ts_stat->num_scanned++;
- ts_stat->index++;
+ if (to_write.nbuffers > 0)
+ WriteBuffers(&to_write, ioq, &wb_context);
+
/* Have all the buffers from the tablespace been processed? */
if (ts_stat->num_scanned == ts_stat->num_to_scan)
@@ -3442,15 +3542,23 @@ BufferSync(int flags)
*
* (This will check for barrier events even if it doesn't sleep.)
*/
- CheckpointWriteDelay(flags, (double) num_processed / num_to_scan);
+ CheckpointWriteDelay(ioq, flags, (double) num_processed / num_to_scan);
}
+ Assert(to_write.nbuffers == 0);
+ io_queue_wait_all(ioq);
+
/*
* Issue all pending flushes. Only checkpointer calls BufferSync(), so
* IOContext will always be IOCONTEXT_NORMAL.
*/
IssuePendingWritebacks(&wb_context, IOCONTEXT_NORMAL);
+ io_queue_wait_all(ioq); /* IssuePendingWritebacks might have added
+ * more */
+ io_queue_free(ioq);
+ BuffersToWriteEnd(&to_write);
+
pfree(per_ts_stat);
per_ts_stat = NULL;
binaryheap_free(ts_heap);
@@ -3476,7 +3584,7 @@ BufferSync(int flags)
* bgwriter_lru_maxpages to 0.)
*/
bool
-BgBufferSync(WritebackContext *wb_context)
+BgBufferSync(IOQueue *ioq, WritebackContext *wb_context)
{
/* info obtained from freelist.c */
int strategy_buf_id;
@@ -3519,6 +3627,9 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ BuffersToWrite to_write;
+ int max_combine;
+
/*
* Find out where the freelist clock sweep currently is, and how many
* buffer allocations have happened since our last call.
@@ -3539,6 +3650,8 @@ BgBufferSync(WritebackContext *wb_context)
return true;
}
+ max_combine = Min(io_bounce_buffers, io_combine_limit);
+
/*
* Compute strategy_delta = how many buffers have been scanned by the
* clock sweep since last time. If first time through, assume none. Then
@@ -3695,11 +3808,25 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
+ BuffersToWriteInit(&to_write, ioq, wb_context);
+
/* Execute the LRU scan */
while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
{
- int sync_state = SyncOneBuffer(next_to_clean, true,
- wb_context);
+ int sync_state;
+
+ sync_state = PrepareToWriteBuffer(&to_write, next_to_clean + 1,
+ true, ioq, wb_context);
+ if (sync_state & BUF_CANT_MERGE)
+ {
+ Assert(to_write.nbuffers > 0);
+
+ WriteBuffers(&to_write, ioq, wb_context);
+
+ sync_state = PrepareToWriteBuffer(&to_write, next_to_clean + 1,
+ true, ioq, wb_context);
+ Assert(sync_state != BUF_CANT_MERGE);
+ }
if (++next_to_clean >= NBuffers)
{
@@ -3710,6 +3837,13 @@ BgBufferSync(WritebackContext *wb_context)
if (sync_state & BUF_WRITTEN)
{
+ Assert(sync_state & BUF_REUSABLE);
+
+ if (to_write.nbuffers == max_combine)
+ {
+ WriteBuffers(&to_write, ioq, wb_context);
+ }
+
reusable_buffers++;
if (++num_written >= bgwriter_lru_maxpages)
{
@@ -3721,6 +3855,11 @@ BgBufferSync(WritebackContext *wb_context)
reusable_buffers++;
}
+ if (to_write.nbuffers > 0)
+ WriteBuffers(&to_write, ioq, wb_context);
+
+ BuffersToWriteEnd(&to_write);
+
PendingBgWriterStats.buf_written_clean += num_written;
#ifdef BGW_DEBUG
@@ -3759,8 +3898,66 @@ BgBufferSync(WritebackContext *wb_context)
return (bufs_to_lap == 0 && recent_alloc == 0);
}
+static inline bool
+BufferTagsSameRel(const BufferTag *tag1, const BufferTag *tag2)
+{
+ return (tag1->spcOid == tag2->spcOid) &&
+ (tag1->dbOid == tag2->dbOid) &&
+ (tag1->relNumber == tag2->relNumber) &&
+ (tag1->forkNum == tag2->forkNum)
+ ;
+}
+
+static bool
+CanMergeWrite(BuffersToWrite *to_write, BufferDesc *cur_buf_hdr)
+{
+ BlockNumber cur_block = cur_buf_hdr->tag.blockNum;
+
+ Assert(to_write->nbuffers > 0); /* can't merge with nothing */
+ Assert(to_write->start_at_tag.relNumber != InvalidOid);
+ Assert(to_write->start_at_tag.blockNum != InvalidBlockNumber);
+
+ Assert(to_write->ioh != NULL);
+
+ /*
+ * First check if the blocknumber is one that we could actually merge,
+ * that's cheaper than checking the tablespace/db/relnumber/fork match.
+ */
+ if (to_write->start_at_tag.blockNum + to_write->nbuffers != cur_block)
+ return false;
+
+ if (!BufferTagsSameRel(&to_write->start_at_tag, &cur_buf_hdr->tag))
+ return false;
+
+ /*
+ * Need to check with smgr how large a write we're allowed to make. To
+ * reduce the overhead of the smgr check, only inquire once, when
+ * processing the first to-be-merged buffer. That avoids the overhead in
+ * the common case of writing out buffers that definitely not mergeable.
+ */
+ if (to_write->nbuffers == 1)
+ {
+ SMgrRelation smgr;
+
+ smgr = smgropen(BufTagGetRelFileLocator(&to_write->start_at_tag), INVALID_PROC_NUMBER);
+
+ to_write->max_combine = smgrmaxcombine(smgr,
+ to_write->start_at_tag.forkNum,
+ to_write->start_at_tag.blockNum);
+ }
+ else
+ {
+ Assert(to_write->max_combine > 0);
+ }
+
+ if (to_write->start_at_tag.blockNum + to_write->max_combine <= cur_block)
+ return false;
+
+ return true;
+}
+
/*
- * SyncOneBuffer -- process a single buffer during syncing.
+ * PrepareToWriteBuffer -- process a single buffer during syncing.
*
* If skip_recently_used is true, we don't write currently-pinned buffers, nor
* buffers marked recently used, as these are not replacement candidates.
@@ -3769,22 +3966,50 @@ BgBufferSync(WritebackContext *wb_context)
* BUF_WRITTEN: we wrote the buffer.
* BUF_REUSABLE: buffer is available for replacement, ie, it has
* pin count 0 and usage count 0.
+ * BUF_CANT_MERGE: can't combine this write with prior writes, caller needs
+ * to issue those first
*
* (BUF_WRITTEN could be set in error if FlushBuffer finds the buffer clean
* after locking it, but we don't care all that much.)
*/
static int
-SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
+PrepareToWriteBuffer(BuffersToWrite *to_write, Buffer buf,
+ bool skip_recently_used,
+ IOQueue *ioq, WritebackContext *wb_context)
{
- BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
+ BufferDesc *cur_buf_hdr = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
int result = 0;
- uint32 buf_state;
- BufferTag tag;
+ XLogRecPtr cur_buf_lsn;
+ LWLock *content_lock;
+ bool may_block;
+
+ /*
+ * Check if this buffer can be written out together with already prepared
+ * writes. We check before we have pinned the buffer, so the buffer can be
+ * written out and replaced between this check and us pinning the buffer -
+ * we'll recheck below. The reason for the pre-check is that we don't want
+ * to pin the buffer just to find out that we can't merge the IO.
+ */
+ if (to_write->nbuffers != 0)
+ {
+ if (!CanMergeWrite(to_write, cur_buf_hdr))
+ {
+ result |= BUF_CANT_MERGE;
+ return result;
+ }
+ }
+ else
+ {
+ to_write->start_at_tag = cur_buf_hdr->tag;
+ }
/* Make sure we can handle the pin */
ReservePrivateRefCountEntry();
ResourceOwnerEnlarge(CurrentResourceOwner);
+ /* XXX: Should also check if we are allowed to pin one more buffer */
+
/*
* Check whether buffer needs writing.
*
@@ -3794,7 +4019,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
* don't worry because our checkpoint.redo points before log record for
* upcoming changes and so we are not required to write such dirty buffer.
*/
- buf_state = LockBufHdr(bufHdr);
+ buf_state = LockBufHdr(cur_buf_hdr);
if (BUF_STATE_GET_REFCOUNT(buf_state) == 0 &&
BUF_STATE_GET_USAGECOUNT(buf_state) == 0)
@@ -3803,40 +4028,294 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
}
else if (skip_recently_used)
{
+#if 0
+ elog(LOG, "at block %d: skip recent with nbuffers %d",
+ cur_buf_hdr->tag.blockNum, to_write->nbuffers);
+#endif
/* Caller told us not to write recently-used buffers */
- UnlockBufHdr(bufHdr, buf_state);
+ UnlockBufHdr(cur_buf_hdr, buf_state);
return result;
}
if (!(buf_state & BM_VALID) || !(buf_state & BM_DIRTY))
{
/* It's clean, so nothing to do */
- UnlockBufHdr(bufHdr, buf_state);
+ UnlockBufHdr(cur_buf_hdr, buf_state);
return result;
}
+ /* pin the buffer, from now on its identity can't change anymore */
+ PinBuffer_Locked(cur_buf_hdr);
+
+ /*
+ * Acquire IO, if needed, now that it's likely that we'll need to write.
+ */
+ if (to_write->ioh == NULL)
+ {
+ /* otherwise we should already have acquired a handle */
+ Assert(to_write->nbuffers == 0);
+
+ to_write->ioh = io_queue_acquire_io(ioq);
+ pgaio_io_get_wref(to_write->ioh, &to_write->iow);
+ }
+
/*
- * Pin it, share-lock it, write it. (FlushBuffer will do nothing if the
- * buffer is clean by the time we've locked it.)
+ * If we are merging, check if the buffer's identity possibly changed
+ * while we hadn't yet pinned it.
+ *
+ * XXX: It might be worth checking if we still want to write the buffer
+ * out, e.g. it could have been replaced with a buffer that doesn't have
+ * BM_CHECKPOINT_NEEDED set.
*/
- PinBuffer_Locked(bufHdr);
- LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
+ if (to_write->nbuffers != 0)
+ {
+ if (!CanMergeWrite(to_write, cur_buf_hdr))
+ {
+ elog(LOG, "changed identity");
+ UnpinBuffer(cur_buf_hdr);
+
+ result |= BUF_CANT_MERGE;
+
+ return result;
+ }
+ }
+
+ may_block = to_write->nbuffers == 0
+ && !pgaio_have_staged()
+ && io_queue_is_empty(ioq)
+ ;
+ content_lock = BufferDescriptorGetContentLock(cur_buf_hdr);
+
+ if (!may_block)
+ {
+ if (LWLockConditionalAcquire(content_lock, LW_SHARED))
+ {
+ /* done */
+ }
+ else if (to_write->nbuffers == 0)
+ {
+ /*
+ * Need to wait for all prior IO to finish before blocking for
+ * lock acquisition, to avoid the risk a deadlock due to us
+ * waiting for another backend that is waiting for our unsubmitted
+ * IO to complete.
+ */
+ pgaio_submit_staged();
+ io_queue_wait_all(ioq);
+
+ elog(DEBUG2, "at block %u: can't block, nbuffers = 0",
+ cur_buf_hdr->tag.blockNum
+ );
+
+ may_block = to_write->nbuffers == 0
+ && !pgaio_have_staged()
+ && io_queue_is_empty(ioq)
+ ;
+ Assert(may_block);
+
+ LWLockAcquire(content_lock, LW_SHARED);
+ }
+ else
+ {
+ elog(DEBUG2, "at block %d: can't block nbuffers = %d",
+ cur_buf_hdr->tag.blockNum,
+ to_write->nbuffers);
- FlushBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
+ UnpinBuffer(cur_buf_hdr);
+ result |= BUF_CANT_MERGE;
+ Assert(to_write->nbuffers > 0);
- LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+ return result;
+ }
+ }
+ else
+ {
+ LWLockAcquire(content_lock, LW_SHARED);
+ }
- tag = bufHdr->tag;
+ if (!may_block)
+ {
+ if (!StartBufferIO(cur_buf_hdr, false, !may_block))
+ {
+ pgaio_submit_staged();
+ io_queue_wait_all(ioq);
- UnpinBuffer(bufHdr);
+ may_block = io_queue_is_empty(ioq) && to_write->nbuffers == 0 && !pgaio_have_staged();
+
+ if (!StartBufferIO(cur_buf_hdr, false, !may_block))
+ {
+ elog(DEBUG2, "at block %d: non-waitable StartBufferIO returns false, %d",
+ cur_buf_hdr->tag.blockNum,
+ may_block);
+
+ /*
+ * FIXME: can't tell whether this is because the buffer has
+ * been cleaned
+ */
+ if (!may_block)
+ {
+ result |= BUF_CANT_MERGE;
+ Assert(to_write->nbuffers > 0);
+ }
+ LWLockRelease(content_lock);
+ UnpinBuffer(cur_buf_hdr);
+
+ return result;
+ }
+ }
+ }
+ else
+ {
+ if (!StartBufferIO(cur_buf_hdr, false, false))
+ {
+ elog(DEBUG2, "waitable StartBufferIO returns false");
+ LWLockRelease(content_lock);
+ UnpinBuffer(cur_buf_hdr);
+
+ /*
+ * FIXME: Historically we returned BUF_WRITTEN in this case, which
+ * seems wrong
+ */
+ return result;
+ }
+ }
/*
- * SyncOneBuffer() is only called by checkpointer and bgwriter, so
- * IOContext will always be IOCONTEXT_NORMAL.
+ * Run PageGetLSN while holding header lock, since we don't have the
+ * buffer locked exclusively in all cases.
*/
- ScheduleBufferTagForWriteback(wb_context, IOCONTEXT_NORMAL, &tag);
+ buf_state = LockBufHdr(cur_buf_hdr);
+
+ cur_buf_lsn = BufferGetLSN(cur_buf_hdr);
+
+ /* To check if block content changes while flushing. - vadim 01/17/97 */
+ buf_state &= ~BM_JUST_DIRTIED;
+
+ UnlockBufHdr(cur_buf_hdr, buf_state);
+
+ to_write->buffers[to_write->nbuffers] = buf;
+ to_write->nbuffers++;
+
+ if (buf_state & BM_PERMANENT &&
+ (to_write->max_lsn == InvalidXLogRecPtr || to_write->max_lsn < cur_buf_lsn))
+ {
+ to_write->max_lsn = cur_buf_lsn;
+ }
+
+ result |= BUF_WRITTEN;
+
+ return result;
+}
+
+static void
+WriteBuffers(BuffersToWrite *to_write,
+ IOQueue *ioq, WritebackContext *wb_context)
+{
+ SMgrRelation smgr;
+ Buffer first_buf;
+ BufferDesc *first_buf_hdr;
+ bool needs_checksum;
+
+ Assert(to_write->nbuffers > 0 && to_write->nbuffers <= io_combine_limit);
+
+ first_buf = to_write->buffers[0];
+ first_buf_hdr = GetBufferDescriptor(first_buf - 1);
+
+ smgr = smgropen(BufTagGetRelFileLocator(&first_buf_hdr->tag), INVALID_PROC_NUMBER);
+
+ /*
+ * Force XLOG flush up to buffer's LSN. This implements the basic WAL
+ * rule that log updates must hit disk before any of the data-file changes
+ * they describe do.
+ *
+ * However, this rule does not apply to unlogged relations, which will be
+ * lost after a crash anyway. Most unlogged relation pages do not bear
+ * LSNs since we never emit WAL records for them, and therefore flushing
+ * up through the buffer LSN would be useless, but harmless. However,
+ * GiST indexes use LSNs internally to track page-splits, and therefore
+ * unlogged GiST pages bear "fake" LSNs generated by
+ * GetFakeLSNForUnloggedRel. It is unlikely but possible that the fake
+ * LSN counter could advance past the WAL insertion point; and if it did
+ * happen, attempting to flush WAL through that location would fail, with
+ * disastrous system-wide consequences. To make sure that can't happen,
+ * skip the flush if the buffer isn't permanent.
+ */
+ if (to_write->max_lsn != InvalidXLogRecPtr)
+ XLogFlush(to_write->max_lsn);
+
+ /*
+ * Now it's safe to write buffer to disk. Note that no one else should
+ * have been able to write it while we were busy with log flushing because
+ * only one process at a time can set the BM_IO_IN_PROGRESS bit.
+ */
+
+ for (int nbuf = 0; nbuf < to_write->nbuffers; nbuf++)
+ {
+ Buffer cur_buf = to_write->buffers[nbuf];
+ BufferDesc *cur_buf_hdr = GetBufferDescriptor(cur_buf - 1);
+ Block bufBlock;
+ char *bufToWrite;
+
+ bufBlock = BufHdrGetBlock(cur_buf_hdr);
+ needs_checksum = PageNeedsChecksumCopy((Page) bufBlock);
+
+ /*
+ * Update page checksum if desired. Since we have only shared lock on
+ * the buffer, other processes might be updating hint bits in it, so
+ * we must copy the page to a bounce buffer if we do checksumming.
+ */
+ if (needs_checksum)
+ {
+ PgAioBounceBuffer *bb = pgaio_bounce_buffer_get();
+
+ pgaio_io_assoc_bounce_buffer(to_write->ioh, bb);
+
+ bufToWrite = pgaio_bounce_buffer_buffer(bb);
+ memcpy(bufToWrite, bufBlock, BLCKSZ);
+ PageSetChecksumInplace((Page) bufToWrite, cur_buf_hdr->tag.blockNum);
+ }
+ else
+ {
+ bufToWrite = bufBlock;
+ }
+
+ to_write->data_ptrs[nbuf] = bufToWrite;
+ }
+
+ pgaio_io_set_handle_data_32(to_write->ioh,
+ (uint32 *) to_write->buffers,
+ to_write->nbuffers);
+ pgaio_io_register_callbacks(to_write->ioh, PGAIO_HCB_SHARED_BUFFER_WRITEV);
+
+ smgrstartwritev(to_write->ioh, smgr,
+ BufTagGetForkNum(&first_buf_hdr->tag),
+ first_buf_hdr->tag.blockNum,
+ to_write->data_ptrs,
+ to_write->nbuffers,
+ false);
+ pgstat_count_io_op(IOOBJECT_RELATION, IOCONTEXT_NORMAL,
+ IOOP_WRITE, 1, BLCKSZ * to_write->nbuffers);
+
+
+ for (int nbuf = 0; nbuf < to_write->nbuffers; nbuf++)
+ {
+ Buffer cur_buf = to_write->buffers[nbuf];
+ BufferDesc *cur_buf_hdr = GetBufferDescriptor(cur_buf - 1);
+
+ UnpinBuffer(cur_buf_hdr);
+ }
+
+ io_queue_track(ioq, &to_write->iow);
+ to_write->total_writes++;
- return result | BUF_WRITTEN;
+ /* clear state for next write */
+ to_write->nbuffers = 0;
+ to_write->start_at_tag.relNumber = InvalidOid;
+ to_write->start_at_tag.blockNum = InvalidBlockNumber;
+ to_write->max_combine = 0;
+ to_write->max_lsn = InvalidXLogRecPtr;
+ to_write->ioh = NULL;
+ pgaio_wref_clear(&to_write->iow);
}
/*
@@ -4212,6 +4691,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
error_context_stack = errcallback.previous;
}
+
/*
* RelationGetNumberOfBlocksInFork
* Determines the current number of pages in the specified relation fork.
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index 91da73dda8b..c4a78dc96d2 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1480,6 +1480,16 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
return true;
}
+bool
+PageNeedsChecksumCopy(Page page)
+{
+ if (PageIsNew(page))
+ return false;
+
+ /* If we don't need a checksum, just return the passed-in data */
+ return DataChecksumsEnabled();
+}
+
/*
* Set checksum for a page in shared buffers.
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index d084c476ec8..d1d8758566e 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -348,6 +348,7 @@ BufferManagerRelation
BufferStrategyControl
BufferTag
BufferUsage
+BuffersToWrite
BuildAccumulator
BuiltinScript
BulkInsertState
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0027-Temporary-Increase-BAS_BULKREAD-size.patch (1.3K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/28-v2.4-0027-Temporary-Increase-BAS_BULKREAD-size.patch)
download | inline diff:
From 7ed7dc075a64fb31d68218324b1143ce74794674 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sun, 1 Sep 2024 00:42:27 -0400
Subject: [PATCH v2.4 27/29] Temporary: Increase BAS_BULKREAD size
Without this we only can execute very little AIO for sequential scans, as
there's just not enough buffers in the ring. This isn't the right fix, as
just increasing the ring size can have negative performance implications in
workloads where the kernel has all the data cached.
Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
src/backend/storage/buffer/freelist.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 336715b6c63..b72a5957a20 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -555,7 +555,12 @@ GetAccessStrategy(BufferAccessStrategyType btype)
return NULL;
case BAS_BULKREAD:
- ring_size_kb = 256;
+
+ /*
+ * FIXME: Temporary increase to allow large enough streaming reads
+ * to actually benefit from AIO. This needs a better solution.
+ */
+ ring_size_kb = 2 * 1024;
break;
case BAS_BULKWRITE:
ring_size_kb = 16 * 1024;
--
2.48.1.76.g4e746b1a31.dirty
[text/x-diff] v2.4-0028-WIP-Use-MAP_POPULATE.patch (1.1K, ../../clt7rl56kxjcnjtqd7fsajkst232c3yh57ggtmppwp5hmtl4os@i3iibeftfrsp/29-v2.4-0028-WIP-Use-MAP_POPULATE.patch)
download | inline diff:
From 3c501ebbb63c591688546b1e3ff62aaf30b05bc0 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 31 Dec 2024 13:25:56 -0500
Subject: [PATCH v2.4 28/29] WIP: Use MAP_POPULATE
For benchmarking it's quite annoying that the first time a memory is touched
has completely different perf characteristics than subsequent accesses. Using
MAP_POPULATE reduces that substantially.
---
src/backend/port/sysv_shmem.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index 197926d44f6..a700b02d5a1 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -620,7 +620,7 @@ CreateAnonymousSegment(Size *size)
allocsize += hugepagesize - (allocsize % hugepagesize);
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS | MAP_POPULATE | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
--
2.48.1.76.g4e746b1a31.dirty
^ permalink raw reply [nested|flat] 2+ messages in thread
* Re: AIO v2.4
2025-02-19 19:10 Re: AIO v2.4 Andres Freund <andres@anarazel.de>
@ 2025-02-24 10:50 ` Andres Freund <andres@anarazel.de>
0 siblings, 0 replies; 2+ messages in thread
From: Andres Freund @ 2025-02-24 10:50 UTC (permalink / raw)
To: pgsql-hackers; +Cc: Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Jakub Wartak <jakub.wartak@enterprisedb.com>
Hi,
On 2025-02-19 14:10:44 -0500, Andres Freund wrote:
> I'm planning to push the first two commits soon, I think they're ok on their
> own, even if nothing else were to go in.
I did that for the lwlock patch.
But I think I might not do the same for the "Ensure a resowner exists for all
paths that may perform AIO" patch. The paths for which we are missing
resowners are concerned WAL writes - but it'll be a while before we get
AIO WAL writes.
It'd be fairly harmless to do this change before, but I found the justifying
code comments hard to rephrase. E.g.:
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -361,8 +361,15 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
BaseInit();
bootstrap_signals();
+
+ /* need a resowner for IO during BootStrapXLOG() */
+ CreateAuxProcessResourceOwner();
+
BootStrapXLOG(bootstrap_data_checksum_version);
+ ReleaseAuxProcessResources(true);
+ CurrentResourceOwner = NULL;
+
/*
* To ensure that src/common/link-canary.c is linked into the backend, we
* must call it from somewhere. Here is as good as anywhere.
Given that there's no use of resowners inside BootStrapXLOG() today and not
for the next months it seems confusing?
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 2+ messages in thread
end of thread, other threads:[~2025-02-24 10:50 UTC | newest]
Thread overview: 2+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2025-02-19 19:10 Re: AIO v2.4 Andres Freund <andres@anarazel.de>
2025-02-24 10:50 ` Andres Freund <andres@anarazel.de>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox